Model Compression Method, Device, Electronic Device and Storage Medium

By replacing floating-point arithmetic with matrix quantization and inverse quantization of deep learning models, the problem of insufficient computing resources of end devices is solved, and the model volume compression and inference speed is improved, which is suitable for a variety of model structures and architectures.

CN112529189BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011247207.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-10
Publication Date
2025-07-22
Estimated Expiration
2040-11-10

AI Technical Summary

Technical Problem

The inference computing requirements of deep learning models on end devices do not match hardware resources, and the existing model compression methods have poor operability and repeatability, which limits their application scenarios.

Method used

During the model inference process, the left matrix in the two multiplied matrices is quantized by rows and the right matrix is quantized by columns to obtain fixed-point operation results, and the final matrix operation result is obtained through inverse quantization, replacing floating-point operation.

Benefits of technology

It effectively compresses the model volume, improves the model inference speed, and is suitable for various model structures and architectures, with universal applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112529189B_ABST
    Figure CN112529189B_ABST
Patent Text Reader

Abstract

The present application discloses a model compression method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence such as deep learning and speech recognition. The method may include: when matrix operations of single-precision floating-point numbers are required during the model inference process, quantizing the left matrix in the two multiplied matrices row by row to obtain a first quantized matrix, and quantizing the right matrix in the two multiplied matrices column by column to obtain a second quantized matrix; multiplying the first quantized matrix and the second quantized matrix to obtain a third matrix as the result of fixed-point operations; performing dequantization according to the third matrix to obtain a fourth matrix, and using the fourth matrix as the result of the matrix operation. Applying the solution of the present application can improve the model inference speed and has general applicability, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to model compression methods, devices, electronic devices, and storage media in the fields of deep learning and speech recognition, etc. Background Art

[0002] With the development of technology, deep learning models have been more and more widely applied. In order to continuously improve the model accuracy, the depth and volume of the models are continuously increasing. Taking speech recognition as an example, from feedforward deep neural networks to recurrent neural networks, and then to encoder-decoder models, each technological change has brought greater computational requirements for model inference.

[0003] Currently, the deployment of deep learning applications is gradually migrating from cloud servers to end devices. Although the computing performance of end devices is also constantly improving, in actual applications, it still often fails to meet the requirements, and the mismatch problem between model inference and the hardware resources of end devices needs to be solved urgently.

[0004] To solve the above problems, the following implementation methods are mostly adopted: on the basis of the original large model, a more streamlined model structure is obtained by reducing the number of network nodes or connections, so as to achieve model compression. However, this method is closely related to the model structure, and its operability and repeatability are poor, thus limiting its application scenarios, etc. Summary of the Invention

[0005] This application provides a model compression method, device, electronic device, and storage medium.

[0006] A model compression method includes:

[0007] When matrix operations of single-precision floating points are required during the model inference process, the left matrix in the two multiplied matrices is quantized row by row to obtain a first quantization matrix, and the right matrix in the two multiplied matrices is quantized column by column to obtain a second quantization matrix;

[0008] Multiply the first quantization matrix and the second quantization matrix to obtain a third matrix as the result of fixed-point operation;

[0009] According to the third matrix, perform dequantization to obtain a fourth matrix, and use the fourth matrix as the result of the matrix operation.

[0010] A model compression device includes: a quantization module, an operation module, and a dequantization module;

[0011] The quantization module is configured to, when matrix operations in single-precision floating point are required during model inference, quantize the left matrix among the two matrices to be multiplied row by row to obtain a first quantized matrix, and quantize the right matrix among the two matrices to be multiplied column by column to obtain a second quantized matrix;

[0012] The operation module is configured to multiply the first quantized matrix and the second quantized matrix to obtain a third matrix as the result of fixed-point operation;

[0013] The dequantization module is configured to perform dequantization based on the third matrix to obtain a fourth matrix, and use the fourth matrix as the result of the matrix operation.

[0014] An electronic device includes:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0018] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above.

[0019] One embodiment of the above application has the following advantages or beneficial effects: Through quantization processing, single-precision floating point can be converted to fixed-point, so that during matrix operations in the model inference process, fixed-point operations can replace floating-point operations, thereby effectively compressing the model volume and improving the model inference speed. Moreover, it can be applied to various different model structures and has general applicability, etc.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings are used to better understand the solution and do not constitute a limitation to this application. Among them:

[0022] Figure 1 is a flowchart of an embodiment of the model compression method described in this application;

[0023] Figure 2 is a schematic diagram of the quantization fine-tuning training process described in this application;

[0024] Figure 3Schematic diagram of the composition structure of the model compression device 30 according to an embodiment of the present application;

[0025] Figure 4 Block diagram of an electronic device for the method according to an embodiment of the present application. Detailed implementation manners

[0026] The following describes exemplary embodiments of the present application with reference to the accompanying drawings. Various details of the embodiments of the present application are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0027] In addition, it should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0028] Figure 1 Flowchart of the method embodiment of the model compression method according to the present application. As Figure 1 shown, it includes the following specific implementation manners.

[0029] In step 101, when matrix operations in single-precision floating point are required during the model inference process, the left matrix in the two matrices to be multiplied is quantized row by row to obtain a first quantization matrix, and the right matrix in the two matrices to be multiplied is quantized column by column to obtain a second quantization matrix.

[0030] In step 102, the first quantization matrix and the second quantization matrix are multiplied to obtain a third matrix as the result of fixed-point operation.

[0031] In step 103, dequantization is performed according to the third matrix to obtain a fourth matrix, and the fourth matrix is used as the result of the matrix operation.

[0032] Matrix operations are the most core and costly operations during the model inference process. In the solution described in the above method embodiment, through quantization processing, single-precision floating point can be converted to fixed-point, so that during matrix operations in the model inference process, fixed-point operations can replace floating-point operations, thereby effectively compressing the model volume and improving the model inference speed. Moreover, it can be applied to various different model structures and various different architectures, such as Advanced RISC Machine (ARM) and X86, etc., and has general applicability.

[0033] The matrix operation generally refers to the operation of multiplying two matrices. The left matrix in the two matrices to be multiplied can be quantized row by row to obtain a first quantization matrix, and the right matrix in the two matrices to be multiplied can be quantized column by column to obtain a second quantization matrix. For example, for the matrix operation W*X, W is the left matrix and X is the right matrix.

[0034] For the left matrix, it can be quantized row by row to obtain a first quantization matrix. Preferably, the reference value corresponding to each row element in the left matrix can be determined respectively, and for each element in the left matrix, the quantization value of the element can be determined respectively according to the reference value corresponding to the row where the element is located and a predetermined bit width. The values of the elements in the first quantization matrix are all quantization values.

[0035] Among them, for each row of elements in the left matrix, the maximum value among the absolute values of the values of the elements included therein can be used as the reference value corresponding to that row. That is, for each row of elements in the left matrix, fabsmax(w i,* ) can be used as the reference value corresponding to that row, and fabsmax(w i,* ) represents the maximum value among the absolute values of the values of the elements included in the i-th (representing any row) row. For example, if there are 10 elements in the i-th row, and the absolute values of the values of these 10 elements are absolute value 1 - absolute value 10 respectively, and the value of absolute value 6 is the largest, then absolute value 6 can be used as the reference value corresponding to that row.

[0036] After that, for each element in the left matrix, the quotient of the value of the element and the reference value corresponding to the row where the element is located can be calculated respectively, and the product of the calculated quotient and 2 B-1 is calculated, where B represents the predetermined bit width, and then the obtained product can be used as the quantization value of the element.

[0037] For example, for element a, there is:

[0038]

[0039] Among them, w i,j represents the value of element a, fabsmax(w i,* ) represents the reference value corresponding to the row where element a is located, w′ i,j represents the quantization value of element a, and the specific value of B can be determined according to actual needs, such as 8 or 4 or other values, etc.

[0040] For the right matrix, it can be quantized column by column to obtain a second quantization matrix. Preferably, the reference value corresponding to each column element in the right matrix can be determined respectively, and for each element in the right matrix, the quantization value of the element can be determined respectively according to the reference value corresponding to the column where the element is located and a predetermined bit width. The values of the elements in the second quantization matrix are all quantization values.

[0041] Among them, for each column element in the right matrix, the maximum value among the absolute values of the values of each included element can be used as the reference value corresponding to that column. That is, for each column element in the right matrix, the absolute value maximum of the values of each included element can be used as the reference value corresponding to that column. For fabsmax(X *,j ), it is used as the reference value corresponding to that column. fabsmax(X *,j ) represents the maximum value among the absolute values of the values of each included element in the j-th (representing any column) column.

[0042] After that, for each element in the right matrix, the quotient of the value of the element and the reference value corresponding to the column where the element is located can be calculated respectively, and the product of the obtained quotient and 2 B-1 is calculated, and then the obtained product can be used as the quantization value of the element.

[0043] For example, for the element b, there is:

[0044]

[0045] Among them, x i,j represents the value of the element b, fabsmax(x *,j ) represents the reference value corresponding to the column where the element b is located, x′ i,j represents the quantization value of the element b, and the specific value of B can be determined according to actual needs.

[0046] After obtaining the first quantization matrix and the second quantization matrix respectively in the above manner, the first quantization matrix and the second quantization matrix can be multiplied to obtain the third matrix as the result of fixed-point operation.

[0047] That is: O′ = W′ * X′; (3)

[0048] Among them, W′ represents the first quantization matrix, X′ represents the second quantization matrix, and O′ represents the third matrix.

[0049] Through quantization processing, single-precision floating-point is converted to fixed-point, so that in matrix operations, fixed-point operations replace floating-point operations, thereby effectively compressing the model volume and improving the model inference speed, etc.

[0050] Due to the quantization processing, correspondingly, it is also necessary to perform dequantization according to the third matrix to obtain the fourth matrix, and use the fourth matrix as the result of the final required matrix operation.

[0051] Preferably, for each element in the fourth matrix, the following processing can be performed respectively: calculate the product of the value of the corresponding element of the element in the third matrix, the reference value corresponding to the row where the element is located, and the reference value corresponding to the column where the element is located, and use the obtained product as the value of the element, and the corresponding element is the element in the same position.

[0052] For example, for element c, there is:

[0053] o i,j = o' i,j *fabsmax(w i,* )*fabsmax(x *,j ); (4)

[0054] Wherein, o' i,j represents the value of element c in the third matrix, fabsmax(w i,* ) represents the reference value corresponding to the row where element c is located, fabsmax(x *,j ) represents the reference value corresponding to the column where element c is located, and o i,j represents the value of element c in the fourth matrix.

[0055] During model inference, the above left matrix is usually the weight matrix and is fixed. Therefore, the quantization process for the left matrix can be completed offline. For the right matrix, it is related to the input. Therefore, the quantization process for the right matrix needs to be completed online.

[0056] As mentioned above, the specific value of the bit width B can be determined according to actual needs, such as 8 or 4 or other values, etc. In this application, a symmetric fixed-point representation method can be adopted. The representation range of the bit width B is (-2 B-1 , 2 B-1 ). Taking int4 as an example, its representation range can be [-7, 7]. If int4 is adopted, compared with the single-precision floating-point fp32 representation, it can be compressed by 8 times.

[0057] The model compression method described in this application can also be used in combination with other compression methods. For example, after compressing the model according to the model compression method described in this application, the Huffman compression method can also be adopted to achieve further compression of the model, etc.

[0058] As the representation bit width decreases, the gap between the fixed-point model accuracy and the floating-point model accuracy gradually increases. The low-bit model needs to be fine-tuned to obtain the original model accuracy. For this reason, in this application, a progressive low-precision model training method is further proposed to improve the model accuracy, stability, and reproducibility, etc.

[0059] Preferably, a single-precision model can be first trained according to the single-precision training method as the initial model. Then, the model parameters of the initial model can be quantized and fine-tuned to obtain the final model. Using the final model, model inference and other processes can be performed.

[0060] First, a tuned single-precision model can be obtained according to the existing normal single-precision training method as an initial model, which serves as the basis for subsequent quantization fine-tuning training.

[0061] For the model parameters of the initial model, the following first processing can be performed: Quantize the weight matrix (i.e., the aforementioned left matrix) in the model parameters row by row, and dequantize the quantization result to obtain the processed model parameters; perform forward calculation and backward calculation based on the processed model parameters to obtain the model parameter gradients; update the model parameters according to the model parameter gradients, and for the updated model parameters, repeat the first processing until a predetermined end condition is met.

[0062] Based on the above introduction, Figure 2 This is a schematic diagram of the quantization fine-tuning training process described in this application. As Figure 2 shown, before the start of training for each batch of training data, first, for the latest model parameters, quantize the weight matrix row by row, that is, respectively determine the reference value corresponding to each row element in the weight matrix, and for each element in the weight matrix, respectively determine the quantization value of the element according to the reference value corresponding to the row where the element is located and the predetermined bit width. Among them, for each row element in the weight matrix, the maximum value of the absolute values of the values of the elements included therein can be used as the reference value corresponding to that row. In addition, for each element in the weight matrix, the quotient of the value of the element and the reference value corresponding to the row where the element is located can be calculated respectively, and the product of the calculated quotient and 2 B-1 is used as the quantization value of the element, where B represents the bit width. Then, the floating-point parameter value can be obtained through dequantization to obtain the processed model parameters. Then, forward calculation and backward calculation can be performed in sequence in combination with the processed model parameters and training data, etc., to obtain the model parameter gradients. Further, the model parameters can be updated according to the model parameter gradients, and for the updated model parameters, the above processing can be repeated until a predetermined end condition is met.

[0063] In the above process, how to perform dequantization, how to obtain the model parameter gradients, and how to update the model parameters according to the model parameter gradients are all prior arts. In addition, in practical applications, preferably, the above process can be sequentially performed in the order from the input layer to the output layer until a predetermined end condition is met.

[0064] What the specific predetermined end condition is can be determined according to actual needs. For example, it can refer to that the model accuracy or the model volume reaches the expectation, etc.

[0065] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0066] The above is the introduction to the method embodiments. The following further illustrates the solution of this application through device embodiments.

[0067] Figure 3 It is a schematic structural diagram of the composition of the model compression device 30 embodiment of this application. As Figure 3 shown, it includes: a quantization module 301, an operation module 302, and an inverse quantization module 303.

[0068] The quantization module 301 is used to quantize the left matrix in the two matrices to be multiplied row by row when single-precision floating-point matrix operations are required during the model inference process, to obtain a first quantization matrix, and quantize the right matrix in the two matrices to be multiplied column by column, to obtain a second quantization matrix.

[0069] The operation module 302 is used to multiply the first quantization matrix and the second quantization matrix to obtain a third matrix as the result of fixed-point operations.

[0070] The inverse quantization module 303 is used to perform inverse quantization according to the third matrix to obtain a fourth matrix, and use the fourth matrix as the result of the matrix operation.

[0071] Among them, when quantizing the left matrix, the quantization module 301 can respectively determine the reference value corresponding to each row element in the left matrix, and for each element in the left matrix, respectively determine the quantization value of the element according to the reference value corresponding to the row where the element is located and the predetermined bit width.

[0072] When quantizing the right matrix, the quantization module 301 can respectively determine the reference value corresponding to each column element in the right matrix, and for each element in the right matrix, respectively determine the quantization value of the element according to the reference value corresponding to the column where the element is located and the predetermined bit width.

[0073] Preferably, the quantization module 301 can respectively use the maximum value among the absolute values of the values of each element included in each row element in the left matrix as the reference value.

[0074] Similarly, the quantization module 301 can respectively use the maximum value among the absolute values of the values of each element included in each column element in the right matrix as the reference value.

[0075] In addition, for each element in the left matrix, the quantization module 301 can separately calculate the quotient of the value of the element and the reference value corresponding to the row where the element is located, and calculate the product of the quotient and 2 B-1 The product is used as the quantization value of the element, where B represents a predetermined bit width.

[0076] Similarly, for each element in the right matrix, the quantization module 301 can separately calculate the quotient of the value of the element and the reference value corresponding to the column where the element is located, and calculate the product of the quotient and 2 B-1 The product is used as the quantization value of the element.

[0077] The operation module 302 can multiply the first quantization matrix and the second quantization matrix to obtain a third matrix as the fixed-point operation result.

[0078] After that, for each element in the fourth matrix, the dequantization module 303 can separately perform the following processing: calculate the product of the value of the corresponding element of the element in the third matrix, the reference value corresponding to the row where the element is located, and the reference value corresponding to the column where the element is located, and use the product as the value of the element, where the corresponding element is the element in the same position.

[0079] As Figure 3 shown, the device may further include: a preprocessing module 300, configured to train a single-precision model in a single-precision training manner as an initial model, and perform quantization fine-tuning training on the model parameters of the initial model to obtain a final model.

[0080] Generally speaking, the above left matrix is a weight matrix. The preprocessing module 300 can perform the following first processing on the model parameters: perform quantization on the weight matrix in the model parameters row by row, and perform dequantization on the quantization result to obtain the processed model parameters; perform forward calculation and backward calculation according to the processed model parameters to obtain the model parameter gradient; update the model parameters according to the model parameter gradient, and repeat the first processing for the updated model parameters until a predetermined end condition is met.

[0081] Figure 3 For the specific working process of the device embodiment shown, please refer to the relevant description in the foregoing method embodiment, and details are not described herein again.

[0082] In summary, by adopting the solution of the device embodiment of the present application, through quantization processing, single-precision floating-point can be converted to fixed-point, so that in matrix operations during model inference, fixed-point operations can be used instead of floating-point operations, thereby effectively compressing the model volume and improving the model inference speed. Moreover, it can be applied to various different model structures and various different architectures, and has general applicability, etc.

[0083] The solution described in this application can be applied to the field of artificial intelligence, especially in the fields of deep learning and speech recognition. Artificial intelligence is a discipline that studies how to make computers simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0084] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.

[0085] As Figure 4 shown, it is a block diagram of an electronic device according to the method described in an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0086] As Figure 4 shown, the electronic device includes: one or more processors Y01, a memory Y02, and an interface for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a graphical user interface on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as, as a server array, a group of blade servers, or a multi-processor system). Figure 4 One processor Y01 is taken as an example herein.

[0087] Memory Y02 is the non-transitory computer-readable storage medium provided by this application. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the method provided by this application. The non-transitory computer-readable storage medium of this application stores computer instructions, and these computer instructions are used to cause a computer to execute the method provided by this application.

[0088] As a non-transitory computer-readable storage medium, memory Y02 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the embodiments of this application. By running the non-transitory software programs, instructions, and modules stored in memory Y02, processor Y01 executes various functional applications and data processing of the server, that is, implements the method in the above method embodiments.

[0089] Memory Y02 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, memory Y02 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, flash memory device, or other non-transitory solid-state storage devices. In some embodiments, memory Y02 may optionally include a memory remotely set relative to processor Y01, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranet, blockchain network, local area network, mobile communication network, and combinations thereof.

[0090] The electronic device may further include: an input device Y03 and an output device Y04. Processor Y01, memory Y02, input device Y03, and output device Y04 can be connected through a bus or other means, Figure 4 taking the connection through the bus as an example.

[0091] The input device Y03 can receive input digital or character information, and generate key signal inputs related to the user settings and function control of the electronic device, such as input devices like touch screens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, and joysticks. The output device Y04 may include a display device, an auxiliary lighting device, and a tactile feedback device (such as a vibration motor), etc. The display device may include but is not limited to a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0092] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, application specific integrated circuits, computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0093] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a disk, optical disk, memory, programmable logic device) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0094] For purposes of providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0095] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area networks, wide area networks, blockchain networks, and the Internet.

[0096] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0097] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitation is imposed herein.

[0098] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A model compression method, which is applied to a speech recognition scenario and includes: When single-precision floating-point matrix operations are required during model inference, the left matrix among the two matrices to be multiplied is quantized row by row to obtain a first quantization matrix, including: for each row of elements in the left matrix, respectively taking the maximum value among the absolute values of the values of each element included therein as the reference value; for each element in the left matrix, respectively calculating the quotient of the value of the element and the reference value corresponding to the row where the element is located, and calculating the product of the quotient and 2 B-1 of the product, where B represents a predetermined bit width, and taking the product as the quantization value of the element; quantizing the right matrix among the two matrices to be multiplied column by column to obtain a second quantization matrix, including: for each column of elements in the right matrix, respectively taking the maximum value among the absolute values of the values of each element included therein as the reference value; for each element in the right matrix, respectively calculating the quotient of the value of the element and the reference value corresponding to the column where the element is located, and calculating the product of the quotient and the 2 B-1 of the product, and taking the product as the quantization value of the element; Multiplying the first quantization matrix and the second quantization matrix to obtain a third matrix as the result of fixed-point operation; Performing inverse quantization according to the third matrix to obtain a fourth matrix, including: for each element in the fourth matrix, respectively performing the following processing: calculating the product of the value of the corresponding element of the element in the third matrix, the reference value corresponding to the row where the element is located, and the reference value corresponding to the column where the element is located, and using the product as the value of the element, where the corresponding element is the element in the same position; using the fourth matrix as the result of the matrix operation; It further includes: training a single-precision model in a single-precision training manner as an initial model, and performing quantization fine-tuning training on the model parameters of the initial model to obtain a final model; Wherein, the left matrix is a weight matrix; the quantization fine-tuning training of the model parameters of the initial model includes: in each batch of training, for the latest obtained model parameters, respectively performing the following first processing: quantizing the weight matrix in the model parameters row by row, and performing inverse quantization on the quantization result to obtain processed model parameters; performing forward calculation and backward calculation according to the processed model parameters to obtain model parameter gradients; updating the model parameters according to the model parameter gradients, and for the updated model parameters, repeating the first processing until a predetermined end condition is met.

2. A model compression device, which is applied to the speech recognition scenario and includes: A preprocessing module, a quantization module, an operation module, and an inverse quantization module; The quantization module is used to quantize the left matrix in the two multiplied matrices row by row to obtain a first quantization matrix when single-precision floating-point matrix operations are required during model inference, including: for each row element in the left matrix, respectively taking the maximum value among the absolute values of the values of each element included therein as a reference value; for each element in the left matrix, respectively calculating the quotient of the value of the element and the reference value corresponding to the row where the element is located, and calculating the product of the quotient and 2 B-1 where B represents a predetermined bit width, and taking the product as the quantization value of the element; quantizing the right matrix in the two multiplied matrices column by column to obtain a second quantization matrix, including: for each column element in the right matrix, respectively taking the maximum value among the absolute values of the values of each element included therein as a reference value; for each element in the right matrix, respectively calculating the quotient of the value of the element and the reference value corresponding to the column where the element is located, and calculating the product of the quotient and the 2 B-1 and taking the product as the quantization value of the element; The operation module is used to multiply the first quantization matrix and the second quantization matrix to obtain a third matrix as the result of fixed-point operation; The inverse quantization module is used to perform inverse quantization according to the third matrix to obtain a fourth matrix, including: for each element in the fourth matrix, respectively performing the following processing: calculating the product of the value of the corresponding element of the element in the third matrix, the reference value corresponding to the row where the element is located, and the reference value corresponding to the column where the element is located, and using the product as the value of the element, where the corresponding element is the element in the same position; using the fourth matrix as the result of the matrix operation; The preprocessing module is used to train a single-precision model in a single-precision training manner as an initial model, and perform quantization fine-tuning training on the model parameters of the initial model to obtain a final model; Wherein, the left matrix is a weight matrix; in each batch of training, the preprocessing module, for the latest obtained model parameters, respectively performs the following first processing: quantizing the weight matrix in the model parameters row by row, and performing inverse quantization on the quantization result to obtain processed model parameters; performing forward calculation and backward calculation according to the processed model parameters to obtain model parameter gradients; updating the model parameters according to the model parameter gradients, and for the updated model parameters, repeating the first processing until a predetermined end condition is met.

3. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method recited in claim 1.

4. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method recited in claim 1.

Citation Information

Patent Citations

  • Method, apparatus and device for processing floating-point number matrix, and computer readable storage medium

    CN108628807A