Model Compression Method, Training Method, Multimedia Data Processing Method and Device

By generating weight scaling factors in the language generation large model and performing weight parameter scaling processing, the activation value distribution of the model is more concentrated, solving the problem of large-scale deployment resource occupancy and reducing the accuracy loss in the quantization process.

CN117370798BActive Publication Date: 2025-05-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311235188.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-05-30
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

Because the language generates large parameters for large models, it requires a large memory and computing resources to be consumed during deployment, and Cache KV technology causes language generation large models to lose accuracy during quantization.

Method used

By generating weight scaling factors for the activation value distribution information in the activation value matrix of the model, the weight parameter scaling process is performed, so that the calculation result distribution is more concentrated, thereby reducing accuracy loss during quantization.

Benefits of technology

It effectively reduces the accuracy loss of the model compression process, saves data storage, and reduces the demand for hardware resources for model deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117370798B_ABST
    Figure CN117370798B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model compression method, a training method, a multimedia data processing method and apparatus, relating to the field of artificial intelligence technologies, and particularly to technologies such as deep learning, natural language processing, and computer vision. The specific implementation solution of the model compression method is as follows: input multimedia data into a model to be compressed to obtain the activation value matrix of each of the N processing layers cascaded in the model to be compressed, where the activation value matrix of the nth processing layer represents the output features obtained by processing the multimedia data by the n - 1 processing layers located before the nth processing layer; generate a weight scaling factor for each processing layer according to the distribution information of the activation values in the activation value matrix; perform a scaling process on the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain the parameter to be quantized for each processing layer; and sequentially quantize the parameter to be quantized for each processing layer based on a predetermined quantization accuracy to obtain a compressed model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and particularly relates to technical fields such as deep learning, natural language processing, and computer vision. In particular, it relates to a model compression method, a training method, a multimedia data processing method, and an apparatus. Background Art

[0002] Large language generation models based on GPT (Generative Pre-trained Transformer) are widely used in the field of natural language processing. However, due to the large number of parameters in the large language generation models, for example, the number of parameters of the GPT-3 model can reach 170 billion. Therefore, when deploying large language generation models, a large amount of memory and computing resources are required. Summary of the Invention

[0003] The present disclosure provides a model compression method, a training method, and a multimedia data processing method.

[0004] According to one aspect of the present disclosure, a model compression method is provided, including: inputting multimedia data into a model to be compressed to obtain activation value matrices of each of N cascaded processing layers in the model to be compressed, where the activation value matrix of the nth processing layer represents the output features obtained by processing the multimedia data by n - 1 processing layers, N is an integer greater than 1, and n = 2,..., N; generating a weight scaling factor for each processing layer according to the distribution information of the activation values in the activation value matrix; scaling the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain quantization parameters to be quantized for each processing layer; and quantizing the quantization parameters to be quantized for each processing layer in sequence based on a predetermined quantization accuracy to obtain a compressed model.

[0005] According to another aspect of the present disclosure, a training method for a compressed model is provided, including: inputting first multimedia data as sample data into the compressed model to obtain a first probability vector for a plurality of predetermined categories; the first multimedia data is labeled with first category information; determining a loss value of the compressed model according to the first category information and the first probability vector; and adjusting the model parameters of the compressed model according to the loss value to obtain a trained compressed model, where the compressed model is a compressed deep learning model obtained by using the model compression method described above.

[0006] According to another aspect of the present disclosure, there is provided a method for processing multimedia data, including: inputting the multimedia data to be processed into a compression model to obtain a probability vector representing the category to which the multimedia data belongs; the probability vector includes the probability values of the multimedia data belonging to each of a plurality of predetermined categories; and determining, according to the probability vector, the target category to which the multimedia data belongs among the plurality of predetermined categories, wherein the compression model includes a compressed deep learning model obtained by using the model compression method described above.

[0007] According to another aspect of the present disclosure, there is provided a model compression device, including: a first obtaining module, a generating module, a scaling module, and a quantization module. The first obtaining module is configured to input multimedia data into a model to be compressed to obtain an activation value matrix of each of N cascaded processing layers in the model to be compressed, where the activation value matrix of the nth processing layer represents the output features obtained by processing the multimedia data by n - 1 processing layers, N is an integer greater than 1, and n = 2,..., N; the generating module is configured to generate a weight scaling factor for each processing layer according to the distribution information of the activation values in the activation value matrix; the scaling module is configured to perform a scaling process on the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain the parameter to be quantized for each processing layer; and the quantization module is configured to sequentially quantize the parameter to be quantized for each processing layer based on a predetermined quantization precision to obtain a compression model.

[0008] According to another aspect of the present disclosure, there is provided a training device for a compression model, including: a second obtaining module, a first determining module, and an adjusting module. The second obtaining module is configured to input first multimedia data as sample data into the compression model to obtain a first probability vector for a plurality of predetermined categories; the first multimedia data is labeled with first category information; the first determining module is configured to determine the loss value of the compression model according to the first category information and the first probability vector; and the adjusting module is configured to adjust the model parameters of the compression model according to the loss value to obtain a trained compression model, where the compression model is a compressed deep learning model obtained by using the model compression method described above.

[0009] According to another aspect of the present disclosure, there is provided a processing device for multimedia data, including: a third obtaining module and a second determining module. The third obtaining module is configured to input the multimedia data to be processed into the compression model to obtain a probability vector representing the category to which the multimedia data belongs; the probability vector includes the probability values of the multimedia data belonging to each of a plurality of predetermined categories; and the second determining module is configured to determine the target category to which the multimedia data belongs among the plurality of predetermined categories according to the probability vector, where the compression model includes a compressed deep learning model obtained by using the model compression method described above.

[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0012] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method as described above.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0015] Figure 1 Schematically shows an exemplary system architecture to which a model compression method, a training method, or a multimedia data processing method and apparatus according to an embodiment of the present disclosure can be applied;

[0016] Figure 2 Schematically shows a flowchart of a model compression method according to an embodiment of the present disclosure;

[0017] Figure 3 Schematically shows the weight scaling in the processing layer P 3 330 in an embodiment of the present disclosure;

[0018] Figure 4A Schematically shows a schematic diagram of determining quantization parameters according to some embodiments of the present disclosure;

[0019] Figure 4B Schematically shows a schematic diagram of determining quantization parameters according to other embodiments of the present disclosure;

[0020] Figure 4C Schematically shows a schematic diagram of determining quantization parameters according to still other embodiments of the present disclosure;

[0021] Figure 5 Schematically shows an exemplary architecture diagram of a model to be compressed according to an embodiment of the present disclosure;

[0022] Figure 6 Schematically shows a flowchart of a compression model training method according to an embodiment of the present disclosure;

[0023] Figure 7 Schematically shows a flowchart of a multimedia data processing method according to an embodiment of the present disclosure;

[0024] Figure 8 Schematically shows a block diagram of a model compression device according to an embodiment of the present disclosure;

[0025] Figure 9 Schematically shows a block diagram of a compression model training device according to an embodiment of the present disclosure;

[0026] Figure 10 Schematically shows a block diagram of a multimedia data processing device according to an embodiment of the present disclosure; and

[0027] Figure 11 Schematically shows a block diagram of an electronic device suitable for implementing a model compression method, a compression model training method, or a multimedia data processing method according to an embodiment of the present disclosure. Detailed Embodiments

[0028] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0029] In large language generation models, the Cache Key-Value (key-value cache, hereinafter referred to as CacheKV) technology is generally introduced to improve the inference performance of the model. The principle of Cache KV is to cache the computable results reusable in the i-th round. When performing the calculation in the (i + 1)-th round, directly read the computable results of the i-th round cached, and splice the computable results of the i-th round cached and the computable results of the (i + 1)-th round to obtain the final computable results of the (i + 1)-th round.

[0030] Introducing Cache KV into large language generation models not only improves the inference speed of the model but also increases the number of model parameters. The size of the model can be reduced by quantifying the model parameters to lower the requirements for hardware.

[0031] However, due to the large dispersion of the computable results cached in Cache KV, when quantifying large language generation models, relatively serious accuracy loss of large language generation models is caused.

[0032] In view of this, embodiments of the present disclosure provide a model compression method, which generates a weight scaling factor for each processing layer according to the distribution information of activation values in the activation value matrix; and scales the weight parameters of each processing layer according to the weight scaling factor of each processing layer, so that the calculation results of each processing layer are more concentratedly distributed, reducing the loss of model accuracy during the quantization process, saving data storage, and making the data storage space required for the compressed model less than a preset value, thereby reducing the storage space and computing resource requirements of the model deployment for the hardware environment.

[0033] Figure 1 Schematically shows an exemplary system architecture to which the model compression method and apparatus according to embodiments of the present disclosure can be applied.

[0034] It should be noted that Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the model compression method and apparatus can be applied may include a terminal device, but the terminal device can implement the model compression method and apparatus provided by embodiments of the present disclosure without interacting with the server.

[0035] As Figure 1 shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0038] Server 105 may be a server that provides various services, such as a background management server (only for example) that supports the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0039] It should be noted that the model compression method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the model compression device provided by the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0040] Alternatively, the model compression method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the model compression device provided by the embodiments of the present disclosure can generally be set in the server 105. The model compression method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the model compression device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0041] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0042] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.

[0043] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user has been obtained.

[0044] Figure 2 Schematically shows a flowchart of the model compression method according to an embodiment of the present disclosure.

[0045] As Figure 2 shown, the method 200 includes operations S210 to S240.

[0046] In operation S210, multimedia data is input into the model to be compressed, and activation value matrices of each of the N cascaded processing layers in the model to be compressed are obtained.

[0047] In operation S220, weight scaling factors for each processing layer are generated according to the distribution information of the activation values in the activation value matrices.

[0048] In operation S230, according to the weight scaling factors of each processing layer, the weight parameters of each processing layer are scaled to obtain quantization parameters to be processed for each processing layer.

[0049] In operation S240, based on a predetermined quantization precision, the quantization parameters to be processed for each processing layer are sequentially quantized to obtain a compressed model.

[0050] According to an embodiment of the present disclosure, the multimedia data may include at least one of the following: images, texts, audios, videos, etc. The model to be compressed may be a model corresponding to the type of the multimedia data.

[0051] For example: If the multimedia data is text, the model to be compressed may be an ERNIE model or a Generative Pre-trained Transformer (GPT) model, etc.

[0052] For example: If the multimedia data is an image, the model to be compressed may be a Residual Network (ResNet) series model (such as ResNet-50), a Vision Transformer (ViT) model, an End-to-End Object Detection with Transformers (DERT) model, etc.

[0053] For example: If the multimedia data is a video, the model to be compressed may be a Swin-Transformer model, etc.

[0054] According to an embodiment of the present disclosure, the N cascaded processing layers in the model to be compressed may correspond to the type of the model to be compressed. For example: The model to be compressed may be a Generative Pre-trained Transformer model, and the N processing layers may include a normalization layer, a multi-head attention layer, a fully connected layer, a stacking layer, a feed-forward neural network layer, etc.

[0055] According to an embodiment of the present disclosure, the activation value matrix of the nth processing layer represents the output features obtained by processing the multimedia data by the n-1 processing layers located before the nth processing layer, where N is an integer greater than 1, and n = 2,...N. For example: the activation value matrix of the multi-head attention layer is the output features obtained by processing the multimedia data by the normalization layer; the activation value matrix of the fully connected layer is the output features obtained by processing the multimedia data by the normalization layer, and then input into the multi-head attention layer and processed based on the multi-head attention mechanism.

[0056] According to an embodiment of the present disclosure, the distribution information of the activation values can represent the degree of dispersion of the activation values in the activation value matrix. When the elements in the activation value matrix are the same, the larger the value range of the activation values, the higher the degree of dispersion of the activation values. For example: the degree of dispersion of the activation value matrix with the value range of the activation values from 1 to 100 is higher than that of the activation value matrix with the value range of the activation values from 1 to 10. The higher the degree of dispersion, the more parameters are truncated during the quantization process, and the greater the loss of the model accuracy.

[0057] For example: for a certain processing layer, the activation value matrix can be a 3x4 matrix, and the smallest activation value in the 3x4 matrix can be 1, and the largest activation value can be 100. Based on the matrix multiplication calculation of the activation value matrix and the weight parameters in the processing layer, the degree of dispersion of the obtained output result matrix (i.e., the activation values of the next processing layer) is relatively large.

[0058] According to an embodiment of the present disclosure, the weight scaling factor can represent the value used to scale the weight parameters in each processing layer. For example: for a certain processing layer, after scaling the weight parameters according to the weight scaling factor, the distribution of the weight parameters can be relatively concentrated. Therefore, when performing matrix multiplication calculation on the 3x4 matrix and the scaled weight matrix of the processing layer, an output result matrix with relatively concentrated data distribution can be obtained.

[0059] According to an embodiment of the present disclosure, based on the activation value matrix with relatively concentrated data distribution, since there are fewer discrete points in the activation value matrix, the loss of feature information in the quantization parameters obtained by truncating the activation value matrix is less, thereby effectively reducing the loss of the model accuracy during the quantization process.

[0060] According to an embodiment of the present disclosure, according to the distribution information of the activation values in the activation value matrix, the weight scaling factor of each processing layer is generated; according to the weight scaling factor of each processing layer, the weight parameters of each processing layer are scaled, so that the calculation results of each processing layer are relatively concentrated, reducing the loss of the model accuracy during the quantization process, saving the data storage amount, making the data storage space required for compressing the model less than the preset value, thereby reducing the storage space and computing resource requirements of the model deployment for the hardware environment.

[0061] The following is a reference to Figures 3 to 5 , and the method shown in the specific embodiments is further described. Figure 2

[0062] Figure 3 Schematically shows the weight scaling in the processing layer P 3 330 in accordance with an embodiment of the present disclosure.

[0063] As Figure 3 shown, in Embodiment 300, the model to be compressed may include cascaded processing layers P 1 310, processing layer P 2 320, processing layer P 3 330, and processing layer P 4 340.

[0064] According to an embodiment of the present disclosure, multimedia data is input into the processing layer P 1 310, and the output features are input into the processing layer P 2 320 to output features, that is, the activation value matrix 321 of the processing layer P 3 330. In the processing layer P 3 330, the activation value matrix 321 is multiplied by the weight matrix 331 to obtain a first output result 332. In the first output result 332, the color of the color block can represent the magnitude of the first output result value, and the darker the color, the larger the output result value. It can be seen that in the first output result 332, the color distribution of the color blocks varies greatly, indicating that the output result values in the first output result 332 are relatively dispersed.

[0065] According to an embodiment of the present disclosure, the weight scaling threshold of each processing layer can be determined according to the distribution information of the activation values of each processing layer; and the weight scaling factor of each processing layer can be generated according to the activation values of each processing layer and the weight scaling threshold of each processing layer.

[0066] For example: the value range of the activation values can be determined according to all the activation values X11~X34 in the activation value matrix 321, and the maximum activation value can be determined as the weight scaling threshold X max , and the weight scaling factor S max of this processing layer can be generated according to the activation values in the activation value matrix 321 and the weight scaling threshold X ij of this processing layer according to Equation (1).

[0067] S ij =X ij / X max (1)

[0068] Wherein, X ij ​Denotes the element value at the \(i\)-th row and \(j\)-th column in the activation value matrix; \(X\) max Denotes the maximum activation value in the activation value matrix; \(S\) ij Denotes the same as \(X\) ij Corresponding weight scaling factor.

[0069] According to an embodiment of the present disclosure, the weight scaling factor \(S\) corresponding to each activation value in the activation value matrix 321 can be obtained according to Equation (1) ij to obtain the weight scaling factor matrix 322.

[0070] Since the distribution information of the activation values of each processing layer has a certain relationship with the multimedia data being processed, different batches of multimedia data can be input into the model to be compressed based on \(M\) rounds, where \(M\) is an integer greater than 1, so as to obtain different activation value distribution information for each processing layer for different batches of multimedia data.

[0071] According to an embodiment of the present disclosure, determining the weight scaling threshold for each processing layer according to the distribution information of the activation values of each processing layer may include the following operations: for each processing layer, determining a first scaling threshold according to the distribution information of the activation values corresponding to the multimedia data input in the \((m - 1)\)-th round, where \(m\) is an integer greater than 1 and less than or equal to \(M\); determining a second scaling threshold according to the distribution information of the activation values corresponding to the multimedia data input in the \(m\)-th round; and in response to the second scaling threshold being greater than the first scaling threshold, determining the second scaling threshold as the weight scaling threshold.

[0072] For example: for a certain processing layer, for the distribution information of the activation values corresponding to the multimedia data input in the first round, the maximum activation value in this processing layer obtained in the first round can be determined as the first scaling threshold abs-max1. For the distribution information of the activation values corresponding to the multimedia data input in the second round, the maximum activation value in this processing layer obtained in the second round can be determined as the second scaling threshold abs-max2. In the case where abs-max1 is less than abs-max2, abs-max2 is determined as the weight scaling threshold. And so on, until the multimedia data processing process of the \(M\)-th round is completed, the maximum value among the \(M\) scaling thresholds corresponding to the \(M\) rounds can be used as the weight scaling threshold.

[0073] According to an embodiment of the present disclosure, the distribution information of the activation values corresponding to the multimedia data input in the \((m - 1)\)-th round can characterize the numerical range of the activation values. For example: 1 to 100; then the maximum activation value 100 can be determined as the first scaling threshold.

[0074] Similarly, the distribution information of the activation values corresponding to the multimedia data input in the m-th round can characterize the numerical range of the activation values. For example, 1 to 50; then the maximum activation value 50 can be determined as the second scaling threshold. Since the second scaling threshold 50 is less than the first scaling threshold 100, the weight scaling threshold is determined to be 100.

[0075] According to an embodiment of the present disclosure, by dividing each element value in the activation value matrix by the maximum element value in the activation value matrix, a weight scaling factor matrix with element value ranges from 0 to 1 is obtained, so as to narrow the value range of the weight parameter matrix, make the weight parameter distribution concentrated, and facilitate quantization.

[0076] According to an embodiment of the present disclosure, according to the weight scaling factor of each processing layer, scaling processing is performed on the weight parameters of each processing layer to obtain the parameter to be quantized for each processing layer, which may include the following operations:

[0077] According to the weight scaling factor of each processing layer, scaling processing is performed on the weight parameters of each processing layer to obtain the target model after weight scaling; the multimedia input is input into the target model after weight scaling to obtain the activation values of each processing layer; and according to the distribution information of the activation values of each processing layer, the parameter to be quantized for each processing layer is obtained.

[0078] As Figure 3 shown, each element in the weight scaling factor matrix 322 can be multiplied by each element in the initial weight parameter matrix 331 according to the element positions to obtain each element in the updated weight parameter matrix 333.

[0079] For example: multiplying the element S at the first row and first column in the weight scaling factor matrix 322 11 by the element W at the first row and first column in the initial weight parameter matrix 331 11 can obtain the element W at the first row and first column in the updated weight parameter matrix 333 11’ .

[0080] According to an embodiment of the present disclosure, after the weight parameter matrix of the processing layer P 3 330 is changed, the target model after weight scaling can be obtained. At this time, when the activation value matrix 321 of the processing layer P 3 330 remains unchanged, the activation value matrix 321 is multiplied by the updated weight parameter matrix 333 to obtain the second output result 334. In the second output result 334, the color of the color block can characterize the magnitude of the second output result value, and the darker the color, the larger the output result value. It can be seen that in the second output result 334, the color distribution of the color blocks is uniform, indicating that the output result values in the second output result 334 are relatively concentrated.

[0081] According to an embodiment of the present disclosure, by comparing the first output result 332 and the second output result 334, it can be understood that for the same processing layer, when the activation value matrix is the same, after the weight parameter matrix is scaled, the dispersion of the output result value distribution significantly decreases, thereby effectively reducing the loss of accuracy during the model compression process, and achieving the savings of data storage space and the occupancy rate of hardware resources with a relatively low loss of model prediction accuracy.

[0082] According to an embodiment of the present disclosure, obtaining the quantization parameters for each processing layer based on the distribution information of the activation values of each processing layer may include the following operations: obtaining the truncation threshold for each processing layer according to the distribution information of the activation values of each processing layer; and based on the truncation threshold, performing truncation processing on the parameters of each processing layer to obtain the quantization parameters for each processing layer.

[0083] According to an embodiment of the present disclosure, the distribution information of the activation values may characterize the value range of the activation values. For example, when quantifying the activation values in the cache unit (Cache) of a certain processing layer, the maximum activation value among all the currently cached activation values in the cache unit may be determined as the parameter truncation threshold. Based on this maximum activation value, all the currently cached activation values of this processing layer are truncated to obtain the quantization parameters for this processing layer. The quantization parameters may be the maximum activation value of this processing layer.

[0084] According to an embodiment of the present disclosure, taking the maximum activation value among all the activation values as the truncation threshold, then each processing layer has only one quantization factor, which can effectively improve the data processing speed during the quantization process.

[0085] However, each processing layer has only one quantization factor, and the granularity of this quantization method is relatively coarse. In some application scenarios, there may be a problem of relatively low quantization accuracy.

[0086] In view of this, the embodiment of the present disclosure may construct a distribution matrix based on the distribution information of the activation values of each processing layer; divide the distribution matrix to obtain the sub-distribution information of the activation values corresponding to each divided region; and determine the parameter truncation threshold corresponding to each divided region according to the sub-distribution information to improve the quantization accuracy.

[0087] According to an embodiment of the present disclosure, taking the multi-head attention layer as an example, the activation values in the cache unit of this processing layer generally include two dimensions: the number of heads and the number of hidden channels in each head.

[0088] For example: the distribution matrix may be divided according to the number of processing channels of each processing layer to obtain the sub-distribution information of the activation values corresponding to each divided region.

[0089] Figure 4A Schematically shows a schematic diagram of determining parameters to be quantized according to some embodiments of the present disclosure.

[0090] like Figure 4A As shown, in the embodiment 400A, the second output result 334 can be used as the activation value matrix of the next processing layer. The first distribution matrix 410 can be constructed based on the distribution information of the activation value of the second output result 334. The first distribution matrix 410 is divided according to the number of channels. The elements of each column in the distribution matrix 410 represent the activation values ​​of the same processing channel. Therefore, the distribution matrix 410 can be divided by column to obtain the first segmentation area 411, the second segmentation area 412, the third segmentation area 413 and the fourth segmentation area 414.

[0091] According to an embodiment of the present disclosure, determining a parameter cutoff threshold corresponding to each segmented region based on sub-distribution information may include the following operations: determining a maximum activation value corresponding to each segmented region based on the sub-distribution information; and determining the maximum activation value corresponding to each segmented region as the parameter cutoff threshold corresponding to each segmented region.

[0092] For example: for the activation value X corresponding to the first segmented area 411 11’ , activation value X 21’ , activation value X 31’ , activate the value X 11’ , activation value X 21’ , activation value X 31’ The maximum activation value X in 11’ Determine as X maxl The activation value corresponding to the first segmented region 411 is truncated, and the first parameter to be quantized 421 obtained can be the activation value X 11’ By analogy, the second parameter to be quantized 422, the third parameter to be quantized 423 and the fourth parameter to be quantized 424 corresponding to the other three segmented regions can be obtained. The first parameter to be quantized 421, the second parameter to be quantized 422, the third parameter to be quantized 423 and the fourth parameter to be quantized 424 are concatenated to obtain the final parameter to be quantized 430.

[0093] According to an embodiment of the present disclosure, the activation value matrix is ​​segmented according to the number of channels of each processing layer, and a parameter to be quantized can be obtained for each segmented area, so that fine-grained quantization of the activation value can be achieved and the quantization accuracy can be improved.

[0094] The dimension of the parameter to be quantified determined in Embodiment 400A is proportional to the number of channels. Although the quantization accuracy can be improved, the computational time consumption increases continuously with the increase in the number of channels. To balance the computational time consumption and the quantization accuracy, the activation values of each processing layer can be clustered to obtain a clustering result; and based on the clustering result, the distribution matrix can be segmented to obtain the sub-distribution information of the activation values corresponding to each segmented region.

[0095] Figure 4B Schematically shows a schematic diagram for determining the parameter to be quantified according to other embodiments of the present disclosure.

[0096] As Figure 4B shown, in Embodiment 400B, a second distribution matrix 440 can be constructed according to the distribution information of the activation values of the second output result 334. The activation values in the second distribution matrix 440 are clustered according to any clustering algorithm to obtain a clustering result. The clustering result may include the number of clusters obtained after clustering the activation values and the activation values corresponding to each cluster.

[0097] According to an embodiment of the present disclosure, 3 clusters can be obtained according to the clustering result, and the second distribution matrix 440 can be segmented according to the activation values corresponding to each cluster to obtain a fifth segmented region 441, a sixth segmented region 442, and a seventh segmented region 443.

[0098] For the activation values X 11’ ~ activation value X 32’ corresponding to the fifth segmented region 441, the maximum activation value X 11’ ~ activation value X 32’ in is determined as X 32’ max5 . The activation values corresponding to the fifth segmented region 441 are truncated, and the fifth parameter to be quantified 451 obtained may be the activation value X 32’ . By analogy, the sixth parameter to be quantified 452 and the seventh parameter to be quantified 453 corresponding to the other two segmented regions can be obtained. The fifth parameter to be quantified 451, the sixth parameter to be quantified 452, and the seventh parameter to be quantified 453 are concatenated to obtain the final parameter to be quantified 460.

[0099] According to an embodiment of the present disclosure, by clustering the activation values, the activation values with relatively concentrated numerical ranges can be segmented into the same region, and a higher-precision quantization of the model can be achieved with less computational time consumption.

[0100] According to an embodiment of the present disclosure, segmenting the distribution matrix to obtain the sub-distribution information of the activation values corresponding to each segmented region may include the following operations: dividing the distribution matrix evenly according to a predetermined number of regions to obtain the sub-distribution information of the activation values corresponding to each segmented region.​

[0101] Figure 4C A schematic diagram of determining parameters to be quantized according to some further embodiments of the present disclosure is schematically shown.

[0102] like Figure 4C As shown, in embodiment 400C, a third distribution matrix 470 can be constructed according to the distribution information of the activation value of the second output result 334. The third distribution matrix 470 can be evenly divided according to the predetermined number of area divisions to obtain an eighth segmentation area 471 and a ninth segmentation area 472, and the number of activation values ​​in each segmentation area is the same.

[0103] The activation value X corresponding to the eighth segmented area 471 11’ Activation value X 32’ , activate the value X 11’ Activation value X 32’ The maximum activation value X in 32’ Determine as X max8 The activation value corresponding to the eighth segmented region 471 is truncated, and the obtained eighth parameter to be quantized 481 may be the activation value X 32’ By analogy, the ninth parameters to be quantized 482 corresponding to the other segmented regions can be obtained. The eighth parameter to be quantized 481 and the ninth parameter to be quantized 482 are concatenated to obtain the final parameter to be quantized 490.

[0104] According to the embodiments of the present disclosure, the number of predetermined areas can be pre-set based on the needs of actual application scenarios, so as to realize the quantification of the model with higher precision based on the needs of actual application scenarios, for example, the needs for computational time consumption.

[0105] According to an embodiment of the present disclosure, based on a predetermined quantization accuracy, the parameters to be quantized of each processing layer are quantized in turn to obtain a compression model, which may include the following operations: for each processing layer, a quantization factor is determined according to the predetermined quantization accuracy and the distribution information of the parameters to be quantized of each processing layer; and according to the quantization factor, the parameters to be quantized of each processing layer are quantized to obtain a compression model.

[0106] According to an embodiment of the present disclosure, for each processing layer, a quantization factor is determined according to a predetermined quantization accuracy and distribution information of the parameters to be quantized of each processing layer, which may include the following operations: determining the maximum parameter value of each processing layer according to the distribution information of the parameters to be quantized of each processing layer; and obtaining the quantization factor according to the predetermined quantization accuracy and the maximum parameter value.

[0107] For example, for a certain processing layer, the parameters to be quantized may include the activation value X 32’ and activation value X 14’ , at the activation value X32' Greater than the activation value X 14' In this case, the maximum parameter value can be determined as the activation value X 32' .

[0108] According to an embodiment of the present disclosure, the parameter to be quantized can be quantized according to formulas (2) and (3) to obtain a compressed model.

[0109]

[0110] x q = clip(round(s·x), -2 b-1 , 2 b-1 ) (3)

[0111] Where s represents the quantization factor, b represents the predetermined quantization precision, α represents the parameter with the largest absolute value among the parameters to be quantized; x q represents the quantized parameter, x represents the parameter to be quantized; round() represents the rounding operation, and clip represents truncation with the maximum and minimum values.

[0112] According to an embodiment of the present disclosure, since the weight parameters of each processing layer are scaled, the distribution of the activation values of each processing layer is relatively concentrated. When quantizing the activation values, the precision loss in the quantization process can be effectively reduced, achieving the compression of the model with a small precision loss, saving the data storage amount of the model parameters and the hardware resource requirements of the deployment environment.

[0113] When applying the model compression method provided by the embodiment of the present disclosure to compress a large language generation model, in order not to change the prediction accuracy of the large language generation model, therefore, the weight scaling process between two processing layers with data interaction should be mathematically equivalent. For example: in the transformer network, the query weight parameter matrix (Q) and the key weight parameter matrix (K) need to perform matrix multiplication operations. Therefore, after reducing the weight parameters in the query weight parameter matrix based on the weight scaling factor, the key weight parameter matrix needs to be enlarged according to the reciprocal of the same weight scaling factor so that the result of the matrix multiplication is mathematically equivalent to the result of the matrix multiplication before the weight scaling.

[0114] Figure 5 Schematically shows an exemplary architecture diagram of the model to be compressed according to an embodiment of the present disclosure.

[0115] Such as Figure 5As shown, the model 500 to be compressed may include a Layer Normalization layer (LayerNorm) 510, a value feature processing layer (V) 520_1, a query feature processing layer (Q) 520_2, a key feature processing layer (K) 520_3, a first cache unit (Cache) 530_1, a second cache unit (Cache) 530_2, a first matrix multiplication processing layer (BatchedMatMul, BMM) 540_1, a second matrix multiplication processing layer (BatchedMatMul, BMM) 540_2, a Softmax layer 550, and an output layer (Out proj) 560.

[0116] For the query feature processing layer (Q) 520_2 and the key feature processing layer (K) 520_3, the updated weight parameters in the key feature processing layer (K) 520_3 are obtained by element-wise multiplication of the initial weight parameter matrix W k and the first weight scaling factor matrix S k The updated weight parameters in the query feature processing layer (Q) 520_2 are obtained by element-wise multiplication of the initial weight parameter matrix W q and the reciprocal of the first weight scaling factor matrix 1 / S k The first weight scaling factor matrix is obtained according to the activation values cached in the first cache unit 530_1 using the weight factor calculation method described above.

[0117] When the second matrix multiplication processing layer 540_2 performs matrix multiplication on the query features and the key features, the operation result remains unchanged, thus achieving that when scaling the weight parameter matrix by inserting scaling factors in the query feature processing layer (Q) 520_2 and the key feature processing layer (K) 520_3, an operation result mathematically equivalent to that before scaling the weight parameter matrix can be obtained, thereby ensuring the accuracy of the model prediction result.

[0118] Similarly, for the value feature processing layer (V) 520_1, the updated weight parameters are obtained by element-wise multiplication of the initial weight parameter matrix W v and the second weight scaling factor matrix S v The updated weight parameters in the output layer (Out proj) 560 are obtained by element-wise multiplication of the initial weight parameter matrix W o and the reciprocal of the second weight scaling factor matrix S v The second weight scaling factor matrix S v is obtained according to the activation values cached in the second cache unit 530_2 using the weight factor calculation method described above.

[0119] When the matrix multiplication operation is performed on the value features and the output result of the Softmax layer 550 in the first matrix multiplication processing layer 540_2, and then the operation result is input to the output layer 560 for operation, the operation result of the output layer 560 does not change. Thus, when the scaling factor is inserted into the value feature processing layer (V) 520_1 and the output layer (Out proj) 560 to scale the weight parameter matrix, an operation result mathematically equivalent to that before the scaling of the weight parameter matrix can be obtained, thereby ensuring the accuracy of the model prediction result.

[0120] According to an embodiment of the present disclosure, since the weight scaling factor matrix is incorporated into the weight parameters of each processing layer, the activation value distributions cached in the first cache unit 530_1 or the second cache unit 530_2 in each round are relatively concentrated, thereby reducing the loss of the model during the quantization process.

[0121] Figure 6 The flowchart of the compression model training method according to an embodiment of the present disclosure is schematically shown.

[0122] As Figure 6 shown, the training method 600 may include operations S610 to S630.

[0123] In operation S610, the first multimedia data serving as sample data is input into the compression model to obtain a first probability vector for multiple predetermined categories.

[0124] In operation S620, the loss value of the compression model is determined according to the first category information and the first probability vector.

[0125] In operation S630, based on the loss value, the model parameters of the compression model are adjusted to obtain a trained compression model.

[0126] According to an embodiment of the present disclosure, the compression model may be obtained by compressing the model to be compressed according to the method described above. The sample data may be part of the sample data in the training set for pre-training the compression model. For example, 5%, 10%, or any other arbitrary proportion of data may be randomly selected from the training set to obtain the sample data. The first multimedia data serving as the sample data may include category information, for example, it may include the first category information. The first category information represents a certain category among multiple predetermined categories.

[0127] According to an embodiment of the present disclosure, the sample data may be the same as the defined scope of the multimedia data described above. After processing the first multimedia data, the compression model may output a first probability vector. The first probability vector includes the probability values predicted by the compression model for each of the multiple predetermined categories. For example, the type of the first multimedia data may be an image, and the predetermined categories may be the attributes of the target objects in the image, such as people, animals, buildings, etc. The type of the first multimedia data may be text, and the predetermined categories may be the semantic attribute words of the text.

[0128] According to an embodiment of the present disclosure, the loss value of the compression model may be calculated by using a predetermined loss function according to the probability value corresponding to the category represented by the first category information in the first probability vector. Among them, the predetermined loss function may be, for example, a cross-entropy loss function, a mean square error loss function (i.e., L2 loss function), or a hinge loss function, etc., and the present disclosure does not limit this.

[0129] According to an embodiment of the present disclosure, with the goal of minimizing the loss value, a gradient descent algorithm may be used to adjust network parameters such as weight parameters in the compression model to achieve the training of the compression model.

[0130] Figure 7 The flowchart of the multimedia data processing method according to an embodiment of the present disclosure is schematically shown.

[0131] As Figure 7 shown, the multimedia data processing method 700 may include operations S710 to S720.

[0132] In operation S710, the multimedia data to be processed is input into the compression model to obtain a probability vector representing the category to which the multimedia data belongs.

[0133] In operation S720, according to the probability vector, the target category to which the multimedia data belongs among the multiple predetermined categories is determined.

[0134] According to an embodiment of the present disclosure, the compression model may be a deep learning model obtained by compressing the model to be compressed according to the model compression method described above. It may also be obtained by training the compressed deep learning model according to the training method of the compression model described above.

[0135] According to an embodiment of the present disclosure, the implementation principle of operation S710 is similar to the implementation principle of operation S610 described above. The probability vector may include the probability values of the multimedia data belonging to each of the multiple predetermined categories, which will not be elaborated here.

[0136] According to an embodiment of the present disclosure, the predetermined category corresponding to the maximum probability value in the probability vector can be used as the target category to which the multimedia data belongs. When the multimedia data is text, the compression model can be, for example, a model obtained by compressing and training a Wenxin model or the like. When the multimedia data is an image, the compression model can be, for example, a model obtained by compressing and training ResNet-50 or the like.

[0137] Based on the model compression method provided by the embodiments of the present disclosure, the embodiments of the present disclosure further provide a model compression apparatus, which will be described in detail below in conjunction with Figure 8 the model compression apparatus.

[0138] Figure 8 The block diagram of the model compression apparatus according to the embodiment of the present disclosure is schematically shown.

[0139] As Figure 8 shown, the model compression apparatus 800 may include: a first acquisition module 810, a generation module 820, a scaling module 830, and a quantization module 840.

[0140] The first acquisition module 810 is configured to input multimedia data into the model to be compressed, and obtain the activation value matrix of each of the N cascaded processing layers in the model to be compressed, where the activation value matrix of the nth processing layer represents the output feature obtained by processing the multimedia data by n - 1 processing layers, N is an integer greater than 1, and n = 2,..., N.

[0141] The generation module 820 is configured to generate a weight scaling factor for each processing layer according to the distribution information of the activation values in the activation value matrix.

[0142] The scaling module 830 is configured to perform a scaling process on the weight parameters of each processing layer according to the weight scaling factor of each processing layer, and obtain the parameter to be quantized for each processing layer.

[0143] The quantization module 840 is configured to quantize the parameter to be quantized for each processing layer in sequence based on a predetermined quantization accuracy, and obtain a compression model.

[0144] According to an embodiment of the present disclosure, the scaling module may include: a scaling sub-module, a first acquisition sub-module, and a second acquisition sub-module. The scaling sub-module is configured to perform a scaling process on the weight parameters of each processing layer according to the weight scaling factor of each processing layer, and obtain a target model after weight scaling. The first acquisition sub-module is configured to input multimedia into the target model after weight scaling, and obtain the activation value of each processing layer. The second acquisition sub-module is configured to obtain the parameter to be quantized for each processing layer according to the distribution information of the activation values of each processing layer.

[0145] According to an embodiment of the present disclosure, the second obtaining sub-module may include: a first obtaining unit and a first processing unit. The first obtaining unit is configured to obtain a parameter truncation threshold for each processing layer according to the distribution information of the activation values of each processing layer. The first processing unit is configured to perform truncation processing on the parameters of each processing layer based on the parameter truncation threshold to obtain the quantifiable parameters of each processing layer.

[0146] According to an embodiment of the present disclosure, the first obtaining unit may include: a construction subunit, a segmentation subunit, and a first determination subunit. The construction subunit is configured to construct a distribution matrix according to the distribution information of the activation values of each processing layer. The segmentation subunit is configured to segment the distribution matrix to obtain sub-distribution information of the activation values corresponding to each segmentation region. The first determination subunit is configured to determine a parameter truncation threshold corresponding to each segmentation region according to the sub-distribution information.

[0147] According to an embodiment of the present disclosure, the segmentation subunit is configured to: segment the distribution matrix according to the number of processing channels of each processing layer to obtain sub-distribution information of the activation values corresponding to each segmentation region.

[0148] According to an embodiment of the present disclosure, the segmentation subunit is configured to: cluster the activation values of each processing layer to obtain a clustering result; and segment the distribution matrix based on the clustering result to obtain sub-distribution information of the activation values corresponding to each segmentation region.

[0149] According to an embodiment of the present disclosure, the segmentation subunit is configured to: evenly segment the distribution matrix according to a predetermined number of regions to obtain sub-distribution information of the activation values corresponding to each segmentation region.

[0150] According to an embodiment of the present disclosure, the first determination subunit is configured to: determine the maximum activation value corresponding to each segmentation region according to the sub-distribution information; and determine the maximum activation value corresponding to each segmentation region as the parameter truncation threshold corresponding to each segmentation region.

[0151] According to an embodiment of the present disclosure, the generation module may include: a first determination sub-module and a generation sub-module. The first determination sub-module is configured to determine a weight scaling threshold for each processing layer according to the distribution information of the activation values of each processing layer. The generation sub-module is configured to generate a weight scaling factor for each processing layer according to the activation value of each processing layer and the weight scaling threshold of each processing layer.

[0152] According to an embodiment of the present disclosure, the first determination sub-module may include: a first determination unit, a second determination unit, and a third determination unit. The first determination unit is configured to determine a first scaling threshold for each processing layer according to the distribution information of the activation values corresponding to the multimedia data input in the (m-1)-th round, where m is an integer greater than 1 and less than or equal to M. The second determination unit is configured to determine a second scaling threshold according to the distribution information of the activation values corresponding to the multimedia data input in the m-th round. The third determination unit is configured to determine the second scaling threshold as the weight scaling threshold in response to the second scaling threshold being greater than the first scaling threshold.

[0153] According to an embodiment of the present disclosure, the first determination unit may include: a second determination sub-unit and a third determination sub-unit. The second determination sub-unit is configured to determine the maximum activation value of each processing layer corresponding to the (m-1)-th round according to the distribution information of the activation values corresponding to the multimedia data input in the (m-1)-th round. The third determination sub-unit is configured to determine the maximum activation value of each processing layer corresponding to the (m-1)-th round as the first scaling threshold.

[0154] According to an embodiment of the present disclosure, the second determination unit may include: a fourth determination sub-unit and a fifth determination sub-unit. The fourth determination sub-unit is configured to determine the maximum activation value of each processing layer corresponding to the m-th round according to the distribution information of the activation values corresponding to the multimedia data input in the m-th round. The fifth determination sub-unit is configured to determine the maximum activation value of each processing layer corresponding to the m-th round as the second scaling threshold.

[0155] According to an embodiment of the present disclosure, the quantization module may include: a second determination sub-module and a quantization sub-module. The second determination sub-module is configured to determine a quantization factor for each processing layer according to a predetermined quantization precision and the distribution information of the parameters to be quantized of each processing layer. The quantization sub-module is configured to quantize the parameters to be quantized of each processing layer according to the quantization factor to obtain a compressed model.

[0156] According to an embodiment of the present disclosure, the second determination sub-module may include: a fourth determination unit and a second obtaining unit. The fourth determination unit is configured to determine the maximum parameter value of each processing layer according to the distribution information of the parameters to be quantized of each processing layer. The second obtaining unit is configured to obtain the quantization factor according to the predetermined quantization precision and the maximum parameter value.

[0157] Based on the training method of the compressed model provided by the embodiment of the present disclosure, the embodiment of the present disclosure also provides a training device for the compressed model, which will be described in detail below in combination with Figure 9 The training device for the compressed model will be described in detail.

[0158] Figure 9 The block diagram of the training device for the compressed model according to the embodiment of the present disclosure is schematically shown.

[0159] As shown in Figure 9 , the training device 900 may include: a second acquisition module 910, a first determination module 920, and an adjustment module 930.

[0160] The second acquisition module 910 is configured to input first multimedia data as sample data into a compression model to obtain a first probability vector for multiple predetermined categories; the first multimedia data is labeled with first category information.

[0161] The first determination module 920 is configured to determine a loss value of the compression model according to the first category information and the first probability vector.

[0162] The adjustment module 930 is configured to adjust model parameters of the compression model according to the loss value to obtain a trained compression model. The compression model is a compressed deep learning model obtained by using the model compression method described above.

[0163] Based on the multimedia data processing method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide a multimedia data processing device, which will be described in detail below in combination with Figure 10 the multimedia data processing device.

[0164] Figure 10 A block diagram of a multimedia data processing device according to an embodiment of the present disclosure is schematically shown.

[0165] As shown in Figure 10 , the processing device 1000 may include a third acquisition module 1010 and a second determination module 1020.

[0166] The third acquisition module 1010 is configured to input multimedia data to be processed into a compression model to obtain a probability vector representing the category to which the multimedia data belongs; the probability vector includes probability values of the multimedia data belonging to each of multiple predetermined categories.

[0167] The second determination module 1020 is configured to determine a target category to which the multimedia data belongs among multiple predetermined categories according to the probability vector. The compression model is a compressed deep learning model obtained by using the model compression method described above. It may also be a model obtained by training the compressed model using the compressed model training method described above.

[0168] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0169] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0170] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method as described above.

[0171] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method described above when executed by a processor.

[0172] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0173] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0174] A plurality of components in the device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0175] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the model compression method, the training method of the compressed model, or the processing method of multimedia data. For example, in some embodiments, the model compression method, the training method of the compressed model, or the processing method of multimedia data can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the model compression method, the training method of the compressed model, or the processing method of multimedia data described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the model compression method, the training method of the compressed model, or the processing method of multimedia data by any other suitable means (e.g., by means of firmware).

[0176] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0177] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0178] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0179] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0180] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0181] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0182] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0183] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A model compression method, comprising: Inputting multimedia data into a model to be compressed, obtaining activation value matrices of each of N processing layers cascaded in the model to be compressed, where the activation value matrix of the nth processing layer represents the output features obtained by processing the multimedia data by the n - 1 processing layers before the nth processing layer, N is an integer greater than 1, and n = 2,..., N; Generating a weight scaling factor for each processing layer to scale the weight parameters in each processing layer according to the distribution information of the activation values in the activation value matrix; wherein, the weight scaling processes between processing layers with data interaction are mathematically equivalent; Scaling the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain quantization parameters to be processed for each processing layer; and Quantizing the quantization parameters to be processed for each processing layer in sequence based on a predetermined quantization precision to obtain a compressed model.

2. The method according to claim 1, wherein, The scaling the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain quantization parameters to be processed for each processing layer includes: Scaling the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain a target model after weight scaling; Inputting the multimedia data into the target model after weight scaling to obtain the activation values of each processing layer; and Obtaining the quantization parameters to be processed for each processing layer according to the distribution information of the activation values of each processing layer.

3. The method according to claim 2, wherein, The obtaining the quantization parameters to be processed for each processing layer according to the distribution information of the activation values of each processing layer includes: Obtaining a truncation threshold for each processing layer according to the distribution information of the activation values of each processing layer; and Based on the truncation threshold, performing truncation processing on the parameters of each processing layer to obtain the quantization parameters to be processed for each processing layer.

4. The method according to claim 3, wherein, The obtaining a truncation threshold for each processing layer according to the distribution information of the activation values of each processing layer includes: Constructing a distribution matrix according to the distribution information of the activation values of each processing layer; Dividing the distribution matrix to obtain sub - distribution information of the activation values corresponding to each divided region; and Determining a truncation threshold corresponding to each divided region according to the sub - distribution information.

5. The method according to claim 4, wherein, The dividing the distribution matrix to obtain sub - distribution information of the activation values corresponding to each divided region includes: Dividing the distribution matrix according to the processing channel identifier of the model to be compressed to obtain the sub - distribution information of the activation values corresponding to each divided region.

6. The method according to claim 4, wherein, The dividing the distribution matrix to obtain sub - distribution information of the activation values corresponding to each divided region includes: Clustering the activation values of each processing layer to obtain a clustering result; and Based on the clustering result, the distribution matrix is segmented to obtain sub-distribution information of activation values corresponding to each segmented region.

7. The method according to claim 4, wherein, the segmenting the distribution matrix to obtain sub-distribution information of activation values corresponding to each segmented region includes: averagely segmenting the distribution matrix according to a predetermined number of regions to obtain sub-distribution information of activation values corresponding to each segmented region.

8. The method according to claim 5, wherein, the determining a truncation threshold corresponding to each segmented region according to the sub-distribution information includes: determining a maximum activation value corresponding to each segmented region according to the sub-distribution information; and determining the maximum activation value corresponding to each segmented region as the truncation threshold corresponding to each segmented region.

9. The method according to claim 1, wherein, the generating a weight scaling factor for each processing layer according to the distribution information of activation values of each processing layer includes: determining a weight scaling threshold for each processing layer according to the distribution information of activation values of each processing layer; and generating a weight scaling factor for each processing layer according to the activation value of each processing layer and the weight scaling threshold of each processing layer.

10. The method according to claim 9, wherein, the distribution information of activation values of each processing layer includes the distribution information of activation values corresponding to the multimedia data input in M rounds, and M is an integer greater than 1; the determining a weight scaling threshold for each processing layer according to the distribution information of activation values of each processing layer includes: for each processing layer, determining a first scaling threshold according to the distribution information of activation values corresponding to the multimedia data input in the (m - 1)-th round, where m is an integer greater than 1 and less than or equal to M; determining a second scaling threshold according to the distribution information of activation values corresponding to the multimedia data input in the m-th round; and in response to the second scaling threshold being greater than the first scaling threshold, determining the second scaling threshold as the weight scaling threshold.

11. The method according to claim 10, wherein, the determining a first scaling threshold according to the distribution information of activation values corresponding to the multimedia data input in the (m - 1)-th round includes: determining a maximum activation value of each processing layer corresponding to the (m - 1)-th round according to the distribution information of activation values corresponding to the multimedia data input in the (m - 1)-th round; and determining the maximum activation value of each processing layer corresponding to the (m - 1)-th round as the first scaling threshold.

12. The method according to claim 10, wherein, the determining a second scaling threshold according to the distribution information of activation values corresponding to the multimedia data input in the m-th round includes: determining a maximum activation value of each processing layer corresponding to the m-th round according to the distribution information of activation values corresponding to the multimedia data input in the m-th round; and Determine the maximum activation value of each processing layer corresponding to the m-th round as the second scaling threshold.

13. The method according to claim 1, wherein, Based on a predetermined quantization precision, quantize the parameters to be quantized of each processing layer in sequence to obtain a compressed model, including: For each processing layer, determine a quantization factor according to the predetermined quantization precision and the distribution information of the parameters to be quantized of each processing layer; Quantize the parameters to be quantized of each processing layer according to the quantization factor to obtain the compressed model.

14. The method according to claim 13, wherein, The step of, for each processing layer, determining a quantization factor according to the predetermined quantization precision and the distribution information of the parameters to be quantized of each processing layer includes: Determine the maximum parameter value of each processing layer according to the distribution information of the parameters to be quantized of each processing layer; and Obtain the quantization factor according to the predetermined quantization precision and the maximum parameter value.

15. A method for training a compressed model, including: Input first multimedia data as sample data into a compressed model to obtain a first probability vector for multiple predetermined categories; The first multimedia data is labeled with first category information; Determine a loss value of the compressed model according to the first category information and the first probability vector; and Adjust model parameters of the compressed model according to the loss value to obtain a trained compressed model, wherein the compressed model is a compressed deep learning model obtained by using the method according to any one of claims 1 to 14.

16. A method for processing multimedia data, including: Input multimedia data to be processed into a compressed model to obtain a probability vector representing the category to which the multimedia data belongs; The probability vector includes probability values for each of the multiple predetermined categories that the multimedia data belongs to; and Determine a target category to which the multimedia data belongs among the multiple predetermined categories according to the probability vector, wherein the compressed model includes a compressed deep learning model obtained by using the method according to any one of claims 1 to 14.

17. A model compression device, including: A first acquisition module, configured to input multimedia data into a model to be compressed, and obtain an activation value matrix of each of N processing layers cascaded in the model to be compressed, wherein the activation value matrix of the n-th processing layer represents output features obtained by processing the multimedia data by n-1 processing layers before the n-th processing layer, N is an integer greater than 1, and n = 2,..., N; A generation module, configured to generate a weight scaling factor for each processing layer to scale the weight parameters in each processing layer according to the distribution information of the activation values in the activation value matrix; wherein, the weight scaling process between processing layers with data interaction is mathematically equivalent; A scaling module, configured to perform a scaling process on the weight parameters of each processing layer according to the weight scaling factor of each processing layer to obtain the parameters to be quantized of each processing layer; and A quantization module, configured to quantize the parameters to be quantized of each of the processing layers in sequence based on a predetermined quantization precision, so as to obtain a compressed model.

18. The apparatus according to claim 17, wherein, the scaling module includes: a scaling sub-module, configured to scale the weight parameters of each of the processing layers according to the weight scaling factor of each of the processing layers, so as to obtain a target model after weight scaling; a first obtaining sub-module, configured to input the multimedia into the target model after weight scaling, so as to obtain the activation values of each of the processing layers; and a second obtaining sub-module, configured to obtain the parameters to be quantized of each of the processing layers according to the distribution information of the activation values of each of the processing layers.

19. The apparatus according to claim 18, wherein, the second obtaining sub-module includes: a first obtaining unit, configured to obtain the truncation threshold of each of the processing layers according to the distribution information of the activation values of each of the processing layers; and a first processing unit, configured to perform truncation processing on the parameters of each of the processing layers based on the truncation threshold, so as to obtain the parameters to be quantized of each of the processing layers.

20. The apparatus according to claim 19, wherein, the first obtaining unit includes: a construction subunit, configured to construct a distribution matrix according to the distribution information of the activation values of each of the processing layers; a segmentation subunit, configured to segment the distribution matrix to obtain sub-distribution information of the activation values corresponding to each segmentation region; and a first determination subunit, configured to determine the truncation threshold corresponding to each segmentation region according to the sub-distribution information.

21. The apparatus according to claim 20, wherein, the segmentation subunit is configured to: segment the distribution matrix according to the number of processing channels of the model to be compressed, so as to obtain the sub-distribution information of the activation values corresponding to each segmentation region.

22. The apparatus according to claim 20, wherein, the segmentation subunit is configured to: cluster the activation values of each of the processing layers to obtain a clustering result; and and segment the distribution matrix based on the clustering result to obtain sub-distribution information of the activation values corresponding to each segmentation region.

23. The apparatus according to claim 20, wherein, the segmentation subunit is configured to: perform average segmentation on the distribution matrix according to a predetermined number of regions, so as to obtain sub-distribution information of the activation values corresponding to each segmentation region.

24. The apparatus according to claim 20, wherein, the first determination subunit is configured to: determine the maximum activation value corresponding to each segmentation region according to the sub-distribution information; and and determine the maximum activation value corresponding to each segmentation region as the truncation threshold corresponding to each segmentation region.

25. The apparatus according to claim 17, wherein, the generation module includes: a first determination sub-module, configured to determine the weight scaling threshold of each of the processing layers according to the distribution information of the activation values of each of the processing layers; and a generation sub-module, configured to generate the weight scaling factor of each of the processing layers according to the activation values of each of the processing layers and the weight scaling threshold of each of the processing layers.

26. The apparatus according to claim 25, wherein, the distribution information of the activation values of each processing layer includes the distribution information of the activation values corresponding to the multimedia data input in M rounds, and M is an integer greater than 1; the first determination sub-module includes: a first determination unit, configured to, for each processing layer, determine a first scaling threshold according to the distribution information of the activation values corresponding to the multimedia data input in the (m-1)-th round, where m is an integer greater than 1 and less than or equal to M; a second determination unit, configured to determine a second scaling threshold according to the distribution information of the activation values corresponding to the multimedia data input in the m-th round; and a third determination unit, configured to, in response to the second scaling threshold being greater than the first scaling threshold, determine the second scaling threshold as the weight scaling threshold.

27. The apparatus according to claim 26, wherein, the first determination unit includes: a second determination sub-unit, configured to determine the maximum activation value of each processing layer corresponding to the (m-1)-th round according to the distribution information of the activation values corresponding to the multimedia data input in the (m-1)-th round; and a third determination sub-unit, configured to determine the maximum activation value of each processing layer corresponding to the (m-1)-th round as the first scaling threshold.

28. The apparatus according to claim 26, wherein, the second determination unit includes: a fourth determination sub-unit, configured to determine the maximum activation value of each processing layer corresponding to the m-th round according to the distribution information of the activation values corresponding to the multimedia data input in the m-th round; and a fifth determination sub-unit, configured to determine the maximum activation value of each processing layer corresponding to the m-th round as the second scaling threshold.

29. The apparatus according to claim 17, wherein, the quantization module includes: a second determination sub-module, configured to, for each processing layer, determine a quantization factor according to the predetermined quantization precision and the distribution information of the parameters to be quantized of each processing layer; a quantization sub-module, configured to quantize the parameters to be quantized of each processing layer according to the quantization factor to obtain the compressed model.

30. The apparatus according to claim 29, wherein, the second determination sub-module includes: a fourth determination unit, configured to determine the maximum parameter value of each processing layer according to the distribution information of the parameters to be quantized of each processing layer; and a second obtaining unit, configured to obtain the quantization factor according to the predetermined quantization precision and the maximum parameter value.

31. A training apparatus for a compressed model, comprising: a second obtaining module, configured to input first multimedia data as sample data into the compressed model to obtain a first probability vector for a plurality of predetermined categories; the first multimedia data is labeled with first category information; a first determination module, configured to determine a loss value of the compressed model according to the first category information and the first probability vector; and An adjustment module, configured to adjust model parameters of the compression model according to the loss value to obtain a trained compression model, where the compression model is a compressed deep learning model obtained by using the method according to any one of claims 1 to 14.

32. A processing device for multimedia data, comprising: A third obtaining module, configured to input the multimedia data to be processed into the compression model to obtain a probability vector representing the category to which the multimedia data belongs; The probability vector includes probability values of the multimedia data belonging to each of a plurality of predetermined categories; and A second determining module, configured to determine a target category to which the multimedia data belongs among the plurality of predetermined categories according to the probability vector, where the compression model includes a compressed deep learning model obtained by using the method according to any one of claims 1 to 14.

33. An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 16.

34. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 16.

35. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Compression method, training method, processing method and device of deep learning model

    CN116611495A