Hardware-aware dynamic model compression method and system

Through the hardware-aware dynamic model compression method, the hardware feature vector and reinforcement learning agent generation adaptation strategy are used to adjust the model structure and accuracy allocation, and the model compression problem caused by hardware differences in the existing technology is solved, and the model compression efficiency and hardware resource utilization are improved.

CN120579593APending Publication Date: 2025-09-02SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510710314.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing model compression method cannot adapt to the hardware differences between different devices, resulting in excessive cropping on low-end devices, resulting in a sharp drop in accuracy, unable to fully utilize computing power on high-end devices, and dynamic pruning based on input data increases inference delay, and low hardware resource utilization.

Method used

By classifying and encoding the hardware parameters of the target device to generate hardware feature vectors, using reinforcement learning agents to output mixed granularity cutting strategies and mixed precision allocation modes, adjust the model structure and quantize it to generate a model compression scheme suitable for hardware.

Benefits of technology

The coordinated optimization of model compression and hardware resources is realized, and a model compression solution suitable for deploying hardware is generated, which improves compression efficiency and hardware resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579593A_ABST
    Figure CN120579593A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a hardware-aware dynamic model compression method and system, and the method comprises the steps: carrying out the classification coding of hardware parameters of target equipment, and generating a unified hardware feature vector; inputting the hardware feature vector and the performance index of the to-be-compressed model into a trained reinforcement learning agent, and outputting a mixed granularity cutting strategy and a mixed precision distribution mode; adjusting a model structure of the to-be-compressed model according to the mixed granularity cutting strategy, deleting redundant nodes and reconnecting a computational graph; and performing quantification processing on the model parameters of the to-be-compressed model according to the mixing precision distribution mode. Through a closed-loop framework of hardware parameter coding, cutting strategy generation and mixing precision optimization, cooperation of model compression and hardware resources is realized, and a model compression scheme suitable for hardware deployment is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, for example, to a hardware-aware dynamic model compression method and system. Background Art

[0002] Current mainstream large language models have a large number of parameters. When deployed locally on devices with limited computing power and memory resources, model compression is often used to reduce storage and computing requirements. Model compression is the process of reducing the size of machine learning models through technical means, reducing computational complexity and memory usage.

[0003] Common model compression methods include model pruning and quantization. Currently, fixed model pruning strategies cannot adapt to hardware differences between different devices, resulting in a sharp drop in accuracy due to excessive pruning on low-end devices, or underutilization of computing power on high-end devices. Dynamic pruning based on input data requires additional computing resources to generate pruning masks during inference, increasing inference latency. Furthermore, quantization schemes designed for specific hardware may not perform well on other devices, such as those with mismatched hardware capabilities, resulting in reduced hardware resource utilization. Therefore, a hardware-aware dynamic model compression method is urgently needed to optimize the model compression process.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] The embodiments of the present disclosure provide a hardware-aware dynamic model compression method and system to solve the technical problems of existing model compression methods having inference delays or incompatibility with hardware.

[0007] In some embodiments, the hardware-aware dynamic model compression method includes:

[0008] Classify and encode the hardware parameters of the target device to generate a unified hardware feature vector;

[0009] Input the hardware feature vector and the performance index of the model to be compressed into the trained reinforcement learning agent, and output the mixed granularity pruning strategy and mixed precision allocation mode;

[0010] Adjust the model structure of the model to be compressed according to the hybrid granularity pruning strategy, delete redundant nodes and reconnect the calculation graph;

[0011] The model parameters of the model to be compressed are quantized according to the mixed precision allocation mode.

[0012] In some embodiments, the data types of the hardware parameters include numerical hardware parameters, categorical hardware parameters, and Boolean hardware parameters. The hardware parameters of the target device are classified and encoded to generate a unified hardware feature vector, including:

[0013] Numerical hardware parameters are standardized and encoded, categorical hardware parameters are One-Hot encoded, and Boolean hardware parameters are binary encoded;

[0014] The vectors obtained after processing the numerical hardware parameters, categorical hardware parameters and Boolean hardware parameters are concatenated to obtain a hardware feature vector.

[0015] In some embodiments, the reward function of the reinforcement learning agent is a multi-objective reward function, which includes: an accuracy loss penalty sub-function, a delay optimization reward sub-function, and a memory saving reward sub-function. The multi-objective reward function uniformly quantifies the accuracy loss penalty sub-function, the delay optimization reward sub-function, and the memory saving reward sub-function through a linear combination of dynamic weight coefficients.

[0016] In some embodiments, the mixed granularity pruning strategy includes:

[0017] Dynamically retain the Transformer layer of the model to be compressed based on the hardware computing power;

[0018] Prune the attention heads and hidden layer dimensions of the retained layers of the compressed model;

[0019] Replace operators not supported by the hardware and merge consecutive linear layers.

[0020] In some embodiments, adjusting the model structure of the model to be compressed according to the hybrid granularity pruning strategy, deleting redundant nodes, and reconnecting the computation graph includes:

[0021] Perform attention head pruning on the multi-head attention mechanism, adjust the parameter dimensions of the original number of attention heads according to the preset retention number, reconstruct the dimensions of the key, query and value projection matrices, and synchronously update the splicing logic of the multi-head attention output;

[0022] Perform neuron pruning on the hidden layer and reduce the dimension of the weight matrix by the number of retained neurons.

[0023] In some embodiments, adjusting the model structure of the model to be compressed according to the hybrid granularity pruning strategy, deleting redundant nodes and reconnecting the computation graph further includes:

[0024] When the target Transformer layer is removed, the layer normalization module and residual connection structure of the target Transformer layer are skipped, and the output nodes of the predecessor layer are directly connected to the input nodes of the successor layer;

[0025] Detect and merge continuous linear layer sequences, disconnect the internal connection edges between adjacent linear layers, merge multiple linear computing nodes into a single composite computing node, retain the input connection edges between the first-layer linear nodes and the upper-layer nodes, and the output connection edges between the last-layer linear nodes and the lower-layer nodes, and perform numerical reconstruction calculations on the weights of the merged connection edges.

[0026] In some embodiments, the mixed-precision allocation mode includes:

[0027] According to the bandwidth, the calculation data precision supported by the target device and the hardware native instruction set acceleration, the decision tree method is used to obtain the target quantization accuracy.

[0028] In some embodiments, quantizing the model parameters of the model to be compressed according to the mixed precision allocation mode includes:

[0029] The original data range is obtained according to the original minimum and maximum values ​​of the data to be quantized, and a symmetrical or asymmetrical target data range is set according to the type of target quantization accuracy;

[0030] A scaling factor is calculated based on a ratio of a difference between an extreme value of the original data range and a difference between an extreme value of the target data range, and a zero offset is determined according to a symmetric or asymmetric quantization mode, wherein the zero offset is zero in the symmetric quantization mode and the zero offset is calculated in the asymmetric quantization mode by a ratio between an extreme lower limit of the target data range and an extreme lower limit of the original data range;

[0031] The original data to be quantized is linearly transformed by the scaling factor and the zero offset and then rounded to obtain a quantized value of the target quantization accuracy.

[0032] In some embodiments, hardware parameters include theoretical computing power, number of computing cores, memory size, bandwidth, and supported data precision.

[0033] In some embodiments, the hardware-aware dynamic model compression system is used to perform any of the above-mentioned hardware-aware dynamic model compression methods.

[0034] The hardware-aware dynamic model compression method and system provided by the embodiments of the present disclosure can achieve the following technical effects:

[0035] The hardware parameters of the target device are classified and encoded to generate a unified hardware feature vector. The obtained hardware feature vector and the performance indicators of the model to be compressed are used as input to enable the trained reinforcement learning agent to dynamically generate an adaptation strategy, that is, output a mixed granularity pruning strategy and a mixed precision allocation mode. According to the mixed granularity pruning strategy, layered and progressive pruning is performed to adjust the model structure of the model to be compressed, delete redundant nodes and reconnect the calculation graph, and quantize the model parameters of the model to be compressed according to the mixed precision allocation mode to improve compression efficiency. This application realizes the coordination of model compression and hardware resources through a closed-loop framework of hardware parameter encoding, pruning strategy generation and mixed precision optimization, and generates a model compression solution suitable for deployed hardware.

[0036] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,

[0038] Figure 1 This is a flowchart of a hardware-aware dynamic model compression method provided by an embodiment of the present disclosure;

[0039] Figure 2 This is a framework diagram of hardware parameter encoding provided by an embodiment of the present disclosure;

[0040] Figure 3 Schematic diagram of the quantization precision decision process of the mixed precision allocation mode provided in an embodiment of the present disclosure;

[0041] Figure 4 is a flow chart of the model parameter quantization process according to an embodiment of the present disclosure;

[0042] Figure 5 This is an overall framework diagram of a hardware-aware dynamic model compression system provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0044] The terms "first," "second," and the like in the embodiments of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to facilitate the description of the embodiments of the present disclosure herein. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.

[0045] Unless otherwise stated, the term "plurality" means two or more.

[0046] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.

[0047] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0048] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0049] The existing mainstream large language models have a huge number of parameters. When deployed locally on devices with limited computing power and memory resources, model compression methods are usually used to reduce storage and computing requirements. Common model compression methods include model pruning and quantization. However, conventional model pruning and quantization have many problems. For example, traditional model pruning uses a fixed strategy that cannot adapt to the hardware differences of different devices. It is easy to cause excessive pruning on low-end devices, resulting in a sharp drop in accuracy, while high-end devices cannot fully utilize the computing power. Dynamic pruning based on input data requires additional computing resources to generate pruning masks during inference, which increases inference latency. In addition, model quantization schemes designed for one hardware architecture will greatly reduce hardware resource utilization under another hardware architecture due to hardware capability mismatch.

[0050] To solve the above problems, the present disclosure provides a hardware-aware dynamic model compression method and system. The model compression method provided by the present disclosure is described below with reference to the accompanying drawings.

[0051] Figure 1 This is a flow chart of a hardware-aware dynamic model compression method provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes the following steps:

[0052] S101: Classify and encode the hardware parameters of the target device to generate a unified hardware feature vector.

[0053] In some embodiments, the hardware parameters include theoretical computing power, number of computing cores, memory size, bandwidth, and supported data accuracy. In S101, the multi-dimensional hardware parameters are abstracted into computable feature vectors to achieve a unified representation of the hardware parameters of the target device, providing a decision basis for subsequent model compression.

[0054] Figure 2 This is a framework diagram of hardware parameter encoding provided by the embodiment of the present disclosure. Figure 2 As shown, the data types of the hardware parameters include numerical hardware parameters, categorical hardware parameters, and Boolean hardware parameters. Since the conversion and encoding methods of the three types of hardware parameters are different, the hardware parameters of the target device are classified and encoded to generate a unified hardware feature vector. Specifically, the numerical hardware parameters are first standardized (Z-Score) encoded, the categorical hardware parameters are One-Hot encoded, and the Boolean hardware parameters are binary encoded. Then, the vectors obtained after processing the numerical hardware parameters, the categorical hardware parameters, and the Boolean hardware parameters are spliced ​​to obtain the hardware feature vector.

[0055] The hardware parameter encoding process is described below with reference to specific examples.

[0056] For example, existing hardware device architectures include GPUs, CPUs, TPUs, and NPUs. Different devices support data precisions including FP64, FP32, FP16, BF16, FP8, INT8, and INT4. Consider a device with a GPU architecture, 500 computing cores, 32GB of video memory, and 400GB / s of bandwidth. The supported data precisions are FP32, FP16, and INT8, and the corresponding theoretical computing power for these precisions is 32TFLOPS, 64TFLOPS, and 128TOPS, respectively.

[0057] When encoding and converting device hardware parameters for this GPU architecture, the four parameters (number of cores, memory size, bandwidth, and theoretical computing power) are numerical parameters, the GPU architecture is a categorical parameter, and the supported data precision is a Boolean parameter. Based on the assumptions in the preconditions, the device type is GPU, and the categories are GPU, CPU, TPU, and NPU, so the encoding is [1,0,0,0]. The device supports FP32, FP16, and INT8, and the precision list is FP64, FP32, FP16, BF16, FP8, INT8, and INT4, so the supported data precision is encoded as [0,1,1,0,0,1,0]. Numerical parameters require standardization for uniform representation, which is specifically explained below.

[0058] Taking the number of computing cores as an example, the original value of the number of computing cores of the GPU architecture device is 500. Assuming that the average number of cores of all GPU architecture devices on the market is 3000 and the standard deviation is 1000, the formula for standardizing the number of computing cores is: Among them, μ is the mean and σ is the standard deviation, so the final value of the number of cores in the eigenvector is The processing method for other numerical parameters is similar to the calculation method for the number of computing cores. Finally, the vectors obtained after all parameter processing are spliced ​​to obtain the hardware feature vector for subsequent decision-making.

[0059] S102: Input the hardware feature vector and the performance index of the model to be compressed into a trained reinforcement learning agent, and output a mixed granularity pruning strategy and a mixed precision allocation mode.

[0060] In some embodiments, the state-action space of the reinforcement learning agent is designed as follows: the state input includes the hardware feature vector and the performance indicators (accuracy, inference latency, memory usage) of the model to be compressed, and the action output is a mixed granularity clipping strategy and a mixed precision allocation mode.

[0061] In some embodiments, the reward function of the reinforcement learning agent is a multi-objective reward function, which includes: an accuracy loss penalty sub-function, a delay optimization reward sub-function and a memory saving reward sub-function.

[0062] The mathematical expression of the multi-objective reward function is:

[0063] R = αR acc +βR latency +γR memory

[0064] In the formula, α, β, γ are dynamic weight coefficients, satisfying α+β+γ=1, R acc , R katency , R memory These are the precision loss penalty sub-function, the latency optimization reward sub-function, and the memory conservation reward sub-function. The multi-objective reward function quantizes these three sub-functions through a linear combination of dynamic weight coefficients. For example, if the device is a memory-sensitive mobile phone, γ > β > α can be set. If the device is a precision-sensitive server, α > β > γ can be set.

[0065] The mathematical expression of the precision loss penalty subfunction is:

[0066]

[0067] Where Acc baseIndicates the accuracy of the original model (model to be compressed) on the validation set, and λ indicates the accuracy sensitivity coefficient. If the business requires high fidelity, the value should be set larger, and low-precision scenarios can be set smaller. For example, if the original accuracy is 90% and the cropped accuracy is 85%, then R acc =-λ·(0.05).

[0068] The mathematical expression of the delay optimization reward subfunction is:

[0069]

[0070] Where, Latency base Indicates the average inference delay of the original model (model to be compressed), Latency current Represents the inference delay of the model after trimming and quantization, μ represents the delay sensitivity coefficient. If real-time interaction is required, the value should be set larger, and if real-time requirements are not high, the value should be set smaller. For example, if the original delay is 100ms and the optimized delay is 80ms, then R latency =μ·(0.2).

[0071] The mathematical expression of the memory saving reward sub-function is:

[0072]

[0073] Where, Memory base Indicates the memory size of the original model (model to be compressed), Memory current Indicates the memory usage of the pruned and quantized model. ν represents the memory sensitivity coefficient. For edge and terminal devices with limited memory, the value should be set larger, and for servers or cluster devices, the value should be set smaller. For example, if the original memory usage is 1GB and the optimized memory usage is 0.6GB, then Rmemory = ν·(0.4)

[0074] For example, if the accuracy of the compressed model on the validation set (which can be a public dataset) decreases compared to the performance indicators of the model to be compressed, points will be deducted, for example, 10 points will be deducted for every 1% decrease. Points will be awarded if the inference speed of the compressed model improves, for example, 5 points will be awarded for every 1 tokens / s increase. Points will be awarded if the memory usage of the compressed model decreases, for example, 3 points will be awarded for every 10% decrease.

[0075] In some embodiments, a hybrid granularity pruning strategy includes two types of granularity: layer-level pruning and operator-level pruning. Layer-level pruning further includes inter-layer pruning and intra-layer pruning. The hybrid granularity pruning strategy is described below.

[0076] In some embodiments, the inter-layer pruning action and the intra-layer pruning action are used as a joint action space for progressive pruning. First, coarse inter-layer pruning is performed, then the attention heads of the retained layers are pruned, and finally the hidden layer dimensions are compressed to approach the optimal configuration in stages. Specifically, inter-layer pruning is to dynamically retain the Transformer layer of the model to be compressed based on the hardware computing power. Here, low-computing power devices may prune more layers but retain the attention heads within the layers. For example, low-end devices can prune 50% of the layers, such as from 24 layers to 12 layers. Furthermore, the attention heads and hidden layer dimensions of the retained layers of the model to be compressed are pruned. For example, the number of attention heads of a certain layer is reduced from 16 to 8, and the hidden layer dimensions are compressed from 1024 to 512.

[0077] In some embodiments, the operator list that is not supported by the hardware is queried, the operator that is not supported by the hardware is replaced, and the continuous linear layers are merged. The following takes two continuous linear layers as an example to introduce the merging of continuous linear layers. Assume that the first layer: y = W1x + b1, the second layer: z = W2y + b2, the equivalent linear layer after the merger is to substitute the two layer formulas to obtain the merged expression: z = W2(W1x + b1) + b2 = (W2W1)x + (W2b1 + b2), the merged weight matrix is ​​W2W1, and the bias vector is W2b1 + b2.

[0078] The mixed-granularity pruning strategy is introduced above. Next, the mixed-precision allocation mode is introduced.

[0079] In some embodiments, a decision tree method is used to obtain the target quantization accuracy based on the bandwidth, the calculation data precision supported by the target device, and the hardware native instruction set acceleration.

[0080] Figure 3 This is a schematic diagram of the quantization precision decision process of the mixed precision allocation mode provided by the embodiment of the present disclosure. Figure 3 Let me explain the decision-making process: If the memory bandwidth is >200GB / s, the activation value retains a higher precision (select FP16 or BF16 based on the calculation data precision supported by the device); if it is <100GB / s, it is quantized to a lower precision (select INT8 or INT4, etc. based on the calculation data precision supported by the device). Next, if the memory bandwidth is high enough, after selecting a higher precision, it is also necessary to determine whether the device supports Tensor Core acceleration. If it does, matrix multiplication uses INT8; if it does not support it, but the device has other acceleration instructions or hardware units, the quantization precision is determined based on the data precision it supports. For example, if it has its own acceleration instructions, it may choose INT4 or other precision.

[0081] S103: Adjust the model structure of the model to be compressed according to the hybrid granularity pruning strategy, delete redundant nodes and reconnect the calculation graph.

[0082] In some embodiments, after model pruning, adjusting the computational graph includes: aligning parameter dimensions, which mainly involves intra-layer pruning; and processing cross-layer dependencies, which mainly involves inter-layer pruning and continuous linear layer merging.

[0083] The first is parameter dimension alignment. In some embodiments, attention head pruning is performed on the multi-head attention mechanism, and the parameter dimensions of the original number of attention heads are adjusted according to the preset retention number. The dimensions of the key, query, and value projection matrices are reconstructed, and the splicing logic of the multi-head attention output is updated synchronously. For example, assuming that the original number of attention heads is h, the key, query, and value projection matrices of each head are d_model*d_head, and k attention heads are pruned and retained, the parameter dimension is adjusted from d_model×(h×d_head) to d_model×(k×d_head), and the concat logic of the multi-head attention output is updated at the same time. For example, the original 16 attention heads are reduced to 8, the dimension of each head is reduced from 64 to 32, and the weight matrix is ​​changed from dmodel×1024 to dmodel×512.

[0084] Perform neuron pruning on the hidden layer, reducing the dimension of the weight matrix by the number of neurons retained. That is, if m neurons are retained, the weight dimension becomes d_model*m. For example, if the hidden layer dimension is reduced from 1024 to 512, the weight matrix becomes 512×512 instead of 1024×1024.

[0085] Next, cross-layer dependencies are processed. In some embodiments, when the target Transformer layer is removed, the layer normalization module and residual connection structure of the target Transformer layer are skipped, and the output nodes of the predecessor layer are directly connected to the input nodes of the successor layer. Continuous linear layer sequences are detected and merged, the internal connection edges between adjacent linear layers are disconnected, and multiple linear computing nodes are merged into a single composite computing node. The input connection edges between the first-layer linear nodes and the upper-layer nodes and the output connection edges between the last-layer linear nodes and the lower-layer nodes are retained, and the weights of the merged connection edges are numerically reconstructed. For example, if the original layer order is Layer1→Layer2→Layer3, if Layer2 is cut, the output of Layer1 is directly connected to the input of Layer3.

[0086] S104: quantizing the model parameters of the model to be compressed according to the mixed precision allocation mode.

[0087] Figure 4 This is a flow chart of the model parameter quantization process according to the embodiment of the present disclosure. Figure 4 , the quantization process includes the following steps:

[0088] S401: obtaining an original data range according to the original minimum value and maximum value of the data to be quantized, and setting a symmetrical or asymmetrical target data range according to the type of the target quantization accuracy.

[0089] The first step is to determine the quantization range. The original data range needs to count the minimum value X of the data to be quantified. min and the maximum value X max ; The target precision range is determined by the data precision. For example, the symmetric quantization range of INT8 is [-127,127], and the asymmetric quantization range is [0,255]. The target minimum and maximum values ​​are set as Q min and Q max .

[0090] S402: Calculate a scaling factor based on the ratio of the extreme value difference between the original data range and the target data range, and determine a zero-point offset according to a symmetric or asymmetric quantization mode, wherein the zero-point offset is zero in the symmetric quantization mode, and the zero-point offset is calculated by the proportional relationship between the extreme lower limit of the target data range and the extreme lower limit of the original data range in the asymmetric quantization mode.

[0091] Then calculate the quantization parameter, and the scaling factor is calculated by dividing the difference between the maximum and minimum values ​​of the original data range by the difference between the maximum and minimum values ​​of the target data range. max -X min ) / (Q max -Q min The zero offset is used to align the original data with the zero point after quantization. In symmetric quantization, the zero offset value is 0. In asymmetric quantization, the zero offset value is the minimum value in the target data range minus the minimum value in the original data range and the quotient of the scaling factor Z = Q. min -X min / scale.

[0092] S403: performing linear transformation on the original data to be quantized by using the scaling factor and the zero offset and then rounding the data to obtain a quantized value of the target quantization accuracy.

[0093] Finally, quantization mapping is performed to map the original data x to the quantized value q, where q=round(x / scale+Z).

[0094] Based on the same inventive concept as the above-mentioned hardware-aware dynamic model compression method, the present application also discloses a hardware-aware dynamic model compression system in some embodiments. Figure 5 This is an overall framework diagram of a hardware-aware dynamic model compression system provided by an embodiment of the present disclosure. Figure 5As shown, first, for the input part, the hardware feature vector is obtained through the hardware feature encoder as the basis for subsequent strategy formulation. Another input is the performance index (accuracy, inference delay, memory usage) of the model to be compressed. Then, the data is input to the reinforcement intelligent learning agent, which feeds back a hybrid granularity clipping strategy and a hybrid precision allocation mode. Here, the reward function of the reinforcement learning agent is a multi-objective reward function, which includes: an accuracy loss penalty sub-function, a delay optimization reward sub-function, and a memory saving reward sub-function. Secondly, the model reconstruction module adjusts the model structure of the model to be compressed according to the hybrid granularity clipping strategy, deletes redundant nodes and reconnects the calculation graph. The model parameters of the model to be compressed are quantized according to the hybrid precision allocation mode. Finally, the compressed model is reconstructed.

[0095] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient storage medium.

[0096] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only used to describe the embodiments and are not used to limit the scope of protection. As used in the description in the text, unless the context clearly indicates otherwise, the singular forms of "a", "an" and "the" are intended to also include plural forms. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be found in the description of the method part.

[0097] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0098] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present disclosure may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

Claims

1. A hardware-aware dynamic model compression method, characterized in that: The method comprises: Classify and encode the hardware parameters of the target device to generate a unified hardware feature vector; Input the hardware feature vector and the performance index of the model to be compressed into a trained reinforcement learning agent, and output a mixed granularity pruning strategy and a mixed precision allocation mode; According to the hybrid granularity pruning strategy, the model structure of the model to be compressed is adjusted, redundant nodes are deleted, and the computation graph is reconnected; According to the mixed precision allocation mode, the model parameters of the model to be compressed are quantized.

2. The hardware-aware dynamic model compression method according to claim 1, characterized in that: The data types of the hardware parameters include numerical hardware parameters, categorical hardware parameters, and Boolean hardware parameters. The hardware parameters of the target device are classified and encoded to generate a unified hardware feature vector, including: Performing standardized encoding on the numerical hardware parameters, One-Hot encoding on the categorical hardware parameters, and binary encoding on the Boolean hardware parameters; The vectors obtained after processing the numerical hardware parameters, the categorical hardware parameters, and the Boolean hardware parameters are concatenated to obtain the hardware feature vector.

3. The hardware-aware dynamic model compression method according to claim 1, characterized in that: The reward function of the reinforcement learning agent is a multi-objective reward function, which includes: an accuracy loss penalty sub-function, a delay optimization reward sub-function and a memory saving reward sub-function. The multi-objective reward function uniformly quantifies the accuracy loss penalty sub-function, the delay optimization reward sub-function and the memory saving reward sub-function through a linear combination of dynamic weight coefficients.

4. The hardware-aware dynamic model compression method according to claim 1, characterized in that: The hybrid granularity clipping strategy includes: Dynamically retain the Transformer layer of the model to be compressed based on the hardware computing power; Pruning the attention heads and hidden layer dimensions of the retained layers of the model to be compressed; Replace operators not supported by the hardware and merge consecutive linear layers.

5. The hardware-aware dynamic model compression method according to claim 4, characterized in that: The adjusting the model structure of the to-be-compressed model according to the hybrid granularity clipping strategy, deleting redundant nodes and reconnecting the computation graph includes: Perform attention head pruning on the multi-head attention mechanism, adjust the parameter dimensions of the original number of attention heads according to the preset retention number, reconstruct the dimensions of the key, query and value projection matrices, and synchronously update the splicing logic of the multi-head attention output; Perform neuron pruning on the hidden layer and reduce the dimension of the weight matrix by the number of retained neurons.

6. The hardware-aware dynamic model compression method according to claim 4, characterized in that: The step of adjusting the model structure of the model to be compressed according to the hybrid granularity pruning strategy, deleting redundant nodes, and reconnecting the computation graph further includes: When the target Transformer layer is removed, the layer normalization module and residual connection structure of the target Transformer layer are skipped, and the output nodes of the predecessor layer are directly connected to the input nodes of the successor layer; Detect and merge continuous linear layer sequences, disconnect the internal connection edges between adjacent linear layers, merge multiple linear computing nodes into a single composite computing node, retain the input connection edges between the first-layer linear nodes and the upper-layer nodes, and the output connection edges between the last-layer linear nodes and the lower-layer nodes, and perform numerical reconstruction calculations on the weights of the merged connection edges.

7. The hardware-aware dynamic model compression method according to claim 1, characterized in that: The mixed precision allocation mode includes: According to the bandwidth, the calculation data precision supported by the target device and the hardware native instruction set acceleration, the decision tree method is used to obtain the target quantization accuracy.

8. The hardware-aware dynamic model compression method according to claim 7, characterized in that: The quantizing of the model parameters of the to-be-compressed model according to the mixed precision allocation mode includes: Obtaining an original data range according to the original minimum and maximum values ​​of the data to be quantized, and setting a symmetrical or asymmetrical target data range according to the type of the target quantization accuracy; Calculating a scaling factor based on a ratio of an extreme value difference between the original data range and the target data range, and determining a zero point offset according to a symmetric or asymmetric quantization mode, wherein the zero point offset is zero in the symmetric quantization mode, and the zero point offset is calculated by a proportional relationship between an extreme value lower limit of the target data range and an extreme value lower limit of the original data range in the asymmetric quantization mode; The original data of the data to be quantized is linearly transformed by the scaling factor and the zero offset and then rounded to obtain a quantized value of the target quantization accuracy.

9. The hardware-aware dynamic model compression method according to claim 1, characterized in that: The hardware parameters include theoretical computing power, number of computing cores, memory size, bandwidth, and supported data accuracy.

10. A hardware-aware dynamic model compression system, characterized in that: The system is used to execute the hardware-aware dynamic model compression method described in any one of claims 1 to 9.