Mixing precision quantification of artificial intelligence models

By performing perturbation and sensitivity analysis on the multi-layer weights of the AI ​​model, dynamic allocation bit accuracy is quantified through mixed accuracy, which solves the balance problem of accuracy and efficiency in model quantization in the prior art, and realizes efficient computing and storage on electronic devices.

CN120051779APending Publication Date: 2025-05-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070280.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-03
Filing Date
2023-12-29
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to optimize computing efficiency, memory usage and power consumption while maintaining high accuracy, especially when directly training and deployment on electronic devices.

Method used

By perturbing the weights of multiple layers of the AI ​​model for a predefined number of times, the output variation and sensitivity of each layer are determined, the appropriate bit accuracy is allocated based on this information, and the mixed precision quantization is performed using this bit accuracy.

Benefits of technology

The optimal performance of each layer within the multi-layer of the AI ​​model on the electronic device is achieved, reducing computing requirements and power consumption, while optimizing memory usage and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051779A_ABST
    Figure CN120051779A_ABST
Patent Text Reader

Abstract

A method for mixing precision quantification of an artificial intelligence (AI) model by an electronic device is included. The method includes: performing, by an electronic device, a predefined number of perturbations on a weight of each of a plurality of layers of an AI model; determining, by the electronic device, a change in an output of each of the plurality of layers of the AI model based on the perturbation of the weight of each of the plurality of layers; determining, by the electronic device, a sensitivity metric for each of the plurality of layers of the AI model as a metric for a change in the output of each layer; assigning, by the electronic device, a bit accuracy to each of a plurality of layers of the AI model based on the determined sensitivity metric; and performing, by the electronic device, mixing precision quantization of the AI model using the bit precision allocated to each of the plurality of layers of the AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to artificial intelligence (AI) model compression for electronic devices. More specifically, the present disclosure relates to mixed precision quantization of an AI model. Background Art

[0002] The scale of neural networks and the speed / capacity of their reasoning have posed significant challenges to their deployment in electronic devices. A promising solution to these problems is quantization. However, uniformly quantizing a model to ultra-low precision can result in considerable loss of precision. The parameters of a neural network, which are typically over-parameterized and represented by high bit precision (such as 32-bit floating point numbers), provide ample opportunity to reduce the bit precision to values ​​such as 16-bit, 8-bit, or 4-bit. Reducing the bit precision of parameters can improve memory usage, improve performance, and reduce power consumption. Quantization can be applied uniformly across all layers of a neural network, which may not always lead to an optimal solution, or unique bit precision can be applied to each layer of the network to obtain more ideal results.

[0003] Prior art for implementing mixed precision quantization has attempted to use either search-based or criterion-based approaches. Search-based approaches (such as those utilizing reinforcement learning) can be time-consuming, while criterion-based approaches rely on second-order approximations (such as approximating the Hessian trajectory).

[0004] Mixed precision techniques can potentially reduce the running time and memory requirements of neural networks by appropriately assigning the right data type to each operation. This can be achieved through reinforcement learning or regression analysis. Reinforcement learning (the discipline that studies decision-making techniques) is often used in this process. A unique aspect of the mixed precision quantization service is its ability to iteratively increase object size. The service can initially assume that all objects (including weights) should be converted to smaller data types (such as 8-bit integers).

[0005] The mixed precision quantization service can select objects in an iterative manner to increase from a smaller data type to a larger data type (e.g., from an 8-bit integer to a 12-bit integer) while ensuring that the target object bandwidth is not exceeded. The service uses a carefully ordered set of objects generated in a block to determine which objects should be size-enhanced. The mixed precision quantization service can start with objects that consume less bandwidth and gradually increase the size of objects that consume higher bandwidth.

[0006] Neural networks, such as artificial neural networks (ANNs), typically use various normal precision floating point formats, including 16-bit, 32-bit, 64-bit, and 80-bit floating point formats, for their internal calculations. The process of training an ANN can be demanding in terms of both computation and storage, requiring billions of operations and gigabytes of storage. There are methods to optimize neural network performance, power consumption, and storage requirements, such as using quantized precision floating point formats during training and / or inference. These formats may require a reduction in bit width, involving using fewer bits to represent the mantissa and / or exponent of a number, or a block floating point (BFP) format that uses a limited mantissa of 3, 4, or 5 bits and an exponent shared by two or more numbers.

[0007] The use of quantized precision formats can have adverse effects on neural networks, resulting in reduced accuracy and other potential damage. The method requires the use of second-order techniques such as the hessian trajectory defined as the divergence of the gradient field in Riemannian geometry. The network can learn the most effective behavior to perform in a specific environment to maximize the reward. However, it is worth noting that mixed precision quantization does not rely on first-order techniques.

[0008] The above information is presented as background information only to assist with understanding the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above may be applicable as prior art with respect to the present disclosure. Summary of the invention

[0009] Various aspects of the present disclosure are to at least solve the above problems and / or disadvantages and to provide at least the advantages described below. Therefore, one aspect of the present disclosure is to provide mixed precision quantization of artificial intelligence models.

[0010] Another aspect of the present disclosure is to determine an AI model for mixed precision quantization based on the bit precision assigned to each layer of a plurality of layers.

[0011] Another aspect of the present disclosure is to achieve optimal performance for each layer within multiple layers of an AI model in all aspects of power consumption, memory allocation, computational efficiency, and on-device learning on an electronic device.

[0012] Another aspect of the present disclosure is to provide an optimal configuration for quantization.

[0013] Another aspect of the present disclosure is to reduce computational requirements and result in lower power consumption.

[0014] Another aspect of the present disclosure is to facilitate direct training quantization and deployment on electronic devices, thereby minimizing memory usage and optimizing computational and power efficiency.

[0015] Another aspect of the present disclosure is to perturb the weight of each layer of a plurality of layers of an AI model a predefined number of times, and estimate a change in the output of each layer by calculating an average gradient value as a result of the perturbation.

[0016] Another aspect of the present disclosure is to determine the sensitivity of each layer within the AI ​​model, which is a measure of the expected variation in output observed at each layer.

[0017] Another aspect of the present disclosure is to assign specific bit precisions to various layers based on their determined sensitivity, and then employ the bit precision assignments to quantize the AI ​​model.

[0018] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.

[0019] According to one aspect of the present disclosure, a method for performing mixed precision quantization on an artificial intelligence (AI) model by an electronic device is provided. The method includes: performing a predetermined number of perturbations on the weights of each of the multiple layers of the AI ​​model by the electronic device; determining the change in the output of each of the multiple layers of the AI ​​model based on the perturbation of the weights of each of the multiple layers by the electronic device; determining the sensitivity metric of each of the multiple layers of the AI ​​model by the electronic device as a measure of the change in the output of each layer; allocating bit precision to each of the multiple layers of the AI ​​model by the electronic device based on the determined sensitivity metric; and performing mixed precision quantization of the AI ​​model by the electronic device using the bit precision allocated to each of the multiple layers of the AI ​​model.

[0020] In another aspect, a change in output is determined by determining a loss gradient for each of the multiple layers based on a perturbation weight for each of the multiple layers, and determining a change in output for each of the multiple layers of the AI ​​model based on the loss gradient for each of the multiple layers. The change in output indicates a loss for each of the multiple layers.

[0021] In another aspect, a sensitivity metric for each of a plurality of layers of the AI ​​model is determined based on a loss gradient.

[0022] On the other hand, the step of allocating bit precision to each of the multiple layers of the AI ​​model by the electronic device based on the sensitivity metric includes: the electronic device uses the sensitivity metric and the net compression ratio of each of the multiple layers to construct a constrained optimization problem model, and the electronic device allocates bit precision to each of the multiple layers based on the constrained optimization problem model.

[0023] On the other hand, the step of performing mixed precision quantization of the AI ​​model by an electronic device using the bit precision assigned to each of the multiple layers of the AI ​​model includes: performing post-training quantization of the AI ​​model by using the bit precision assigned to each of the multiple layers, and enabling each of the multiple layers to be at the assigned bit precision to obtain an optimal mixed precision quantized AI model by the electronic device.

[0024] On the other hand, the optimal AI model obtains the best performance for each of the multiple layers of the AI ​​model in terms of power levels, memory usage, computational efficiency levels, and / or on-device learning on the electronic device.

[0025] In another aspect, bit precision is assigned to each of a plurality of layers of the AI ​​model by selecting at least one bit from a set of bit precisions based on a sensitivity metric.

[0026] According to another aspect of the present disclosure, an electronic device for mixed precision quantization of an artificial intelligence (AI) model is provided. The electronic device includes a memory, one or more processors, and a mixed precision quantization controller communicatively coupled to the memory and the one or more processors, wherein the memory stores one or more computer programs including computer executable instructions, which, when executed by the one or more processors, cause the electronic device to perform the following operations: performing a predefined number of perturbations on the weights of each of the multiple layers of the AI ​​model; determining the change in the output of each of the multiple layers of the AI ​​model based on the perturbation of the weights of each of the multiple layers; determining the sensitivity metric of each of the multiple layers of the AI ​​model as a measure of the change in the output of each layer; and assigning bit precision to each of the multiple layers of the AI ​​model based on the determined sensitivity metric; and performing quantization of the AI ​​model using the bit precision assigned to each of the multiple layers of the AI ​​model.

[0027] According to another aspect of the present disclosure, one or more non-transitory computer-readable storage media are provided, which store one or more computer programs including computer executable instructions, which, when executed by one or more processors of an electronic device for mixed precision quantization of an artificial intelligence (AI) model, cause the electronic device to perform operations. The operations include: performing a predetermined number of perturbations on the weights of each of the multiple layers of the AI ​​model by the electronic device; determining the change in the output of each of the multiple layers of the AI ​​model based on the perturbation of the weights of each of the multiple layers by the electronic device; determining the sensitivity metric of each of the multiple layers of the AI ​​model by the electronic device as a measure of the change in the output of each layer; assigning bit precision to each of the multiple layers of the AI ​​model by the electronic device based on the determined sensitivity metric; and performing mixed precision quantization of the AI ​​model by the electronic device using the bit precision assigned to each of the multiple layers of the AI ​​model.

[0028] Other aspects, advantages, and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description which, in conjunction with the accompanying drawings, discloses various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings, in which:

[0030] Figure 1 A block diagram of an electronic device for performing mixed precision quantization of an AI model according to an embodiment of the present disclosure is shown;

[0031] Figure 2 is a block diagram showing a mixed precision quantization controller integrated in an electronic device according to an embodiment of the present disclosure, demonstrating advanced mixed precision quantization technology applied to AI models;

[0032] Figure 3 is a schematic diagram illustrating mixed precision quantization of an AI model according to an embodiment of the present disclosure;

[0033] Figure 4 shows a graphical depiction of a loss function of a loss curve according to an embodiment of the present disclosure, illustrating convergence points within the loss function;

[0034] Figure 5A shows a graphical representation of a loss curve with respect to weights associated with a particular layer (i.e., layer-1) according to an embodiment of the present disclosure;

[0035] Figure 5B shows a graphical representation of a loss curve associated with weights of layer-2 according to an embodiment of the present disclosure;

[0036] Figure 5C shows a graphical representation of computing the average gradient norm by perturbations of weights in N random directions according to an embodiment of the present disclosure;

[0037] Figure 6 A graphical representation showing an estimate of the variation in loss across different bit precision levels applicable to a given layer according to an embodiment of the present disclosure is shown;

[0038] Figure 7 is a flowchart illustrating a method for mixed precision quantization of an AI model based on average gradient norm according to an embodiment of the present disclosure; and

[0039] Figure 8 is a flowchart illustrating a method for mixed precision quantization of an AI model according to an embodiment of the present disclosure.

[0040] Throughout the drawings, like reference numerals will be understood to refer to like parts, components and structures.

[0041] In addition, it will be understood by those of ordinary skill in the art that the elements in the drawings are shown for simplicity and may not necessarily be drawn to scale. For example, the dimensions of some elements in the drawings may be exaggerated relative to other elements to help improve understanding of various aspects of the present disclosure. In addition, one or more elements may have been represented in the drawings by conventional symbols, and the drawings may only show those specific details relevant to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that would be readily understood by those of ordinary skill in the art having the benefit of the description herein. DETAILED DESCRIPTION

[0042] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details are considered to be exemplary only. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0043] The terms and words used in the following description and claims are not limited to the bibliographical meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it is apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only and not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.

[0044] It will be understood that singular forms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.

[0045] The various embodiments described herein are not necessarily mutually exclusive, as some embodiments may be combined with one or more other embodiments to form new embodiments. Unless otherwise specified, the term "or" used herein refers to a non-exclusive or. The examples used herein are intended only to facilitate understanding of the manner in which the embodiments of this article may be practiced, and also to enable those skilled in the art to practice the embodiments of this article. The examples should not be interpreted as limiting the scope of the embodiments of this article.

[0046] According to conventional practice in the art, embodiments may be described and illustrated according to blocks that perform one or more functions described. These blocks (which may be referred to herein as managers, units, modules, hardware components, etc.) are physically implemented by analog circuits and / or digital circuits (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, etc.), and may optionally be driven by firmware. The circuit may be implemented in one or more semiconductor chips, or implemented on a substrate support (such as a printed circuit board, etc.). The circuits constituting the blocks may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuits), or by a combination of dedicated hardware for performing some functions of the blocks and processors for performing other functions of the blocks. Without departing from the scope of the present disclosure, each block of the embodiment may be physically divided into two or more interacting and discrete blocks. Without departing from the scope of the present disclosure, the blocks of the embodiment may be physically combined into more complex blocks.

[0047] The accompanying drawings are used to help easily understand various technical features, and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. In addition to the content specifically described in the accompanying drawings, the present disclosure should also be interpreted as extending to any changes, equivalents and alternative forms. Although the terms "first", "second" etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are usually only used to distinguish one element from another element.

[0048] An embodiment discloses a method for mixed precision quantization of an AI model. The method includes performing a predefined number of perturbations on the weights of each of the multiple layers of the AI ​​model. The method includes determining a change in the output of each of the multiple layers based on the perturbation of the weights of each of the multiple layers of the AI ​​model. For example, the method includes determining a sensitivity metric for each of the multiple layers of the AI ​​model as a measure of the change in the output of each layer, and assigning bit precision to each of the multiple layers of the AI ​​model based on the determined sensitivity metric. For example, the method includes determining a mixed precision quantized AI model based on the bit precision assigned to each of the multiple layers.

[0049] Based on the proposed method, quantization of the model requires allocating different bit widths to different layers through computationally efficient first-order techniques such as linear computational complexity. This approach ensures optimal performance in terms of power, memory, and latency while maintaining accuracy. The first-order method only employs gradient information to determine the mixed precision setting of the neural network, enabling the generation of quantized networks within 1.5% of the original model accuracy in much faster computation time. The proposed method does not assume that the convergence points of the pre-trained neural network represent local minima. It is more general than prior art methods because it considers the convergence points as both convergence points and local minima.

[0050] Quantization is a widely adopted technique in deep learning that effectively minimizes neural network memory requirements and computational complexity while maintaining an acceptable level of accuracy. The method entails encoding the parameters and activations of deep neural networks (DNNs) using lower precision numerical representations, thereby achieving a significant reduction in memory bandwidth and usage. Quantization has the potential to significantly enhance the performance of electronic devices such as mobile phones, servers, and smart watches, as well as more resource-constrained devices such as IoT devices and autonomous vehicles.

[0051] Mixed Precision Quantization (MPQ) is a novel quantization technique that allows different bit widths to be used for different layers within a DNN model. This is different from traditional quantization methods, in which all layers are restricted to the same bit width. MPQ technology helps achieve optimal performance, for example in terms of power consumption, memory usage, and latency, without compromising the accuracy of the model. Each layer can be configured to operate with a suitable bit precision selected from a set of options (such as 2 bits, 4 bits, 8 bits, 16 bits, etc.). This flexibility enables the creation of highly optimized models in which each layer operates at the most appropriate level of bit precision.

[0052] MPQ exploits the trade-off between accuracy and computational efficiency while deploying DNN models. Given that DNNs require millions or even billions of parameters, the need for large amounts of memory storage is essential. The use of MPQ can potentially provide the best configuration for quantization. By quantizing parameters and activations to a lower format, memory resources are significantly saved, which is particularly useful for resource-constrained devices.

[0053] The proposed method and electronic device demonstrate that lower precision formats may require less storage and computational resources. DNN models can be operated faster, ultimately enhancing reasoning. This aspect is of vital importance in real-time applications requiring low latency.

[0054] DNN models typically require significant computational power, resulting in increased power consumption. As described in the proposed method, the utilization of the MPQ method effectively alleviates the computational requirements, thereby reducing power consumption. This aspect is particularly critical for electronic devices to adapt AI models while running DNN models seamlessly in the background.

[0055] Compared to cloud-based servers, electronic devices tend to be memory-constrained and have relatively limited computational resources. The proposed method entails the use of MPQ, which facilitates training, quantization, and deployment directly on electronic devices due to its compact memory footprint and improved computational and power efficiency. MPQ supports on-device learning methods and can be used for model deployment on dedicated hardware. According to the proposed method, various hardware supports provide different accuracy levels on electronic devices.

[0056] The proposed disclosure demonstrates versatility for various multi-model applications, including but not limited to selfie enhancement, bokeh implementation, gaming applications, expert-level raw denoising, and image restoration.

[0057] It should be understood that the blocks in each flowchart and the combination of flowcharts can be performed by one or more computer programs comprising instructions. The entirety of one or more computer programs can be stored in a single memory, or one or more computer programs can be divided into different parts, wherein the different parts are stored in different multiple memories.

[0058] Any of the functions or operations described herein may be processed by a processor or a combination of processors. A processor or a combination of processors is a circuit that performs processing and includes circuits such as an application processor (AP, such as a central processing unit (CPU)), a communication processor (CP, such as a modem), a graphics processing unit (GPU), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a Wi-Fi chip, Chips, Global Positioning System (GPS) chips, Near Field Communication (NFC) chips, connectivity chips, sensor controllers, touch controllers, fingerprint sensor controllers, display driver integrated circuits (ICs), audio codec chips, Universal Serial Bus (USB) controllers, camera controllers, image processing ICs, microprocessor units (MPUs), systems on chip (SoCs), integrated circuits (ICs), etc.

[0059] Figure 1 A block diagram of an electronic device (101) for performing mixed precision quantization of an AI model according to an embodiment of the present disclosure is shown. The electronic device (101) includes a memory (102), a processor (103), a communicator (104), and a mixed precision quantization controller (105).

[0060] The memory (102) stores instructions to be executed by the processor (103). The memory (102) includes a non-volatile storage element. Examples of such non-volatile storage elements include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or an electrically programmable memory (EPROM) or an electrically erasable programmable memory (EEPROM). In addition, in some examples, the memory (102) is considered to be a non-transitory storage medium. The term "non-transitory" indicates that the storage medium is not embodied in a carrier or propagating signal. The term "non-transitory" is not to be interpreted as the memory (102) being non-removable. The memory (102) stores larger amounts of information. In a specific example, the non-transitory storage medium stores data that can change over time (e.g., in a random access memory (RAM) or a cache). The processor (103) includes one or more processors.

[0061] The one or more processors (103) are general-purpose processors (such as a central processing unit (CPU), an application processor (AP), etc.), graphics processing units (such as graphics processing units (GPUs), visual processing units (VPUs)), and / or AI-specific processors (such as neural processing units (NPUs)). In an embodiment, the processor (103) includes multiple cores and runs instructions stored in the memory (102).

[0062] One or more processors control the processing of input data according to predefined operating rules or AI models stored in non-volatile memory and volatile memory. Predefined operating rules or artificial intelligence models are provided through training or learning.

[0063] Providing by learning means forming a predefined operating rule or AI model of desired characteristics by applying a learning method to a plurality of learning data. Learning may be performed in the device itself that executes the AI ​​according to the embodiment, and / or may be implemented by a separate server / system.

[0064] In another embodiment, the AI ​​model may be composed of multiple neural network layers. Each layer has multiple weight values, and the layer operation is performed by the calculation of the previous layer and the operation of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.

[0065] A learning method is a method for training a predetermined target device (e.g., an edge device) using a plurality of learning data to enable, allow, or control the target device to make a determination or prediction. Examples of learning methods include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0066] In an embodiment, the communicator (104) includes electronic circuits dedicated to implementing standard wired or wireless communication. The communicator (104) communicates internally between internal hardware components of the electronic device (101) and communicates with external devices via one or more networks.

[0067] In another embodiment, the mixed precision quantization controller (105) is a complex hardware entity that is configured to be responsible for directing and managing the flow of data between two different entities. In the computing field, this device may take the form of a microchip, card, or other separate hardware mechanism, each designed to supervise the operation of various electronic devices (101). The mixed precision quantization controller (105) acts as an intermediary linking the two electronic devices, actively managing and directing the communication between the devices.

[0068] The mixed precision quantization controller (105) controls each layer (301a-n) (or layers (301a-n)) (i.e., Figure 3 The weights of the electronic device (101) (301a, 301b, 301c, 301d, 301e ... 301n) are perturbed for a predefined number of cycles. Such a cycle is established by the electronic device (101) or its user.

[0069] Based on each of the multiple layers (301 a-n ), the mixed precision quantization controller (105) determines each layer (301) of the multiple layers of the AI ​​model a-n ) output changes. In an embodiment, the mixed precision quantization controller (105) is based on each layer (301) of the multiple layers. a-n ) to determine the perturbation weight of each layer in the plurality of layers (301 a-n ). In addition, the mixed precision quantization controller (105) is based on each layer (301) in the plurality of layers. a-n ) to determine the loss gradient of each of the multiple layers of the AI ​​model (301 a-n ). The change in the output indicates the loss for each of the multiple layers.

[0070] The mixed precision quantization controller (105) determines each of the multiple layers of the AI ​​model (301 a-n ) as a measure of the change in the output of each layer. Determine each layer (301) of the multiple layers of the AI ​​model based on the loss gradient a-n ) is a sensitivity measure.

[0071] By utilizing the determined sensitivity metric, the mixed precision quantization controller (105) intelligently assigns precise bit precision to each layer (301) within the plurality of layers comprising the AI ​​model.a-n In an embodiment, the mixed precision quantization controller (105) uses the sensitivity metric of each layer in combination with the net compression ratio to formulate a constrained optimization problem model. Thereafter, the mixed precision quantization controller (105) uses the optimization model to assign appropriate bit precision to each layer within the plurality of layers (301 a-n ). The sensitivity measure is crucial to this process as it allows the selection of at least one bit to be assigned to each layer from a set of possible bit accuracies.

[0072] The mixed precision quantization controller (105) performs quantization of the AI ​​model using the allocated bit precision assigned to each of the plurality of layers constituting the AI ​​model. The mixed precision quantization controller (105) helps each of the plurality of layers (301 a-n ) is run with its specified bit precision, and finally produces the best mixed precision quantized AI model. The mixed precision quantization controller (105) uses the bit precision for each layer (301) of the multiple layers. a-n ) to run the post-training quantization of the AI ​​model with the allocated bit precision specified. The optimal AI model generated thereby achieves each of the multiple layers (301) in one or more aspects such as power usage, memory consumption, computational efficiency, and on-device learning capability on the electronic device (101). a-n )’s peak performance level.

[0073] Figure 2 1 is a block diagram showing a mixed precision quantization controller (105) integrated in an electronic device (101) according to an embodiment of the present disclosure, showing an advanced mixed precision quantization technique applied to an AI model. The mixed precision quantization controller (105) includes an AI model (201) (e.g., an FP32 model), a layer-by-layer sensitivity calculator (202), a bit precision allocator (203), a trained quantization model (204), and a mixed precision quantization model (205). The layer-by-layer sensitivity calculator (202) includes a layer iterator (206), a weight perturbator (207), and a gradient calculator (208). The bit precision allocator (203) includes a loss estimator (209) and a bit allocator (210). The loss estimator (209) (locally) estimates the possible loss for allocating different bit precisions to each layer. In the example, if the loss estimator (209) allocates 2 bits, 4 bits, or 8 bits to the AI ​​model (201), layer-2, layer-L4, and layer-L8 represent possible losses. The bit allocator (210) is a constraint-based optimizer. The bit allocator (210) considers the total loss from the loss estimator (209) and the six constraints of the model to decide the final bit configuration for each layer.

[0074] The AI ​​model provides input to a layer-by-layer sensitivity calculator (202). The layer-by-layer sensitivity calculator (202) determines the sensitivity of each layer of the AI ​​model one by one. Sensitivity refers to the gradient norm. In an embodiment, Figure 5A and Figure 5B The sensitivity of each layer is explained in . The layer iterator (206) iterates over all the layers of the AI ​​model one by one to calculate the sensitivity. The weight perturbator (207) modifies the weights of the selected layer by adding some random noise in the weight vector. On the AI ​​model with perturbed weights, the gradient calculator (208) performs forward propagation and then performs back propagation and calculates the gradient norm of the selected layer through the layer iterator. The average gradient norm (sensitivity) is calculated as follows:

[0075]

[0076] The sensitivity of each layer of the AI ​​model is fed to the bit precision allocator (203). The bit precision allocator (203) calculates, for example, the average gradient norm (i.e., sensitivity), which represents the general rate of change of the loss due to small perturbations. Figure 6 The operation of the bit precision allocator (203) is explained in. The calculated average gradient norm is provided to the trained quantized model (204).

[0077] Model post-training quantization (204) is based on the provided bit configuration, combined with layer-by-layer bit allocation to perform quantization of the AI ​​model. As a result, the system generates a mixed precision quantized model (205), which represents a neural network model quantized with different bit widths for different layers.

[0078] Figure 3 1 is an example schematic diagram (S300) illustrating mixed precision quantization of an AI model according to an embodiment of the present disclosure. The electronic device (101) perturbs each layer (301) of the AI ​​model a predefined number of times. a-n ). In addition, the electronic device (101) estimates the change of the output of each layer as a result of the perturbation. For example, the change of the output of each layer is estimated by calculating the average gradient value after perturbing the weight of the layer a predefined number of times. Each layer (301 a-n ) is proportional to the average gradient value and indicates the slope of the loss curve for each layer. The slope of the loss curve provides the sensitivity of the layer. As a measure of the estimated change in the output of each layer, the electronic device (101) determines each layer (301) of the AI ​​model a-n ). In response to the determined sensitivity, the electronic device (101) allocates bit precision to each layer and performs quantization on the AI ​​model using the allocated bit precision. The bit precision allocation is proportional to the sensitivity of the layer.

[0079] Mixed precision quantization of the AI ​​model is performed by using downsampling, fully connected (FC) layers, Softmax, and convolutional layers (conv). Downsampling reduces the resolution of the image. In one example, 128×128 reduces the resolution of the image to 64×64. Fully connected (FC) layers and convolutional layers are used in neural networks. In Softmax, a type of layer in a neural network normalizes the values ​​of some intermediate outputs. Downsampling, FC layers, Softmax, and conv are known to those skilled in the art. For the sake of brevity, we do not explain this in the patent disclosure.

[0080] Figure 4 A graphical depiction of a loss function according to an embodiment of the present disclosure is shown (S400), showing convergence points within the loss function. Convergence points are also called saddle points.

[0081] Figure 4 The loss curve (401) and the convergence point (402) are shown. The loss curve (401) indicates that the loss function has a more obvious drop near the convergence point (402), indicating an increase in sensitivity. In addition, Figure 4 A depiction of a loss curve (401) associated with weights for a particular layer is shown.

[0082] Reference Figure 4 The proposed disclosure uniquely enables more accurate measurement of multiple layers (301 a-n ) at the convergence point (402), surpassing the limitations of conventional prior art techniques that rely solely on linear computations.

[0083] Loss function: L = f(w) = w 3 …The gradient of the loss function in equation 1:

[0084] Hessian (2nd order derivative) of the loss function:

[0085] According to the prior art, the sensitivity of the loss function around ∈ is calculated using the average Hessian trajectory as follows:

[0086]

[0087] Sensitivity using the proposed method:

[0088]

[0089] The sensitivity metrics proposed in Equation 4 and Equation 5 capture the actual sensitivity, while current prior art metrics exhibit a false zero sensitivity, which is inaccurate.

[0090] Conventional Hessian-based techniques usually compute a zero sensitivity convergence point for a given loss curve scenario, which is an erroneous result. The proposed method produces remarkable sensitivity that exceeds that of Hessian-based methods.

[0091] Figure 5A A graphical representation of a loss curve for weights associated with a particular layer (i.e., layer-1) according to an embodiment of the present disclosure is shown (S500A). The proposed disclosure captures the curvature of the loss function affected by the weights of layer-1. It has several key components, including a small weight perturbation (Δw) (501a), a change in loss (ΔL) caused by the perturbation (502a), a gradient value (g) (503a), and an attraction loss curve for layer-1 (504a).

[0092] As disclosed in the Examples, Figure 5B The graphical representation (S500B) in is used to depict the loss curve associated with the weights of layer-2. The curvature of the loss function in layer-2 is shown, along with other key elements such as small weight perturbations (Δw) (501b), the change in loss due to the perturbation (ΔL) (502b), the gradient value (g) (503b), and the loss curve of layer-2 itself (505b). It is observed that the gradient value (g) in the loss curve of layer-2 is steeper than the gradient value in the loss curve (504a) of layer-1, indicating greater sensitivity to weight perturbations. More simply, for the same level of weight perturbation, the loss increase in layer-2 is more pronounced compared to layer-1. In the example, for the same amount of perturbation in the weights, the loss increase in layer L2 > the loss increase in layer L1. Therefore, layer-2 (505b) is more sensitive than layer-1 (504a).

[0093] Figure 5C A simplified diagram (S500C) showing the calculation of the average gradient norm is presented. As described in detail in the disclosed embodiments, this is achieved by perturbing the weights in N random directions. The graphical view includes four gradients (506a, 506b, 506c, 506d) and the loss curve (504) of layer-1 and the loss curve (505) of layer-2. Each gradient (506a, 506b, 506c, 506d) is used for a different purpose, where gradient 11, gradient 12, gradient 21 and gradient 22 are represented by 506a, 506b, 506c and 506d, respectively.

[0094] The average gradient norm of the perturbed weights in N random directions (here two directions are visualized) is calculated as follows:

[0095]

[0096] Average 2 > Average 1

[0097] The average gradient norm of the loss curve (505) for layer L2 is greater than the average gradient norm of the loss curve (504) for layer L1, so layer L2 (505) is more sensitive than layer L1 (504).

[0098] Figure 6 A graphical representation (600) of an estimate of the variation in loss across different levels of bit precision applicable to a given layer according to an embodiment of the present disclosure is shown. The view includes a small perturbation of the layer at b1 bit precision (Δwb1) (601), a small perturbation of the layer at b2 bit precision (Δwb2) (602), a variation in loss due to the weight of the layer at b1 bit precision (ΔLb1) (603), a variation in loss due to the weight of the layer at b2 bit precision (ΔLb2) (604), gradient values ​​(g) (605), and a loss curve (606). An average gradient norm (sensitivity) G is calculated by a layer-by-layer sensitivity calculator or controller, indicating the general rate of change of loss due to small perturbations.

[0099]

[0100] When calculating the average gradient norm G, the weights are perturbed by a random amount, that is, Δw is random.

[0101] This means ΔL=G*Δw.

[0102] For estimating the actual ΔL due to the actual weight quantization perturbation in different bit precision scenarios, it is as follows:

[0103] ΔL b =G*Δw b , where b represents the bit precision in the supported bit precision set (such as B = {2, 4, 8, 16, 32}) set on the hardware, which can be obtained from Figure 6 Refer to this content.

[0104] For a particular layer-i, the user of the electronic device (101) may set five estimated loss terms (from set B) for each bit precision allowed, as follows:

[0105] ΔL i 2 ,ΔL i 4 ,ΔL i 8 ,ΔL i 16 ,ΔL i 32 , and similarly for all layers i∈{1…L}.

[0106] For each layer, only one of the bit precisions can be selected from the allowed set B. Furthermore, picking a b from the set B results in:

[0107] Loss change = ΔL i b

[0108] Size change = b*P i , where P i represents the number of parameters in layer i.

[0109] The objective function aims to minimize the total loss represented by ΔL (which is the sum of all ΔLib, where b bits of precision are selected for layer i). On the other hand, the constraint requires that the total size of the model represented by S (which is the sum of b*Pi, where b bits of precision are selected for layer i) remains within the specified limit.

[0110] Figure 7 is a flowchart (S700) showing a method for mixed precision quantization of an AI model based on average gradient norm according to an embodiment of the present disclosure. Operations (S702-S710) are processed by a mixed precision quantization controller (105).

[0111] At operation 702 of the method, the sensitivity of the layer is calculated by using the average gradient norm (or average gradient value). Proceeding to operation 704, the method involves estimating the loss change of each layer by considering the sensitivity of each layer and the allowed bit precision setting. Subsequently, at operation 706, the method uses the calculated loss value change and net compression ratio of each layer to construct a constrained optimization problem model. The constrained optimization problem model attempts to optimize a given objective function and provides decision variables as a result, and considers some constraints. The constrained optimization problem model regards the target model size as a constraint and the total loss caused by model quantization as an objective function for minimization. The method maintains the bit precision of each layer as a decision variable after solving the optimization problem model. The constrained optimization problem model ultimately provides a decision variable as an output, which represents the bit precision to be assigned to each layer. Operation 708 involves assigning bit precision to the layer using the optimized solution from the constrained optimization problem model. Finally, at operation 710, post-training quantization is performed based on the assigned bit precision.

[0112] Figure 8 is a flowchart (S800) showing a method for mixed precision quantization of an AI model according to an embodiment of the present disclosure. Operations S802-S812 are processed by a mixed precision quantization controller.

[0113] At operation 802, the method requires a layer (301 across the AI ​​model) a-n ) multiple iterations of running the weight perturbation. At operation 804, the method involves evaluating each layer in the plurality of layers (301) based on the perturbation weights.a-n ) is the loss gradient of .

[0114] At operation 806 of the method, the loss gradient of each layer is analyzed to determine the loss gradient of each layer for each of the plurality of layers (301 a-n ) of the AI ​​model. In one embodiment, the output change reflects each layer (301) within the plurality of layers. a-n ) experienced losses.

[0115] Furthermore, at operation 808 of the method, each layer within the plurality of layers of the AI ​​model (301) is determined by measuring the output change of each layer. a-n ) is a sensitivity measure of the stratum. This measure is used as an indicator of the degree of responsiveness of each stratum within the stratum.

[0116] At operation 810, the method includes assigning a bit precision to each of a plurality of layers of the AI ​​model based on the determined sensitivity metric (301 a-n In an embodiment, each layer (301) of the plurality of layers is used. a-n ) sensitivity metric and net compression ratio to construct a constrained optimization problem model. Based on the constrained optimization problem model, bit precision is allocated to each of the multiple layers of the AI ​​model (301 a-n In another embodiment, bit precision is assigned to each of the plurality of layers of the AI ​​model by selecting at least one bit from a bit precision set based on a sensitivity metric (301). a-n ).

[0117] At operation 812, the method includes performing quantization of the AI ​​model using the bit precision assigned to each of the plurality of layers of the AI ​​model. In one embodiment, by using the bit precision assigned to each of the plurality of layers (301 a-n ) bit precision to perform post-training quantization of the AI ​​model, with each of the multiple layers (301 a-n ) is in the allocated bit precision to obtain the best mixed precision quantized AI model. The best AI model obtains each of the multiple layers of the AI ​​model in terms of at least one of power level, memory usage, computational efficiency level, and on-device learning on the electronic device (101) (301) a-n ) for optimal performance.

[0118] The proposed disclosure provides a cost-saving and time-saving solution for performing mixed-precision quantization of AI models while maximizing resource utilization. The adoption of an efficient first-order-based mechanism enables optimization of the run time between model development and deployment on devices without compromising quality. This approach surpasses other second-order-based techniques and provides a fast and efficient solution.

[0119] By utilizing only gradient information, the proposed method reduces computational requirements and accelerates the quantization process, thereby facilitating the operation of quantized models through on-device learning. This method allows for increased flexibility in running complex and deeper models at lower bit widths, creating opportunities for deploying additional models on-device. This, in turn, can lead to significant improvements in functional accuracy and performance.

[0120] The various actions, behaviors, blocks, operations, etc. in the flowcharts (S700 and S800) may be performed in the order presented, in a different order, or simultaneously. In addition, in some embodiments, some actions, behaviors, blocks, operations, etc. may be omitted, added, modified, skipped, etc. without departing from the scope of the present disclosure.

[0121] Any such software may be stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more computer programs (software modules) including computer-executable instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of the present disclosure.

[0122] Any such software may be stored in the form of volatile or non-volatile memory (such as, for example, a storage device like a read-only memory (ROM), whether erasable or rewritable), or in the form of a memory (such as, for example, a random access memory (RAM), a memory chip, a device or an integrated circuit), or stored on an optical or magnetically readable medium (such as, for example, a compact disk (CD), a digital versatile disk (DVD), a disk or a tape, etc.). It should be understood that the storage device and the storage medium are various embodiments of non-transitory machine-readable storage, which are suitable for storing one or more computer programs including instructions that, when executed, implement various embodiments of the present disclosure. Therefore, various embodiments provide a program and a non-transitory machine-readable storage storing such a program, the program including code for implementing an apparatus or method as claimed in any one of the claims of this specification.

[0123] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.

Claims

1. A method for performing mixed precision quantization on an artificial intelligence (AI) model by an electronic device, the method include: performing, by the electronic device, a predefined number of perturbations on the weights of each of the plurality of layers of the AI ​​model; determining, by the electronic device, a change in an output of each of the plurality of layers of the AI ​​model based on a perturbation of a weight of each of the plurality of layers; determining, by the electronic device, a sensitivity metric for each of the plurality of layers of the AI ​​model as a measure of a variation in an output of each layer; assigning, by the electronic device, a bit precision to each of the plurality of layers of the AI ​​model based on the determined sensitivity metric; as well as Mixed-precision quantization of the AI ​​model is performed, by the electronic device, using the bit precision assigned to each of the plurality of layers of the AI ​​model.

2. The method according to claim 1, in, The step of determining, by the electronic device, a change in the output of each of the plurality of layers of the AI ​​model comprises: determining, by the electronic device, a loss gradient for each of the plurality of layers based on a perturbation weight for each of the plurality of layers; and The electronic device determines a change in the output of each of the multiple layers of the AI ​​model based on the loss gradient of each of the multiple layers, wherein the change in the output indicates a loss with respect to each of the multiple layers.

3. The method according to claim 2, in, The sensitivity metric for each of the plurality of layers of the AI ​​model is determined based on the loss gradient.

4. The method according to claim 1, in, The step of assigning, by the electronic device, the bit precision to each of the plurality of layers of the AI ​​model based on the sensitivity metric comprises: constructing, by the electronic device, a constrained optimization problem model using the sensitivity metric and the net compression ratio for each of the plurality of layers; and The bit precision is assigned, by the electronic device, to each of the plurality of layers based on the constrained optimization problem model.

5. The method according to claim 1, in, The step of performing, by the electronic device, mixed precision quantization of the AI ​​model using the bit precision assigned to each of the plurality of layers of the AI ​​model includes: The electronic device performs post-training quantization of the AI ​​model by using the bit precision allocated to each of the multiple layers, so that each of the multiple layers can be at the allocated bit precision to obtain an optimal mixed-precision quantized AI model.

6. The method according to claim 5, in, The optimal AI model obtains optimal performance for each of the multiple layers of the AI ​​model in terms of at least one of power level, memory usage, computational efficiency level, and on-device learning on the electronic device.

7. The method according to claim 1, in, The bit precision is assigned to each layer of the plurality of layers of the AI ​​model by selecting at least one bit from a set of bit precisions based on the sensitivity metric.

8. An electronic device for mixed precision quantization of an artificial intelligence (AI) model, the electronic device include: Memory; one or more processors; as well as a mixed precision quantization controller communicatively coupled to the memory and the one or more processors, The memory stores one or more computer programs including computer executable instructions, and when the computer executable instructions are executed by the one or more processors, the electronic device performs the following operations: performing a predefined number of perturbations on the weights of each of the plurality of layers of the AI ​​model, determining a change in an output of each of the plurality of layers of the AI ​​model based on a perturbation of a weight of each of the plurality of layers, determining a sensitivity metric for each of the plurality of layers of the AI ​​model as a measure of a change in an output of each layer, assigning bit precision to each of the plurality of layers of the AI ​​model based on the determined sensitivity metric, and Quantization of the AI ​​model is performed using the bit precision assigned to each of the plurality of layers of the AI ​​model.

9. The electronic device as claimed in claim 8, in, To determine a change in the output of each of the plurality of layers of the AI ​​model, the one or more computer programs further include computer executable instructions to: determining a loss gradient for each of the plurality of layers based on a perturbation weight for each of the plurality of layers, and Determine a change in an output of each of the multiple layers of the AI ​​model based on a loss gradient of each of the multiple layers, wherein the change in the output indicates a loss with respect to each of the multiple layers.

10. The electronic device as claimed in claim 9, in, The sensitivity metric for each of the plurality of layers of the AI ​​model is determined based on the loss gradient.

11. The electronic device as claimed in claim 8, in, To assign the bit precision to each of the plurality of layers of the AI ​​model based on the sensitivity metric, the one or more computer programs further include computer executable instructions to: constructing a constrained optimization problem model using the sensitivity metric and the net compression ratio for each of the plurality of layers, and The bit precision is assigned to each of the plurality of layers based on the constrained optimization problem model.

12. The electronic device as claimed in claim 8, in, To perform quantization of the AI ​​model using the bit precision assigned to each of the plurality of layers of the AI ​​model, the one or more computer programs further include computer executable instructions to: By performing post-training quantization of the AI ​​model using the bit precision assigned to each of the multiple layers, each of the multiple layers can be at the assigned bit precision to obtain an optimal mixed-precision quantized AI model.

13. The electronic device as claimed in claim 12, in, The optimal AI model obtains optimal performance for each of the multiple layers of the AI ​​model in terms of at least one of power level, memory usage, computational efficiency level, and on-device learning on the electronic device.

14. The electronic device as claimed in claim 8, in, The bit precision is assigned to each layer of the plurality of layers of the AI ​​model by selecting at least one bit from a set of bit precisions based on the sensitivity metric.

15. The electronic device as claimed in claim 8, in, To assign the bit precision to each of the plurality of layers of the AI ​​model based on the sensitivity metric, the one or more computer programs further include computer executable instructions to: The final bit configuration of each layer is determined by considering the total loss from the loss estimator and the six constraints of the model.