Method and apparatus for multi-criteria decision
Binarization of AI models using a binarization-friendly NAS and conditional computation addresses the memory and computational challenges of floating-point precision, enabling efficient deployment on resource-constrained devices with improved accuracy and speed.
Patent Information
- Application Number
- PCT/KR2025/000306
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-07
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-10
AI Technical Summary
The high memory and computational demands of floating-point precision in AI models, particularly in resource-constrained environments, hinder efficient deployment and inference speed, which is critical for real-time applications.
A method and apparatus for multi-criteria decision making in binarized AI models, utilizing a binarization-friendly neural architecture search (NAS) to convert weights and activations to binary precision, with a two-stage training process and conditional computation to enhance accuracy and reduce computational overhead.
This approach significantly reduces memory usage and computational load, enabling efficient deployment of AI models on resource-constrained devices with minimal accuracy loss, enhancing inference speed and power efficiency.
Smart Images

Figure KR2025000306_10072025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR MULTI-CRITERIA DECISION
[0001] The disclosure is related to a method and an apparatus for multi-criteria decision making for the search and design of expert binary neural networks.
[0002] Artificial intelligence (AI) has seen substantial advancements in recent years, largely due to the development and utilization of sophisticated AI models. These models are typically constructed using floating-point precision (FP), a numerical representation that is used for handling a wide range of values efficiently. Floating-point precision is crucial for the complex calculations involved in AI mechanisms, as it allows for the representation of real numbers in a manner that balances the need for accuracy with the limitations of computer memory and processing power.
[0003] In the floating-point format, a number is expressed as a combination of two main components: the significand (also known as the mantissa) and the exponent. The significand represents the significant digits of the number, while the exponent indicates the scale or magnitude of the number by specifying the power of a predetermined base, commonly base 2 in binary systems. This format enables the representation of large or small numbers in a compact form, which is particularly beneficial for the intricate calculations required in AI model training and inference.
[0004] Despite the advantages of floating-point precision in enhancing the accuracy of AI models, there are several problems and limitations associated with its use. One significant issue is the substantial amount of memory required to store floating-point numbers. This increased memory requirement leads to a higher volume of calculations, which in turn demands more computational resources. The total floating-point operations per second (FLOPs) necessary for a single forward pass with FP can be extensive, contributing to the overall computational burden.
[0005] Furthermore, the high computational demands of floating-point precision result in significant power consumption. AI models often require considerable power for these computations, which can be a limiting factor in various applications. High power consumption affects operational costs and poses challenges for deploying AI models in resource-constrained environments, such as edge devices or mobile platforms.
[0006] Another issue is the impact on inference speed. The extensive calculations and high memory usage associated with floating-point precision can hinder the speed at which AI models perform inference. This slowdown can be detrimental in real-time applications where rapid decision-making is crucial, such as autonomous driving, real-time language translation, or interactive AI systems.
[0007] Hence, it is desirable to address the aforementioned problems and disadvantages or at least provide a useful alternative.
[0008] In an embodiment, provided is a method for inference processing in a binarized artificial intelligence (AI) model in an electronic device. The method includes inputting by the electronic device an input content to the binarized AI model. Further, the method includes determining by the electronic device a pre-defined cluster to which the input content belongs using the binarized AI model. Further, the method includes activating by the electronic device a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster. Further, the method includes generating by the electronic device a final output by processing the input content received using the activated binary variant.
[0009] In an embodiment, provided is an electronic device for inference processing of a binarized AI model. The electronic device includes a processor, a memory, and an inference processing controller communicatively coupled to the memory and the processor. The inference processing controller receives an input content from a full-precision AI model. Further, the inference processing controller determines the pre-defined cluster to which the input content belongs using the binarized AI model. Further, the inference processing controller activates a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster. Further, the inference processing controller generates a final output by processing the input content received using the activated binary variant.
[0010] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings.
[0011] These and other features, aspects, and advantages of the embodiment are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiment herein will be better understood from the following description with reference to the drawings, in which:
[0012] Fig. 1 is a schematic diagram illustrating an image captured in a night mode enhanced using Vision-based AI mode according to an embodiment of the disclosure.
[0013] Fig. 2 is a schematic diagram illustrating an exemplary large language mode (LLM)-based Smart reply use case on a smartwatch device according to an embodiment of the disclosure.
[0014] Fig. 3 is a block diagram illustrating an optimization of models using Neural Network Architecture Search (NAS) according to an embodiment of the disclosure.
[0015] Fig. 4 is a block diagram of an electronic device for inference processing in a binarized AI model according to an embodiment of the disclosure.
[0016] Fig. 5 is a flow diagram illustrating a method for inference processing in a binarized AI model according to an embodiment of the disclosure.
[0017] Fig. 6 is a block diagram illustrating a schematic of an inference processing controller according to an embodiment of the disclosure.
[0018] Fig. 7 is a block diagram illustrating agile neural architecture search for binarization-friendly high-capacity AI networks according to an embodiment of the disclosure.
[0019] Fig. 8A is a schematic diagram illustrating conditional computation with expert AI blocks according to an embodiment of the disclosure.
[0020] Fig. 8B is a schematic diagram illustrating conditional computation with expert AI blocks according to an embodiment of the disclosure.
[0021] Fig. 9A is a block diagram illustrating a schematic of an exploded view of an expert binarization block of the inference processing controller of Fig. 4 according to an embodiment of the disclosure.
[0022] Fig. 9B is a block diagram illustrating a first-stage training solution for binary weights and activation according to an embodiment of the disclosure.
[0023] Fig. 9C is a block diagram illustrating a second-stage training solution for binary weights and activation according to an embodiment of the disclosure.
[0024] Fig. 10 is a block diagram illustrating the training of a Binary Neural Network according to an embodiment of the disclosure.
[0025] Fig. 11 is a schematic diagram illustrating a use case of binarization of large language models (LLMs) according to an embodiment of the disclosure.
[0026] Fig. 12 is a schematic diagram illustrating a use case of binarization of vision-based stable diffusion models according to an embodiment of the disclosure.
[0027] Terms used in the disclosure will be briefly described, and an embodiment of the disclosure will be described in detail.
[0028] Although general terms being currently widely used were selected as terminology used in the disclosure while considering the functions in the disclosure, they may vary according to intentions of one of ordinary skill in the art, judicial precedents, the advent of new technologies, and the like Also, terms arbitrarily selected by the applicant may also be used in a specific case. In this case, their meanings will be described in detail in the corresponding embodiments of the disclosure. Hence, the terms used in the disclosure must be defined based on the meanings of the terms and the entire content of the disclosure, not by simply stating the terms themselves.
[0029] In the disclosure, the expression "at least one of a, b or c" indicates "a", "b", "c", "a and b", "a and c", "b and c", "all of a, b, and c", or variations thereof.
[0030] Throughout the disclosure, it will be understood that when a certain part "includes" a certain component, the part does not exclude another component but can further include another component, unless the context clearly dictates otherwise. As used herein, the terms "part", "portion", "module", or the like refers to a unit that can perform at least one function or operation, and may be implemented as hardware, software, or a combination of hardware and software.
[0031] Hereinafter, the embodiment of the disclosure will be described in detail with reference to the accompanying drawings such that one of ordinary skill in the technical field to which the disclosure belongs may easily embody the disclosure. However, an embodiment of the disclosure can be implemented in various different forms, and is not limited to the embodiments described herein. Also, in the drawings, portions irrelevant to the description are not shown in order to definitely describe an embodiment of the disclosure, and throughout the disclosure, similar components are assigned similar reference numerals.
[0032] The object of the disclosure is to provide methods and devices for inference processing in a binarized AI model of an electronic device.
[0033] The object of the disclosure is to provide a method that provides a binarization-friendly multi-objective NAS to redesign the AI blocks in full-precision models with quantization-friendly alternatives.
[0034] The object of the disclosure is to provide a Binary Neural Network (BNN) with an Ensemble of Experts where expert specializes in a portion of data, thus ensuring the accuracy.
[0035] The object of the disclosure is to provide a cascaded training mechanism with Entropy Regularization for ensuring dynamic input content-based activation of experts during runtime.
[0036] The object of the disclosure is to provide a lightweight gating mechanism based on a SoftMax function to ensure the dynamic and automated switching between the experts during the inference based on the input content.
[0037] The object of the disclosure is to provide a quantization-aware training mechanism with expert specialization on user's data allowing for on-device binarization ensuring personalization without breach of privacy.
[0038] The object of the disclosure is to provide a method for deploying extremely memory-intensive AI models (e.g., GenAI) on extremely memory-constrained devices (e.g., wearables).
[0039] The embodiment and the various features and details are explained more fully with reference to the non-limiting embodiment that is illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiment herein. Also, the embodiment described herein is not necessarily mutually exclusive, as the embodiment disclosed herein can be combined with another embodiment to form a new embodiment. The term "or" as used herein, refers to a non-exclusive or, unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiment herein. Accordingly, the examples are not be construed as limiting the scope of the embodiment herein.
[0040] The embodiment is described and illustrated in terms of blocks that carry out a described function or functions. These blocks, which referred to herein as managers, units, modules, hardware components or the like, may be physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, and the like, and optionally be driven by firmware and software. The circuits, for example, may be embodied in a plurality of semiconductor chips, or on substrate supports such as printed circuit boards, and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., a plurality of programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiment may be physically separated into two or more interacting and discrete blocks without departing from the scope of the invention. Likewise, the blocks of the embodiment may be physically combined into more complex blocks without departing from the scope of the invention.
[0041] The accompanying drawings are used to help easily understand various technical features and it is understood that the embodiment presented herein is not limited by the accompanying drawings. Although the terms first, second, etc. used herein to describe various elements, these elements are not be limited by these terms. These terms are generally used to distinguish one element from another.
[0042] Binary variant is the AI model whose weights and activations in full precision have been quantized into a binary precision. Quantization is a process that reduces the precision of the weights and activations in a neural network to lower memory usage and computational demand, making it more suitable for deployment on resource-constrained devices. In neural networks, activation refers to the process of applying an activation function to the output of a neuron or layer.
[0043] During the quantization aware training, the data is grouped into clusters and a binary variant is trained to detect patterns in the data. This cluster of data is called predefined cluster during the inference.
[0044] Binarization friendly layers are the layers of the neural network that reduce the inference time and complexity after binarization.
[0045] The AI blocks which consume memory power or have high latency may result in bad on-device Key Performance Indicators (KPI).
[0046] Sub-Optimal AI block are the AI blocks which consume memory power or have high latency, i.e. the AI blocks that result in bad on-device Key Performance Indicators (KPI).
[0047] Expert blocks are interchangeably used as binary variants.
[0048] Input contents to neural network can include but not limited to text, image, video, audio, and numeric.
[0049] Fig. 1 is a schematic diagram illustrating an image captured in night mode enhanced using Vision-based AI mode according to an embodiment of the disclosure. Generally, the weights of the AI models are stored in the ROM. Further, the stored weights are accessed by RAM for processing input images or text. For example, as shown in FIG. 1, the input image (101) is inputted to the AI model (103). The AI model may be, but not limited to, a vision-based AI model. A set of weights (107) of the AI model (103) is stored in memory. The memory may be, but not limited to, the Read Only Memory (ROM) (109). For example, the memory may be a readable and writable memory such as a flash rom. Further, the stored set of weights are accessed by the RAM (111) of the mobile device (113) for processing the input image (101). However, the AI model sizes have been increasing exponentially over the past decades and are expected to grow by 1000X in decades. Thus, the storage of the weights of the AI model cannot be stored in the ROM due to the requirement of high storage capacity. Hence, it is necessary to reduce the memory footprint of growing AI model to make it fit into the limited memory on embedded devices.
[0050] Fig. 2 is a schematic diagram illustrating an example LLM-based Smart Reply use case on a smartwatch device according to an embodiment of the disclosure. A Large Language Model (LLM) is a type of computational model designed for natural processing tasks such as language generation and the like. The blocks (2011-2014) indicate the LLM-based smart reply use case on a smartwatch with support from an AI model hosted on smartphones. The LLMs are commercialized on embedded devices or smart devices in the floating-point 32-bit precision, having a memory of approximately 3GB. Thus, the floating-point 32 precision is converted to binary precision, which will result in 32X savings in memory and 60X improvement in latency, allowing deployment on further constrained devices such as smartwatches.
[0051] The existing methodology can be applied for vision-based generative AI applications. The deployment on extremely resource-constrained devices such as smartwatches is possible by binarization. However, the binarization results in a massive accuracy drop. Also, the binarization of generative AI prevents it from being a one-for-all foundation model that serves multiple use cases. Further, the binarization is a training-time-consuming method.
[0052] Fig. 3 is a block diagram illustrating an optimization of models using Neural Network Architecture Search (NAS) according to an embodiment of the disclosure. The NAS ensures the construction of architecture with reduced complexity. Further, the NAS optimizes the AI model for on-device parameters. However, the AI models will still use floating point precision values, thus resulting in no saving in memory. Also, the NAS has a very high computational demand and the NAS is cost-intensive.
[0053] The NAS facilitates the development of architectures with lower complexity. The NAS is capable of optimizing AI models for parameters suited for on-device applications. These models remain in floating point precision, which does not contribute to memory savings. Further, the NAS is highly demanding in terms of computational resources and costs. To enhance performance, weights or layers are removed in pre-trained AI models that have minimal impact on accuracy. This pruning process optimizes a pre-trained AI model, contrasting with NAS by reducing floating point operations per second (FLOPs). Nevertheless, the pruned models still utilize floating point precision, resulting in reduced memory savings. Not all pruning techniques lead to improved latency. Further, converting floating point weights and activations to integer representation can yield benefits in both memory usage and latency. However, the memory savings achieved are often insufficient to support generative AI models, and accuracy may be compromised due to the reduction of precision.
[0054] An embodiment of the disclosure addresses two key challenges. The first challenge is enhancing the accuracy of a binary neural network during the training phase. The second challenge is ensuring that the optimization achieved is effectively utilized during the inference phase. Enhancing may begin by analyzing the base AI model in full-precision, improving its learning capacity to minimize accuracy loss when quantized into binary format. This is accomplished through a binarization-aware NAS. This flexible solution allows for a dual focus on both accuracy and complexity, transforming the original AI model into one that is more compatible with quantization. Further, the optimized NAS model serves as a teacher model in the subsequent stage, guiding the binarization process. Instead of creating a single binary variant for the full-precision AI model, this intelligent method clusters the data into distinct regions, assigning a unique binary variant to each cluster. By breaking down the learning problem into smaller manageable tasks, the accuracy degradation may be significantly reduced. Furthermore, the high-capacity network derived from NAS acts as a teacher from which knowledge is transferred to the student binary variants.
[0055] The AI model has been reengineered for optimal precision, specifically designed for binary quantization. Implementing multi-objective NAS on the existing full-precision AI model enhances the network's capacity, allowing the binary representation to maintain its accuracy. This solution minimizes accuracy loss and enhances the convergence of the BNN by functioning as an effective teacher model in knowledge distillation-based quantization-aware training. Given that the NAS process is training-free, it operates at remarkable speed compared to previous methods, resulting in agile solutions that are both highly accurate and efficient in terms of hardware usage. Each AI block is quantized into various versions known as experts, each tailored to specialize in specific data regions. The training of the BNN employs a two-stage method where weights are quantized first, followed by activations, facilitating a cascading improvement in accuracy. Conditional computation is implemented to dynamically activate a single quantized expert AI block during runtime.
[0056] Furthermore, the embodiment leverages advanced techniques to ensure that the binary neural network (BNN) maintains high performance across various tasks. By incorporating a binarization-aware NAS, the architecture is specifically optimized to handle the challenges associated with binary quantization. This includes the ability to adaptively manage different data regions, ensuring that each binary variant is finely tuned for its specific subset of data. This granularity in optimization allows the BNN to perform more effectively, reducing the overall accuracy loss typically associated with binary quantization.
[0057] Furthermore, the implementation of conditional computation introduces a significant improvement in efficiency. By dynamically activating the quantized expert AI block during runtime, the system avoids the computational overhead associated with processing multiple blocks simultaneously. This conserves memory and enhances the speed and responsiveness of the AI model during inference. The two-stage training method, which focuses first on weights and then on activations, ensures that each stage of the quantization process is meticulously optimized, leading to a more robust and accurate BNN. This holistic solution to model optimization, from NAS to conditional computation, represents a significant advancement in the field of AI model efficiency and performance.
[0058] Fig. 4 is a block diagram of the electronic device (401) for inference processing in a binarized AI model according to an embodiment of the disclosure. According to an embodiment of the disclosure, the electronic device (401) may include, but are not limited to, Consumer Electronics (such as Mobile Phones and Smartphones), Tablets, Wearable Devices, Computing Devices (such as Laptops, Notebooks, Desktops, Workstations, etc.), IoT Devices, Automotive Systems (such as connected cars, Autonomous Vehicles, Vehicle-to-Everything (V2X) communication devices, etc.), Enterprise Devices such as robotics, Specialized Equipment (such as Medical Devices, Public Safety Devices, etc.), Media Devices (such as Gaming Consoles, Streaming Devices, etc.).
[0059] In an embodiment, the electronic device (401) may include a processor (403), memory (405), an I / O interface (407), and an inference processing controller (409) communicatively coupled to the processor (403) and the memory (405). Although components of the electronic device 401 are described in a singular form, the components of the electronic device 401 may include multiple components. Each component is explained in further detail below.
[0060] The processor (403) may include various processing circuitry and / or multiple processors. For example, as used herein, including the claims, the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "processor", "at least one processor" and "one or more processors"are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited / disclosed functions, e.g., in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions. The processor (403) communicates with the memory (405), the I / O interface (407), and the inference processing controller (409). The processor (403) may execute instructions stored in the memory (405) and perform various processes. The processor (403) may be a general-purpose processor such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Artificial Intelligence (AI) dedicated processor such as a neural processing unit (NPU).
[0061] The memory (405) may include storage locations to be addressable through the processor (403). The memory (405) is not limited to a volatile memory and / or a non-volatile memory. Further, the memory (405) may include a plurality of computer-readable storage media. The memory (405) may include non-volatile storage elements. For example, non-volatile storage elements may include magnetic hard disks, optical disks, floppy disks, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories.
[0062] The I / O interface (407) may transmit the information between the memory (405) and external peripheral devices. The peripheral devices are the input-output devices associated with the inference processing controller (409). Further, the inference processing controller (409) communicates with the I / O interface (407) and the memory (405). The inference processing controller (409) may be communicatively coupled to the memory (405) and the processor (403). The inference processing controller (409) is an innovative hardware that is realized through the physical implementation of both analog and digital circuits, including logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive and active electronic components, as well as optical components.
[0063] In an embodiment, although the inference processing controller (409) is depicted as another processor different from the processor (403), the inference processing controller (409) may be combined into the processor (403). In an embodiment, the inference processing controller (409) may be a processor separate from the processor (403). The inference processing controller (409) may include various controlling circuitry and / or multiple controllers. For example, as used herein, including the claims, the term controller' may include various controlling circuitry, including at least one controller, wherein one or more of at least one controller, individually and / or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "controller", "at least one controller" and "one or more controllers"are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one controller performs some of recited functions and another controllers(s) performs other of recited functions, and also situations in which a single controller may perform all recited functions. Additionally, the at least one controller may include a combination of controllers performing various of the recited / disclosed functions, e.g., in a distributed manner. At least one controller may execute program instructions to achieve or perform various functions. The inference processing controller (409) may input the input content to the binarized AI model. Further, the inference processing controller (409) may determine or identify the pre-defined cluster to which the input content belongs using the binarized AI model. The input content can include, but is not limited to, text, image, and numerical representation. Further, the inference processing controller (409) activates the binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster. Each of the binary variants of the plurality of binary variants is personalized based on the pre-defined cluster of the input content. The pre-defined cluster represents the user data. Further, the inference processing controller (409) may generate a final output by processing the input content received using the activated binary variant.
[0064] In an embodiment, the binarized AI model is generated by enhancing the capacity of a full-precision AI model by adding at least one channel and at least one binarization-friendly layer(s) to the full-precision AI model. This enhancement allows the model to maintain high performance while reducing computational complexity and memory usage, making it more efficient for deployment on resource-constrained devices. The binarization process involves converting the trained weights and activations of the AI model in full precision to binary values (binary precision), which simplifies the arithmetic operations required during inference and leads to faster processing times. In the context of AI models, particularly in deep learning and a neural networks, adding a channel generally refers to increasing the number of feature maps in a particular layer of the AI model. Channels in a neural network layer often represent different types of information or features extracted from the input data. For example, in convolutional layers, an input image typically has three channels corresponding to the RGB color mode (Red, Green, Blue) for color images. Each channel captures different aspects of the image. Each channel in a layer is associated with a feature map. When a channels is added, another feature map is effectively added. This added feature map may highlight specific features that the network has learned to identify. For instance, if a layer has 16 filters (channels), the output will have 16 feature maps representing the learned features from the input data.
[0065] In an embodiment, to add the one or more channels and binarization-friendly layers to the full-precision AI model, the inference processing controller (409) profiles the binarized AI model to identify sub-optimal AI blocks that inhibit efficient binarization. The profiling may include the process of analyzing the performance of an AI model, typically referring to understanding the computational characteristics and resource usage of the AI model. Further, the inference processing controller (409) determines the quantization-aware search space to prepare and train quantization-friendly alternatives for the identified sub-optimal AI blocks. In the context of AI models, particularly in neural architecture search (NAS) and model compression techniques, quantization-aware search space refers to a framework or design space that accounts for the effects of quantization during the search for optimal neural network architectures. This involves exploring various architectural modifications and training techniques that can improve the binarization process. Further, the inference processing controller (409) adds one or more channels and binarization-friendly layers to the full-precision AI model based on the quantization-friendly alternatives for the identified sub-optimal AI blocks. These modifications ensure that the binarized AI model retains its accuracy while benefiting from the reduced computational overhead.
[0066] In an embodiment, each binary variant of the plurality of binary variants is relevant to the pre-defined clusters associated with the input content that is processed by the binarized AI model. This relevance is achieved by tailoring each binary variant to handle specific types of input content, ensuring that the model can provide accurate and efficient predictions for a wide range of content. In an embodiment, during the training of the binary variants, the inference processing controller (409) initializes the plurality of binary variants to the full-precision AI model. Further, the inference processing controller (409) assigns random weights to each binary variant of the plurality of binary variants. This initialization process ensures that each binary variant starts with a unique set of parameters, allowing them to specialize in different aspects of the input content.
[0067] Further, the inference processing controller (409) determines the probability for the random weights assigned to each binary variant of the plurality of binary variants. While assigning, a SoftMax-based activation function may be used. This function normalizes the weights, ensuring that the weights may sum to one and can be interpreted as probabilities. Further, the inference processing controller (409) determines the weighted sum of each binary variant of the plurality of binary variants to obtain the first output. The weighted sum is determined as a sum of the product of an output of each of the binary variants and corresponding probabilities. This solution allows the model to combine the strengths of multiple binary variants, leading to more accurate predictions.
[0068] In an embodiment, the inference processing controller (409) determines the second output using the full-precision AI model. Further, the inference processing controller (409) determines the knowledge distillation loss by comparing the first output and the second output. Knowledge distillation is a technique where a smaller, simpler model (the student) is trained to replicate the behavior of a larger, more complex model (the teacher). Further, the inference processing controller (409) determines the training loss by comparing the first output and a ground truth. The training loss may refer to a loss function used during the training of the neural network. The same loss function may be used for the knowledge distillation loss and the entropy regularization loss.
[0069] This loss leads to measure the difference between the model's predictions and the actual values, guiding the training process. Further, the inference processing controller (409) determines the entropy regularization loss. The entropy regularization loss is one of the regularization techniques used in the training process of AI models, particularly deep learning models. This technique helps to control the uncertainty of the model's predictions, preventing overfitting and improving generalization performance. The determined entropy regularization loss leads to ensure that during inference, only one binary variant (or single binary variant) among the plurality of binary variants is activated, promoting model sparsity and efficiency. Further, the inference processing controller (409) determines a total loss. The total loss is a sum of the knowledge distillation loss, the training loss, and the entropy regularization loss. Further, the inference processing controller (409) determines the final output by minimizing the total loss, where the final output is a collection of tuned weights of each binary variant of the plurality of binary variants along with the weights.
[0070] In an embodiment, each of binary variants is trained on a specific region of the input content through the entropy regularization loss. This targeted training solution ensures that each binary variant becomes an expert in handling particular types of input, improving the overall performance of the binarized AI model. In an embodiment, the inference processing controller (409) generates an output of the full-precision AI model by selecting at least one expert block among a plurality of expert blocks having a probability of one while the probability of the other expert blocks of the plurality of expert blocks are zero based on the input content. This selective activation of expert blocks ensures that the model can efficiently process the input content by leveraging the most relevant parts of the network, further enhancing its accuracy and efficiency.
[0071] Fig 5 is a flow diagram illustrating the method for inference processing in a binarized AI model of an electronic device (401) according to an embodiment of the disclosure. This method streamlines the processing of input content by leveraging the efficiency of binarized neural networks, which utilize binary weights and activations to minimize computational complexity and enhance processing speed.
[0072] At block 501, the method includes inputting the input content to the binarized AI model. This input content can be any form of data that the AI model is designed to process, such as images, text, or sensor readings. The binarized AI model is pre-trained to handle such inputs by converting them into binary format, which simplifies the subsequent computational steps. This conversion allows the model to operate with reduced memory usage and faster processing times, making it ideal for deployment in resource-constrained environments like mobile devices or embedded systems.
[0073] At block 503, the method includes determining the pre-defined cluster to which the input content belongs using the binarized AI model. The model employs clustering techniques to categorize the input content into one of several pre-defined clusters. These clusters are created during the training phase of the model and represent different categories or types of input content that the model can recognize. By determining the appropriate cluster for the input content, the model can tailor its processing approach to the specific characteristics of the data, thereby improving accuracy and efficiency.
[0074] At block 505, the method includes activating the binary variant from the plurality of binary variants corresponding to the determined pre-defined cluster. Each cluster has an associated binary variant, which is a specialized version of the binarized AI model optimized for processing the type of data in that cluster. By activating the relevant binary variant, the method ensures that the input content is processed using the suitable model configuration, further enhancing the performance and accuracy of the inference process.
[0075] At block 507, the method includes generating the final output by processing the input content using the activated binary variant. The activated binary variant processes the input content to produce the final output, which could be a classification result, a prediction, or any other form of processed data depending on the application of the AI model. This final output is then used by the electronic device (401) to perform its intended function, such as recognizing objects in an image, translating text, or making decisions based on sensor data. The use of binarized AI models in this method accelerates the processing time and reduces the power consumption, making it highly suitable for real-time applications in various electronic device (401).
[0076] Fig 6 is a block diagram illustrating the schematic view of the inference processing controller (409) according to an embodiment of the disclosure. The inference processing controller (409) includes a base AI model (601), a binarization-friendly NAS model (603), a high-capacity AI block (605), an expert binarization model (609), a lightweight gating mechanism (607), and a classification model (619).
[0077] The base AI model (601) serves as the foundation of the inference processing controller (409). This model ensures that the inference processing controller (409) can handle basic tasks effectively, setting the stage for advanced functionalities. The robustness of the base AI model (601) underpins inference process, ensuring that the system remains stable and reliable even when dealing with diverse and unpredictable data inputs.
[0078] Facilitating NAS, the binarization-friendly NAS model (603) is compatible with binarization techniques. By optimizing the architecture of neural networks for binary operations, this model enhances the efficiency of the inference process, allowing for faster computations and reduced resource consumption. By reducing the complexity of the neural network operations, this model ensures that high-performance AI can be deployed in a wide range of environments, making advanced AI capabilities more accessible and practical.
[0079] Equipped with advanced processing capabilities, the high-capacity AI model (605) can handle larger datasets. Its high capacity ensures that the inference processing controller (409) can perform analysis and make informed decisions based on the data it processes.
[0080] The expert binarization model (609) refines the binarization process. Leveraging expert knowledge and techniques, this model optimizes the conversion of neural network weights and activations into different binary variants. By doing so, it enhances the overall performance of the AI model, ensuring that the inference processing remains accurate and efficient even when operating in a binary mode. The expert binarization model (609) maintains the integrity and precision of the AI model, particularly in applications.
[0081] During the training phase, the AI model is quantized into a binary format, and different binary experts corresponding to specific regions of the data are obtained. In the inference stage, the lightweight gating mechanism model (607) dynamically switches between the binary variants of the full-precision AI model. This dynamic switching capability optimizes the inference process, as it allows the device to be adapted to varying data inputs and processing requirements in real-time. The classification model (619) is responsible for interpreting the processed data and making classifications of the processed data based on the outputs generated by the preceding components. This model ensures that the final output is accurate and relevant to the input data, providing reliable results for end-users.
[0082] In an example, it is assumed that the input content (611) includes an image showing a few dogs. During the training phase, the expert binarization block (609) develops three distinct expert blocks for the classification of objects included in the image. The objects may be classified into pets, MNIST (Modified National Institute of Standards and Technology database) digits, and cars. In the inference stage, the lightweight gating mechanism (607) activates the appropriate binary variant based on the input content (611) provided to the network, which in this case is pets. This process ensures that the final output is produced accurately, thereby enhancing overall performance. The final output determined by the classification model (619) is a pet classification. This example illustrates the practical application of the inference processing controller (409), demonstrating how it can effectively manage and classify diverse types of data inputs, ensuring high accuracy and efficiency in real-world scenarios.
[0083] Fig. 7 is a block diagram illustrating agile neural architecture search for binarization-friendly high-capacity AI networks according to an embodiment of the disclosure. The binarization-friendly NAS model (603) provides the NAS for the high-capacity alternative of the given AI model. In operation S1, a pre-trained AI model, such as an LLM, is given as input to an automated network surgery block. The pre-trained AI model provided as the input is in a format of floating point precision.
[0084] In operation S2, the automated network surgery block performs profiling of the pre-trained input AI model by determining the sub-optimal AI blocks that prevent efficient binarization. Quantization-friendly alternatives are identified using on-device profiling of the pre-trained input AI model.
[0085] In operation S3, the quantization-aware search space design block trains the determined quantization-friendly alternatives using feature-wise knowledge distillation. Further, the quantization-friendly alternatives are pruned.
[0086] In operation S4, an agile search strategy block performs end-to-end training to redesign the AI model with high capacity and make it quantization-friendly. The agile search strategy also solves the combinatorial optimization problem quickly with around 50X lower infrastructure cost. The resultant obtained by the agile search strategy block is the binarization-friendly neural network in full precision, which serves as the ideal starting point for quantization.
[0087] Fig. 8A is a schematic diagram illustrating the conditional computation with expert AI blocks according to an embodiment of the disclosure. Fig. 8B is a schematic diagram illustrating the conditional computation with expert AI blocks according to an embodiment of the disclosure. As shown in Fig. 8A and Fig. 8B, the binarization procedure represents the full-precision AI network into multiple variants. Each AI model (601) of the pre-trained AI model in the network has N number of binary variants, such as binary variant 1 (613), binary variant 2 (615), and binary variant N (617), among others. The binary variants (613, 615, 617) is assigned random weights α1 (801), α2 (803), and αN (805). These random weights are derived from data points. This derivation may be performed by a predetermined algorithm automatically. For N binary variants, the random weights α1 to αn (801, 803, 805) are assigned. These random weights are then processed with a SoftMax activation (701) to generate probabilities P1to Pn(807, 809, 811) each which may be corresponding to each of the random weights. Upon processing, only one of the probabilities will have the value 1, while the rest will have a value of zero. This ensures the dynamic selection of a single binary variant for the given data point. For better understanding of the SoftMax activation, a SoftMax function should be understood first. The SoftMax function is a mathematical function commonly used in AI models, particularly in the context of multi-class classification problems. The SoftMax function converts a vector of raw scores (logits) from the model into probabilities that sum to 1. This makes it particularly useful in the final layer of neural networks for classification tasks. The SoftMax activation refers to the use of the SoftMax function as an activation function in the output layer of a neural network, particularly for multi-class classification tasks. The SoftMax function transforms the raw output logits (i.e., the unbounded real-valued scores produced by the preceding layer) into a probability distribution over the classes.
[0088] The process of deriving the random weights α1 to αn (801, 803, 805) is illustrated in Fig. 9A. For an input content (611), which can be an image, audio, video, or the like, the content is represented in feature maps. These feature maps are inputted into the convolutional layer (901), which includes N kernels or CNN layers through which the feature maps are passed. The convolutional layer (901) generates activation maps containing n channels. A global average pooling is then performed on the n channels. The score obtained as a result of the global average pooling is the random weight α1 to αn (801, 803, 805). Thus, the score obtained from the global average pooling for the n channels is referred to as the random weights α1 to αn (801, 803, 805).
[0089] As shown in Fig. 8B, the probabilities P1to Pn(807, 809, 811) is multiplied by its corresponding binary variant, resulting in the weighted binary variants. The weighted binary variants is then added each other, and the result of the addition represents the substitute for the input AI model (601). In the proposed solution, the AI model (601) is represented as the weighted sum of multiple binary variants α1 to αn (801, 803, 805). Further, only one of the binary variants will be active, reducing the overhead during runtime.
[0090] Fig. 9B is a block diagram illustrating the first-stage training solution for binary weights and activation according to an embodiment of the disclosure. Fig. 9C is a block diagram illustrating the second-stage training solution for generating binary weights and binary activation (activation maps) according to an embodiment of the disclosure. The first stage is associated with binary weight training. During the training procedure at the first stage, only the weights are binarized. At the first stage, the input content (611) is inputted to the binarization-aware NAS-AI model (907), which operates in floating point precision. Further, the input content (611) is inputted to the binary neural network (909) with experts for each AI model (601) where the weights of the binary neural network (909) are binarized but the activation (activation maps) remain in full precision. Thus, the activation maps are not disturbed at the first stage. The binarization-aware NAS-AI model (907) generates first output content (917), and the binary neural network model (909) generates second output content (919). A knowledge distillation loss is determined between the first output content (917) and the second output content (919) generated from the binarization-aware NAS-AI model (907) and the binary neural network (909) respectively. In an example, the knowledge distillation loss may be obtained with a mean squared error (MSE) between the first output content and the second output content. The mean squared error may be derived from the Equation (1) below.
[0091] ... Equation (1), where yiis the ithpixel value in one map and is the ithpixel value in another map.
[0092] The mean squared error (MSE) provides the average difference between two feature maps. In an embodiment, the two maps are 1) the predicted output and the ground truth for a training loss, and 2) the predicted output and the output in the full precision AI model for a knowledge distillation loss.
[0093] As shown above, the training loss is determined based on the MSE between the second output content (919) generated by the binary neural network (909) and the ground truth (921). The ground truth (921) may be a correct answer or a correct value to a specific problem. For instance, the ground truth (921) may include a location of an object and class information in an object recognition for image processing. Or, the ground truth (921) may be the accurate translation of the original sentence in machine translation.
[0094] An entropy regularization loss is also determined using the Equation (2) provided below:
[0095] ... Equation (2)
[0096] After training, only a single probability remains active while other probabilities are zero for the AI model (601). The knowledge distillation (KD) loss, the training loss, and the entropy regularization (ER) loss are summed to obtain a quantization of the weights.
[0097] The second stage is associated with binary activation map training. During the second stage, the proposed solution aims to quantize the activation maps to reduce accuracy loss, as shown in Fig. 9C. The input contents are fed into the binarization-aware NAS-AI model (907) which operates in floating point precision and the binary neural network model (909) during the quantization of activation. Based on the generated first output content (917) and the second output content (919), the KD loss, the training loss, and the ER loss are determined and minimized to obtain the quantized activation and to enable cascaded accuracy gain. The accuracy drop occurs due to binarizing the weights and the activation maps. At the first stage, the accuracy drop is minimized by considering the impact of the weights alone. At the second stage, the impact of activation maps is also considered. Thus, the accuracy gain is maximized in two cascaded stages.
[0098] Fig. 10 is a block diagram illustrating the training of a Binary Neural Network according to an embodiment of the disclosure. In particular, Fig. 10 shows a flow of the quantization aware training process which results in the conversion of the weights and activations in full precision into binary precision. This achieved in two stages that includes the first training stage (1011) and the second training stage (1019). In the first training stage (1011), the activation maps remain in full precision and only the weights are quantized into binary precision. In the second training stage (1019), the activation maps are quantized into binary precision along with further fine-tuning of the binarized weights to minimize the accuracy drop. In both of the first training stage (1011) and the second training stage (1019) , the weights are the parameters that are tuned by the optimizer to minimize the accuracy drop due to quantization of weights and activation maps. Therefore, dividing the problem into two stages helps the optimizer that tunes the weights, thereby ensuring that the problem of accuracy drop is solved efficiently.
[0099] The dataset (1001) is provided as an input to a teacher model (1003). This teacher model (1003) is a binarization-friendly, hardware-aware NAS optimization AI model. At block 1005, an expert is selected by the inference processing controller (409) in the network architecture of the AI model (601). During the inference, only one of the probabilities will be 1 and the rest will be 0. The output from the experts is measured as the weighted sum of the outputs from each expert where the weights are probabilities. Thus, the expert that has the probability as 1 will be active and selected while the effect of the others will be nullified. Before the quantization aware training starts, the AI model in full precision is modified in an optimal fashion using binarization friendly neural architecture search. The NAS problem as previously discussed is a multi-objective optimization problem, thus resulting in a list of Pareto optimal solutions which are obtained by applying non-dominated sorting at block 1007 over the Pareto solutions. Among the non-dominated solutions, a single architecture whose accuracy on the dataset (1001) is the highest may be selected at block 1005. This model is now quantized using the quantization aware training approach. It also acts as a teacher model (1003) from which the knowledge is distilled into the quantized model. At block 1007, the non-dominating sort is performed on the architecture store which refers to a repository or database that stores various neural network architectures along with their associated performance metrics evaluated on specific tasks or datasets. At block 1009, the intermediate full precision model is obtained. The first training stage is performed at block 1011, focusing on binarizing or quantizing the weights but the activation maps remain in full precision. The procedure after the first training stage (1011) and the second training stage (1019) is similar. The teacher model (1003) is quantized into multiple binary variants which serve as experts corresponding to clusters of training data as previously described. During inference, the experts are aggregated (as demonstrated by the aggregation in the block 1015, however, only one expert will be active as the probabilities associated with other experts will be 0 as previously explained. As a result, there will be a selective mechanism, which is represented as the lightweight gate control. Based on the aggregation over entire channels at block 1015, the gate control is performed to select one expert. In block 1013, a gate control is implemented to allow conditional computing on AI experts. The gate control generally refers to mechanism that regulates the flow of information through the layers of the network - particularly in neural networks. All of this procedure shown in Fig. 10 may be performed for an arbitrary layer n, and the procedure has to be sequentially repeated from layer 1 to layer Ln. Once completed, the same training mechanism and inference mechanism is implemented in the second training stage (1019), resulting in the final binarized network. As shown before, Fig. 10 illustrates the quantization of weights (first training stage) and activation (second training stage) of an arbitrary layer n.
[0100] During the second training stage, binarization is performed for the activation. Similar to the first training stage, block 1021 involves gate control (1021) to allow conditional computing on AI experts. At block 1023, aggregation is again performed to generate probabilities from the input dataset (1001). Upon completing the two-stage training at block 1025, a Pareto-optimal full binary AI model with expert ensemble is obtained. This Pareto-optimal full binary AI model balances the trade-off between accuracy and complexity, specialized in input content data.
[0101] Fig. 11 is a schematic diagram illustrating the use case of binarization of large language models (LLMs) according to an embodiment of the disclosure. The binarization of large language models (3B+ parameters) ensures the deployment of use cases based on natural language understanding, such as Smart Reply and Style correction, on embedded devices like smartphones and smartwatches. Vision-based stable diffusion models (generative AI) can be binarized to enable generative wallpaper use cases on smartwatches, which are currently deployed in higher bit precisions on smartphones. Particularly, a demonstration of the style correction use case with a binarized LLM and Long Range Radio (LoRA) model is shown in elements 1101, 1103, 1005, and 1007 of Fig. 11.
[0102] Fig 12 illustrates a demonstration of Stable Diffusion-based wallpaper generation use-case on a smartphone (1201) in INT-16 precision and the same on a smartwatch (1203) when binarized using the proposed AI model. In the proposed solution, an AI model is redesigned in full-precision to tailor it for binary quantization. The multi-objective NAS improves the network's capacity, thereby allowing binary representation to retain accuracy. The proposed solution reduces the loss in accuracy and improves the convergence of the binary neural network by serving as an accurate teacher model in knowledge distillation-based quantization-aware training. The training process is extremely fast, resulting in agile solutions that are high in accuracy and hardware efficient.
[0103] The binary quantization approach according to an embodiment of the disclosure is particularly advantageous for devices with limited computational resources, such as smartwatches. By converting the AI model into a binary format, the computational load is significantly reduced, making it feasible to run complex models on smaller devices without compromising on performance. This enables new possibilities for integrating advanced AI functionalities into a wider range of consumer electronics, enhancing user experiences with personalized and context-aware features. For instance, the ability to generate high-quality wallpapers on-the-fly on both smartphones and smartwatches can lead to more dynamic and engaging user interfaces.
[0104] The AI block in the proposed solution is quantized into multiple variants called experts, which are specialized in a particular region of the data. The training of the binary neural network is performed in two stages, enabling cascaded accuracy gain. Conditional computation is enabled to dynamically activate a single quantized expert AI block during runtime. This selective activation conserves computational resources and ensures that the relevant expert is utilized for a given task, thereby optimizing performance. A novel and unique entropy regularization technique is implemented for the selection and activation of only a single AI expert during training and inference. This technique ensures that the model remains efficient and accurate, as it minimizes redundancy and focuses computational efforts on the pertinent aspects of the data.
[0105] Further, the entropy regularization technique maintains the balance between model complexity and computational efficiency. By regulating the entropy, the system ensures that only the necessary computations are performed, which is particularly beneficial for real-time applications. This method enhances the model's performance and extends the battery life of electronic devices by reducing unnecessary power consumption. As a result, users can enjoy seamless and responsive AI-driven features without the need for frequent recharges, making the technology more practical and user-friendly in everyday scenarios.
[0106] In an embodiment of the disclosure, provided is a method of inference processing in a binarized artificial intelligence (AI) model by an electronic device. In an embodiment, the method includes inputting an input content to the binarized AI model. In an embodiment, the method includes identifying a pre-defined cluster to which the input content belongs using the binarized AI model. In an embodiment, the method includes activating a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster. In an embodiment, the method includes generating, by the electronic device (401), a final output by processing the input content received using the activated binary variant.
[0107] In an embodiment, the inputting of the input content includes inputting the input content to a binarization-aware NAS-AI model (907) operating in floating point precision; generating, by the binarization-aware NAS-AI model, first output, inputting the input content to a binary neural network with experts, and generating, by the binary neural network, second output. In an embodiment, weights of the binary neural network may be binarized.
[0108] In an embodiment, the inputting of the input content further includes determining a knowledge distillation loss, a training loss, and a entropy regularization loss based on the first output and the second output; and obtaining quantized activation.
[0109] In an embodiment, each of binary variant of the plurality of binary variants is personalized based on the pre-defined cluster of the input content.
[0110] In an embodiment, the pre-defined cluster represents user data.
[0111] In an embodiment, the binarized AI model is generated by enhancing a capacity of a full-precision AI model by adding at least one channel and at least one binarization friendly layer to the full-precision AI model.
[0112] In an embodiment, the adding of the at least one channel and the at least one binarization friendly layer to the full-precision AI model includes profiling the binarized AI model to identify sub-optimal AI blocks that inhibit efficient binarization, determining a quantization-aware search space to prepare and train quantization-friendly alternatives for the identified sub-optimal AI blocks, and adding the at least one channel and the at least one binarization friendly layer to the full-precision AI model based on the quantization-friendly alternatives for the identified sub-optimal AI blocks.
[0113] In an embodiment, each binary variant of the plurality of binary variants is relevant to the pre-defined clusters associated with the input content that is processed by the binarized AI model.
[0114] In an embodiment, each binary variant of the plurality of binary variants is pre-trained by the full-precision AI model. The training of the plurality of binary variants includes initializing the plurality of binary variants to the full-precision AI model, assigning random weights to the each binary variant of the plurality of binary variants, determining a probability for each of the random weights assigned to the each binary variant of the plurality of binary variants, and determining a weighted sum of the each binary variant of the plurality of binary variants to obtain a first output. In an embodiment, the weighted sum is determined as a sum of product of an output of the each of the binary variants and corresponding probabilities.
[0115] In an embodiment, the method includes determining a second output using the full-precision AI model. In an embodiment, the method includes determining a knowledge distillation loss by comparing the first output and the second output. In an embodiment, the method includes determining a training loss by comparing the first output and a ground truth. In an embodiment, the method includes determining an entropy regularization loss. In an embodiment, the method includes the entropy regularization loss ensures during inference the activation of only one binary variant among the plurality of binary variants; determining a total loss that is a sum of the knowledge distillation loss, the training loss and the entropy regularization loss. In an embodiment, the method includes determining the final output by minimizing the total loss. In an embodiment, the final output is a collection of tuned weights of each binary variant of the plurality of binary variants along with the weights.
[0116] In an embodiment, each binary variant of the plurality of binary variants are trained on a specific region of the input content through the entropy regularization loss.
[0117] embodiment, the method includes generating an output of the full-precision AI model by selecting one expert block of a plurality of expert blocks having a probability of one. In an embodiment, the probability of the other expert blocks of the plurality of expert blocks except for the one expert block having the probability of one are zero.
[0118] In an embodiment of the disclosure, provided is an electronic device (401) for inference processing of a binarized AI model. The electronic device (401) includes memory (405) storing instructions, and a processor (403), when executing the instructions stored in the memory (405), configured to input an input content to the binarized AI model, determine a pre-defined cluster to which the input content belongs using the binarized AI model, activate a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster, and generate a final output by processing the input content received using the activated binary variant.
[0119] In an embodiment, the binarized AI model is generated by enhancing a capacity of a full-precision AI model by adding at least one channel and at least one binarization-friendly layer to the full precision AI model.
[0120] In an embodiment, the electronic device (401) trains the plurality of binary variants by initializing the plurality of binary variants to the full-precision AI model, assigning random weights to each binary variant of the plurality of binary variants, determining a probability for the random weights assigned to each binary variant of the plurality of binary variants using a SoftMax based activation function, and determining a weighted sum of each binary variant of the plurality of binary variants to obtain the first output. In an embodiment, the weighted sum is determined as a sum of product of an output of each of the binary variants and corresponding probabilities.
[0121] In an embodiment, the processor (403) is configured to determine a second output using the full-precision AI model, determine a knowledge distillation loss by comparing the first output and the second output, determine a training loss by comparing the first output and a ground truth, determine an entropy regularization loss. In an embodiment, the entropy regularization loss ensures during inference the activation of only one binary variant among the plurality of binary variants. In an embodiment, the processor (403) determines a total loss that is a sum of the knowledge distillation loss, the training loss and the entropy regularization loss, and determines the final output by minimizing the total loss. In an embodiment, the final output is a collection of tuned weights of each binary variant of the plurality of binary variants along with the weights.
[0122] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the embodiments as described herein.
[0123] The method according to an embodiment of the disclosure may be implemented in a program command form that can be executed by various computer means, and may be recorded on computer-readable media. The computer-readable media may also include, alone or in combination with program commands, data files, data structures, and the like. Program commands recorded in the media may be the kind specifically designed and constructed for the disclosure or well-known and available to those of ordinary skill in the computer software field. Examples of the computer-readable media include magnetic media, such as hard disks, floppy disks, and magnetic tapes, optical media, such as compact disc read only memory (CD-ROM) and digital versatile disc (DVD), magneto-optical media such as floptical disks, and hardware devices, such as ROM, RAM, flash memory, and the like, specifically configured to store and execute program commands. Examples of the program commands may include high-level language codes that can be executed on a computer through an interpreter or the like, as well as machine language codes produced by a compiler.
[0124] An embodiment of the disclosure may be implemented in the form of a computer-readable recording medium including an instruction that is executable by a computer, such as a program module that is executed by a computer. The computer-readable recording medium may be an arbitrary available medium which can be accessed by a computer, and may include a volatile or non-volatile medium and a separable or non-separable medium. Further, the computer-readable recording medium may include a computer storage medium and a communication medium. The computer storage medium may include volatile and non-volatile media and separable and non-separable media implemented by an arbitrary method or technology for storing information such as a computer readable instruction, a data structure, a program module, or other data. The communication medium may generally include a computer readable instruction, a data structure, a program module, other data of a modulated data signal such as a carrier wave, or another transmission mechanism, and include an arbitrary information transmission medium. Also, an embodiment of the disclosure may be implemented as a computer program including instructions executable by a computer, such as a computer program that is executed by a computer, or as a computer program product.
[0125] The computer-readable storage media may be provided in a form of non-transitory storage media. Herein, the term 'non-transitory storage medium' means that it is a tangible device, and does not include a signal (e.g., an electromagnetic wave), but this term does not differentiate between where data is semi-permanently stored in the storage medium and where the data is temporarily stored in the storage medium. For example, a 'non-transitory storage medium' may include a buffer in which data is temporarily stored.
[0126] According to an embodiment of the disclosure, the method according to various embodiments of the disclosure may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., CD-ROM), or be distributed (e.g., downloadable or uploadable) online via an application store or between two user devices (e.g., smart phones) directly. When distributed online, at least a part of the computer program product (e.g., a downloadable app) may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as a memory of the manufacturer's server, a server of the application store, or a relay server.
Claims
1.A method of inference processing in a binarized artificial intelligence (AI) model by an electronic device (401), comprising:inputting an input content to the binarized AI model;identifying a pre-defined cluster to which the input content belongs using the binarized AI model;activating a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster; andgenerating a final output by processing the input content received using the activated binary variant.2.The method of claim 1,wherein the inputting of the input content comprises:inputting the input content to a binarization-aware NAS-AI model (907) operating in floating point precision;generating, by the binarization-aware NAS-AI model, first output;inputting the input content to a binary neural network with experts; andgenerating, by the binary neural network, second output,wherein weights of the binary neural network are binarized.3.The method of claim 2, wherein the inputting of the input content further comprises:determining a knowledge distillation loss, a training loss, and a entropy regularization loss based on the first output and the second output; andobtaining quantized activation.4.The method of any one of claims 1 to 3,wherein each of binary variant of the plurality of binary variants is personalized based on the pre-defined cluster of the input content,wherein the pre-defined cluster represents user data.5.The method of any one of claims 1 to 4,wherein the binarized AI model is generated by enhancing a capacity of a full-precision AI model by adding at least one channel and at least one binarization friendly layer to the full-precision AI model.6.The method of any one of claims 1 to 5,wherein the adding of the at least one channel and the at least one binarization friendly layer to the full-precision AI model comprises:profiling the binarized AI model to identify sub-optimal AI blocks that inhibit efficient binarization;determining a quantization-aware search space to prepare and train quantization-friendly alternatives for the identified sub-optimal AI blocks; andadding the at least one channel and the at least one binarization friendly layer to the full-precision AI model based on the quantization-friendly alternatives for the identified sub-optimal AI blocks.7.The method of any one of claims 1 to 6, wherein each binary variant of the plurality of binary variants is relevant to the pre-defined clusters associated with the input content that is processed by the binarized AI model.8.The method of any one of claims 1 to 7,wherein each binary variant of the plurality of binary variants is pre-trained by the full-precision AI model, wherein training of the plurality of binary variants comprises:initializing the plurality of binary variants to the full-precision AI model;assigning random weights to the each binary variant of the plurality of binary variants;determining a probability for each of the random weights assigned to the each binary variant of the plurality of binary variants; anddetermining a weighted sum of the each binary variant of the plurality of binary variants to obtain a first output,wherein the weighted sum is determined as a sum of product of an output of the each of the binary variants and corresponding probabilities.9.The method of any one of claims 1 to 8, comprises:determining a second output using the full-precision AI model;determining a knowledge distillation loss by comparing the first output and the second output;determining a training loss by comparing the first output and a ground truth;determining an entropy regularization loss, wherein the entropy regularization loss ensures during inference the activation of only one binary variant among the plurality of binary variants;determining a total loss that is a sum of the knowledge distillation loss, the training loss and the entropy regularization loss; anddetermining the final output by minimizing the total loss, wherein the final output is a collection of tuned weights ofeachbinary variant of the plurality of binary variants along with the weights.10.The method of any one of claims 1 to 9, wherein the each binary variant of the plurality of binary variants are trained on a specific region of the input content through the entropy regularization loss.11.The method of any one of claims 1 to 10, further comprising:generating an output of the full-precision AI model by selecting one expert block of a plurality of expert blocks having a probability of one,wherein the probability of the other expert blocks except for the one expert block having the probability of one are zero.12.An electronic device (401) for inference processing of a binarized AI model, comprising:memory (405) storing instructions;a processor (403), when executing the instructions stored in the memory (405), configured to:input an input content to the binarized AI model;determine a pre-defined cluster to which the input content belongs using the binarized AI model;activate a binary variant from a plurality of binary variants corresponding to the determined pre-defined cluster; andgenerate a final output by processing the input content received using the activated binary variant.13.The electronic device (401) of claim 12, wherein the binarized AI model is generated by enhancing a capacity of a full-precision AI model by adding at least one channel and at least one binarization-friendly layer to the full precision AI model.14.The electronic device (401) of any one of claims 12 to 13, wherein processor (403) trains the plurality of binary variants byinitializing the plurality of binary variants to the full-precision AI model;assigning random weights to each binary variant of the plurality of binary variants;determining a probability for the random weights assigned to each binary variant of the plurality of binary variants using a SoftMax based activation function; anddetermining a weighted sum of each binary variant of the plurality of binary variants to obtain the first output, wherein the weighted sum is determined as a sum of product of an output of each of the binary variants and corresponding probabilities.15.The electronic device (401) of any one of claims 12 to 14, wherein the processor (403) is configured to:determine a second output using the full-precision AI model;determine a knowledge distillation loss by comparing the first output and the second output;determine a training loss by comparing the first output and a ground truth;determine an entropy regularization loss, wherein the entropy regularization loss ensures during inference the activation of only one binary variant among the plurality of binary variants;determine a total loss that is a sum of the knowledge distillation loss, the training loss and the entropy regularization loss; anddetermine the final output by minimizing the total loss, wherein the final output is a collection of tuned weights ofeachbinary variant of the plurality of binary variants along with the weights.
Citation Information
Patent Citations
Activation Functions for Deep Neural Networks
US20190147323A1
Lookup table activation functions for neural networks
US20210397596A1
Hardware accelerator method and device
US20220383103A1