On-device unified inference-training pipeline for hybrid precision forward-back propagation through heterogeneous floating point graphics processing units (GPUs) and fixed point digital signal processors (DSPs)

By employing a mixed-precision processing approach on edge devices, utilizing fixed-point processors for inference and selectively using floating-point processors for training, the challenges of deploying and training neural network models on edge devices are addressed, achieving efficient and stable inference and training.

CN121844349APending Publication Date: 2026-04-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Edge devices, due to limited computing resources, struggle to efficiently deploy and train complex neural network models, especially in real-time inference and training processes, where memory consumption and high latency issues exist.

Method used

A mixed-precision processing approach is adopted, using a fixed-point processor for inference tasks and a floating-point processor for training selectively based on computational workload, system conditions, or neural network model errors, to achieve concurrent execution of inference and training.

Benefits of technology

It improves the accuracy and training stability of neural network models, reduces training time, and enables efficient inference and training on resource-constrained edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844349A_ABST
    Figure CN121844349A_ABST
Patent Text Reader

Abstract

A processor-implemented method for hybrid precision inference and training includes receiving inputs through an artificial neural network (ANN) model. The ANN model processes the input using a fixed point processor. The processing of the input is performed in a fixed point format to compute an inference. A model update for the ANN model is selectively calculated using either a floating point processor or the fixed point processor based on the loss of the ANN model. The model update is calculated in a floating point format or the fixed point format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] All aspects of this disclosure relate in general to processing, and more specifically to an on-device unified inference-training pipeline with mixed-precision forward-backward propagation via heterogeneous floating-point graphics processing units (GPUs) and fixed-point digital signal processors (DSPs). Background Technology

[0002] Artificial neural networks can comprise interconnected groups of artificial neurons (e.g., neuron models). Artificial neural networks can be computing devices or represented as methods to be performed by computing devices. Artificial neural networks, including feedforward neural networks, convolutional neural networks (CNNs), multilayer perceptrons (MLPs), transformers, graph neural networks (GNNs), recurrent neural networks (RNNs), and others, have numerous applications. Specifically, these neural network architectures are used in various technologies such as image recognition, speech recognition, acoustic scene classification, keyword retrieval, autonomous driving, and other classification tasks.

[0003] Given the many useful applications of neural networks, the demand for using neural networks on edge devices such as smartphones is constantly increasing. However, edge devices have limited computing resources, and generalized models can utilize more complex networks and more computation. Therefore, the memory footprint and high latency of neural networks make their use challenging, especially for efficient deployment and inference on resource-constrained devices.

[0004] Deep neural networks are gaining popularity due to their ability to solve complex problems. Similarly, edge devices such as smartphones are widely used due to their portability and practicality. Therefore, the deployment of deep learning on edge devices for real-time inference is an area of ​​interest. However, training on edge devices, which may have limited computing resources, can be challenging due to the size, memory consumption, and complexity of neural network models. Summary of the Invention

[0005] This disclosure is set forth in the independent claims. Some aspects of this disclosure are described in the dependent claims.

[0006] In some aspects of this disclosure, a processor-implemented method executed by at least one processor includes receiving input through an artificial neural network (ANN) model. The processor-implemented method further includes processing the input using a fixed-point processor through the ANN model. This processing of the input is performed in a fixed-point format to compute inference. The processor-implemented method also includes selectively computed model updates for the ANN model using either a floating-point processor or the fixed-point processor based on the loss of the ANN model. The model updates are computed in either a floating-point format or the fixed-point format.

[0007] Various aspects of this disclosure relate to an apparatus including components for receiving input through an artificial neural network (ANN) model. The apparatus also includes components for processing the input using a fixed-point processor through the ANN model. This processing of the input is performed in a fixed-point format to compute inference. The apparatus further includes components for selectively computed model updates for the ANN model using either a floating-point processor or the fixed-point processor based on the loss of the ANN model. The model updates are computed in either a floating-point format or the fixed-point format.

[0008] In some aspects of this disclosure, a non-transitory computer-readable medium is disclosed, on which non-transitory program code is recorded. This program code is executed by a processor and includes program code for receiving input through an artificial neural network (ANN) model. The program code also includes program code for processing the input using a fixed-point processor through the ANN model. This processing of the input is performed in a fixed-point format to compute inference. The program code also includes program code for selectively computed model updates for the ANN model using a floating-point processor or the fixed-point processor based on the loss of the ANN model. The model updates are computed in either floating-point or fixed-point format.

[0009] Various aspects of this disclosure relate to an apparatus having a memory and one or more processors coupled to the memory. The processor is configured to receive input via an artificial neural network (ANN) model. The processor is also configured to process the input via the ANN model using a fixed-point processor. This processing of the input is performed in a fixed-point format to compute inference. The processor is further configured to selectively compute model updates for the ANN model using either a floating-point processor or the fixed-point processor based on the loss of the ANN model. The model updates are computed in either a floating-point format or the fixed-point format.

[0010] Additional features and advantages of this disclosure will be described below. Those skilled in the art will understand that this disclosure can be readily used as the basis for modifying or designing other structures for implementing the same purposes as this disclosure. Those skilled in the art will also recognize that such equivalent constructions do not depart from the teachings of this disclosure as set forth in the appended claims. Novel features considered characteristic of this disclosure, in both their organization and manner of operation, along with further objects and advantages, will be better understood when considered in conjunction with the accompanying drawings. However, it is to be clearly understood that each drawing is provided for illustrative and descriptive purposes only and is not intended to be a definition of a limitation of this disclosure. Attached Figure Description

[0011] The features, substance, and advantages of this disclosure will become more apparent when understood in conjunction with the accompanying drawings, in which the same reference numerals are consistently used for identification.

[0012] Figure 1 Example implementations of neural networks using a system-on-a-chip (SoC) (including a general-purpose processor) according to certain aspects of this disclosure are illustrated.

[0013] Figure 2A , Figure 2B and Figure 2C These are illustrations of neural networks according to various aspects of this disclosure.

[0014] Figure 2D This is a diagram illustrating exemplary deep convolutional networks (DCNs) according to various aspects of this disclosure.

[0015] Figure 3 This is a block diagram illustrating exemplary deep convolutional networks (DCNs) according to various aspects of this disclosure.

[0016] Figure 4 This is a block diagram illustrating exemplary software architectures that enable modularization of artificial intelligence (AI) functions according to various aspects of this disclosure.

[0017] Figure 5 This is a block diagram illustrating example architectures for performing mixed-precision inference and training according to various aspects of this disclosure.

[0018] Figure 6 This is a flowchart illustrating an example processor implementation of a method for mixed-precision inference and training according to various aspects of this disclosure. Detailed Implementation

[0019] The detailed description that follows, taken in conjunction with the accompanying drawings, is intended as a description of various configurations and not as representing only configurations in which the described concepts can be practiced. To provide a comprehensive understanding of the various concepts, the detailed description includes specific details. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, to avoid obscuring such concepts, well-known structures and components are shown in block diagram form.

[0020] Based on the teachings, those skilled in the art will recognize that the scope of this disclosure is intended to cover any aspect of this disclosure, whether implemented independently of or in combination with any other aspect of this disclosure. For example, an apparatus or method may be implemented using any number of the aspects described. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using other structures, functionalities, or structures and functionalities that complement or differ from the various aspects of this disclosure described. It should be understood that any aspect of this disclosure may be embodied by one or more elements of the claims.

[0021] The word “exemplary” is used to mean “serving as an example, instance, or illustration.” Any aspect described as “exemplary” need not be interpreted as superior to or better than other aspects.

[0022] While specific aspects have been described, numerous variations and substitutions of these aspects fall within the scope of this disclosure. Although some benefits and advantages of preferred aspects have been mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or purpose. Rather, aspects of this disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the accompanying drawings and the following description of preferred aspects. The detailed description and drawings are merely illustrative and not limiting of this disclosure, the scope of which is defined by the appended claims and their equivalents.

[0023] As described, the deployment of deep learning models for real-time inference on edge devices is an area of ​​interest. However, deploying deep learning models on such devices can be challenging, as the models may be too large and complex for resource-constrained edge devices.

[0024] An MLP can have an input layer, an output layer, and one or more hidden neuron layers. Inputs can be combined with initial weights in a weighted sum manner and can be processed by activation functions. Each layer provides a linear combination to the next layer of the MLP until an output is generated.

[0025] Transformer architectures have shown improvements in language modeling and natural language processing (NLP) tasks. Transformer models learn context or meaning in sequential data by tracking relationships in sequential data (e.g., words in a sentence). Transformer models can leverage attention to detect relationships between various parts of an input sequence, which can be referred to as tokens.

[0026] Based on traditional converter architectures, such as Bidirectional Encoder Representation from Converter (BERT), Robustly Optimized BERT Method (RoBERTa), XLNet, Converter-XL, and the Generative Pre-trained Converter (GPT) family (e.g., GPT-2), language models can be pre-trained on large corpora of unlabeled text. Therefore, such converter architectures have become common building blocks in regular NLP pipelines as well as in other fields such as computer vision and audio processing.

[0027] While offering performance improvements in many applications, converter-based models and other artificial neural networks (ANNs) can be extremely large, sometimes exceeding billions of parameters. Therefore, efficiently deploying such ANN models on resource-constrained embedded systems, including mobile devices (e.g., smartphones) and Internet of Things (IoT) devices, as well as some systems in data centers, is challenging due to increased latency, energy consumption, and excessive memory footprint.

[0028] Furthermore, in mobile devices or smartphones, on-device training likely involves primarily sensor data processing and user behavior recognition. Due to privacy, security, and real-time power / performance specifications, both neural network model inference and training can be performed within a sensor hub digital signal processor (DSP). Smaller neural network models can be quantized to 8-bit integer (INT8) format and can be trained using forward-backward propagation in INT8 format on a DSP.

[0029] Neural network quantization reduces memory consumption by using low-bit precision (e.g., four-bit integer (INT4) or eight-bit integer (INT8) formats) for the weights and activation tensors. Furthermore, neural network quantization reduces inference time and improves energy efficiency by employing low-bit fixed-point operations instead of floating-point operations.

[0030] However, quantization can introduce additional noise into neural networks, which can lead to decreased model accuracy and increased computational complexity. For example, transformer models may have numerous outliers in their activations. These activation outliers can cause large quantization errors.

[0031] It may be desirable to determine a neural network model topology and structure that satisfies both inference / data processing specifications and training forward-backward propagation. That is, it may be desirable to determine a neural network model in which, for example, backward gradient updates may be stable for resource-constrained edge devices (e.g., smartphones) under computationally limited conditions.

[0032] To address these and other challenges, aspects of this disclosure relate to a unified inference and training pipeline. Neural network models can use fixed-point processors to perform inference tasks. For example, a neural network model can be selectively trained using either a floating-point processor or a fixed-point processor based on one or more of the following: computational workload, system conditions (e.g., power level or memory consumption), or neural network model error (e.g., loss). In some aspects, inference computation and training can be performed concurrently.

[0033] Specific aspects of the subject matter described in this disclosure may be implemented to achieve one or more of the following potential advantages. In some examples, the described techniques (e.g., using a fixed-point processor to perform inference tasks, or using a floating-point processor to selectively train a neural network model) can improve the accuracy and training stability of the neural network model and reduce training time.

[0034] Figure 1 An example implementation of a system-on-a-chip (SOC) 100 is illustrated, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for mixed-precision inference and training. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a weighted neural network), latency, frequency window (bin) information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or may be loaded from memory block 118.

[0035] The SOC 100 may also include additional processing blocks tailored for specific functions, such as a GPU 104, a DSP 106, a connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 capable of, for example, detecting and recognizing gestures. In one specific implementation, an NPU 108 is implemented within a CPU 102, a DSP 106, and / or a GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120, which may include a global positioning system.

[0036] The SOC 100 may be based on the ARM instruction set. In various aspects of this disclosure, instructions loaded into the general-purpose processor 102 may include code for receiving input through an artificial neural network (ANN) model. The general-purpose processor 102 may also include code for processing the input using a fixed-point processor through the ANN model. This processing of the input is performed in fixed-point format to compute inference. The general-purpose processor 102 may also include code for selectively computed model updates for the ANN model using either a floating-point processor or the fixed-point processor based on the loss of the ANN model. The model updates are computed in either floating-point or fixed-point format.

[0037] Deep learning architectures perform object recognition tasks by learning to represent inputs at progressively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses a major bottleneck of traditional machine learning. Before deep learning, machine learning methods for object recognition problems often relied heavily on human-designed features, possibly in conjunction with shallow classifiers. Shallow classifiers could be two-class linear classifiers, where a weighted sum of feature vector components is compared to a threshold to predict which class the input belongs to. Human-designed features could be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, while deep learning architectures can learn to represent features similar to those that human engineers might design, this requires training. Furthermore, deep networks can learn to represent and recognize novel types of features that humans might not have considered.

[0038] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0039] Deep learning architectures perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to identify features such as wheels, windshields, and others. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0040] Neural networks can be designed to have multiple connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the recognition of higher-level concepts can aid in discerning specific lower-level features of the input.

[0041] The connections between layers in a neural network can be fully connected or partially connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, neurons in the first layer can transmit their outputs to each neuron in the second layer, such that each neuron in the second layer receives input from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, neurons in a first layer can connect to a finite number of neurons in a second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in the layer will have the same or similar connectivity pattern, but the connection strength can have different values ​​(e.g., 210, 212, 214, and 216). The connectivity pattern of locally connected layers can produce spatially different receptive fields in higher layers because neurons in higher layers in a given region can receive inputs that are tuned to the characteristics of a restricted portion of the total input to the network through training.

[0042] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strength associated with the input for each neuron in the second layer is shared (e.g., 208). Convolutional neural networks may be well-suited for problems where the spatial location of the input is meaningful.

[0043] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2DA detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as an in-vehicle camera) is illustrated. The DCN 200 in this example can be trained to identify traffic signs and the numbers provided on them. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or traffic lights.

[0044] Supervised learning can be used to train the DCN 200. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and forward passes can then be computed to produce output 222. The DCN 200 may include a feature extraction part and a classification part. Upon receiving image 226, convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel used for convolutional layer 232 may be a 5x5 kernel that generates a 28x28 feature map. In this example, because four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at convolutional layer 232. The convolutional kernel may also be referred to as a filter or convolutional filter.

[0045] The first set of feature maps 218 can be subsampled by a max-pooling layer (not shown) to generate a second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (e.g., 14×14) is smaller than the size of the first set of feature maps 218 (e.g., 28×28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more sets of subsequent feature maps (not shown).

[0046] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature in the second feature vector 228 may include a number corresponding to a possible feature of the image 226, such as "sign", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 can be the probability that the image 226 includes one or more features.

[0047] In this example, the probabilities for "sign" and "60" in output 222 are higher than the probabilities for other numbers in output 222 (such as "30", "40", "50", "70", "80", "90", and "100"). Before training, output 222 generated by DCN 200 may be incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the baseline ground truth (e.g., "sign" and "60") of image 226. The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is more closely aligned with the target output.

[0048] To adjust the weights, the learning algorithm computes the gradient vector of the weights. The gradient indicates by how much the error will increase or decrease as the weights are adjusted. At the top layers, the gradient corresponds directly to the values ​​of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient depends on the values ​​of the weights and the error gradient computed in the higher layers. The weights can then be adjusted to reduce the error. This method of adjusting weights is called "backpropagation" because it involves "passing backward" through the neural network.

[0049] In practice, the error gradient of the weights can be calculated using a small number of examples to make the calculated gradient approximate the true error gradient. This approximation method is called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, a new image (e.g., a speed limit sign in image 226) can be presented to DCN 200, and output 222 can be generated through the forward pass of DCN 200. This output can be considered as an inference or prediction of DCN 200.

[0050] Deep Belief Networks (DBNs) are probabilistic models that include multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of training datasets. DBNs are obtained by stacking layers of Restricted Boltzmann Machines (RBMs). An RBM is a type of artificial neural network that learns a probability distribution from a set of inputs. Because RBMs can learn a probability distribution without information about the class each input should be classified into, they are often used for unsupervised learning. Using a hybrid paradigm of supervised and unsupervised learning, the bottom RBM of a DBN can be trained unsupervised and used as a feature extractor, while the top RBM can be trained supervisedly (on the joint distribution of inputs from the previous layer and the target class) and used as a classifier.

[0051] DCN is a network of convolutional networks configured with additional pooling and normalization layers. DCN has achieved state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both the input and output targets are known for many paradigms and are used to modify the network's weights using gradient descent.

[0052] DCNs can be feedforward networks. Furthermore, as described above, connections from neurons in the first layer of a DCN to a set of neurons in the next higher layer are shared across neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational cost of a DCN may be much smaller than that of a similarly sized neural network that includes recurrent or feedback connections.

[0053] The processing at each layer of a convolutional network can be thought of as a spatially invariant template or base projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of the convolutional connections can be thought of as forming a feature map in the next layer, where each element in the feature map (e.g., 220) receives input from a range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values ​​in the feature map can be further processed with nonlinearities (such as correction, max(0,x)). Values ​​from neighboring neurons can be further pooled, which corresponds to downsampling and provides additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied through lateral inhibition between neurons in the feature map.

[0054] Figure 3 This is a block diagram illustrating a DCN 350. A DCN 350 can include multiple layers of different types based on connectivity and weight sharing. For example... Figure 3 As shown, DCN 350 includes convolutional blocks 354A and 354B. Each convolutional block in convolutional blocks 354A and 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.

[0055] Although only two convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks 354A and 354B may be included in the DCN 350 according to design preferences.

[0056] Convolutional layer 356 may include one or more convolutional filters that can be applied to the input data to generate feature maps. Normalization layer 358 may normalize the output of the convolutional filters. For example, normalization layer 358 may provide whitening or lateral suppression. Max pooling layer 360 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.

[0057] Parallel filter banks of deep convolutional networks can be loaded onto an SOC 100 (e.g., Figure 1 The CPU 102 or GPU 104 of the SOC 100 can be used to achieve high performance and low power consumption. In an alternative implementation, a parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. In addition, the DCN 350 can access other processing blocks that may exist on the SOC 100, such as the sensor processor 114 and navigation module 120, which are dedicated to sensors and navigation, respectively.

[0058] The DCN 350 may also include one or more fully connected layers 362 (FC1 and FC2). The DCN 350 may also include logistic regression (LR) layers 364. Weights (not shown) to be updated are located between each of the layers 356, 358, 360, 362, and 364 of the DCN 350. The output of each layer (e.g., 356, 358, 360, 362, and 364) can be used as input to the next layer in the DCN 350 (e.g., 356, 358, 360, 362, and 364) to learn hierarchical feature representations from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) supplied at the first convolutional block in convolutional block 354A. The output of the DCN 350 is a classification score 366 of the input data 352. The classification score 366 may be a set of probabilities, where each probability is the probability that the input data includes a feature from the feature set.

[0059] Figure 4 This is a block diagram illustrating an exemplary software architecture 400 with modular artificial intelligence (AI) functionality. According to various aspects of this disclosure, using architecture 400, a system-on-a-chip (SoC) 420 (which may be similar to...) can be designed... Figure 1 Various processing blocks of the SoC 400 (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) support mixed-precision inference and training for AI applications 402. The architecture 400 can be included, for example, in a computing device such as a smartphone.

[0060] AI application 402 can be configured to invoke functions defined in user space 404, which may, for example, provide the detection and recognition of a scene indicating the current location of the computing device (including architecture 400). For example, AI application 402 may configure microphones and cameras differently depending on whether the recognized scene is an office, lecture hall, restaurant, or an outdoor environment such as a lake. AI application 402 may make requests to compiled program code associated with libraries defined in the AI ​​Function Application Programming Interface (API) 406. This request may ultimately rely on the output of a deep neural network configured to provide inferred responses based on, for example, video and location data.

[0061] Runtime engine 408 (which may be compiled code of a runtime framework) may be further accessible to AI application 402. AI application 402 may cause runtime engine 408 to request inference, for example, at specific time intervals or when triggered by events detected by the user interface of AI application 402. Upon causing runtime engine 408 to provide an inference response, the runtime engine may then signal to the operating system (OS) space 410 running on SOC 420, such as kernel 412. In some examples, kernel 412 may be a LINUX kernel. The operating system may then enable sequential relaxation of quantization to be performed on CPU 422, DSP 424, GPU 426, NPU 428, or some combination thereof. CPU 422 may be directly accessible by the operating system, while other processing blocks may be accessed via drivers, such as drivers 414, 416, or 418 for DSP 424, GPU 426, or NPU 428, respectively. In an exemplary example, the deep neural network may be configured to run on a combination of processing blocks such as CPU 422, DSP 424 and GPU 426, or on NPU 428.

[0062] As described, aspects of this disclosure relate to a unified inference and training pipeline. Neural network models can use fixed-point processors to perform inference tasks. For example, a neural network model can be selectively trained using either a floating-point processor or a fixed-point processor based on one or more of the following: computational workload, system conditions (e.g., power level or memory consumption), or neural network model errors (such as loss). In some aspects, inference computation and training can be performed concurrently.

[0063] Figure 5 This is a block diagram illustrating an example architecture 500 for performing mixed-precision inference and training according to various aspects of this disclosure. References Figure 5Example architecture 500 may include an artificial neural network (ANN) model 502. The ANN model 502 may be implemented on an edge device, such as a smartphone, tablet computer, autonomous vehicle, IoT device, or other device. Edge devices may have limited computing resources, memory capacity, or power resources. However, this disclosure is not limited to edge devices, but may also be implemented on servers, such as cloud servers.

[0064] ANN model 502 can be represented as a graph, such as a directed acyclic graph (DAG). A directed graph is a tuple G = (V, E), where V is a set of vertices or nodes, and E... V × V is the set of edges between vertices. A tuple is a finite, ordered list or sequence of elements. A cycle is a sequence of edges i→ with at least one directed edge. →i. Therefore, a directed acyclic graph (DAG) is a finite directed graph that does not have direct cycles. Each node in these nodes can represent a task or operation to be performed. These nodes can represent operations to be performed, and these edges can indicate the relationships between nodes represented by weight parameters.

[0065] In some respects, the ANN model 502 may include (but is not limited to) such as multilayer perceptrons (MLPs), convolutional neural networks (CNNs) (e.g., Figure 3 ANN models can include 350), graph neural networks (GNNs), transformer models, or other types of neural networks. An ANN model 502 can include multiple layers. For example... Figure 5 As shown, the ANN model 502 may include an input layer 504, a hidden layer 506, and an output layer 508. Figure 5 The example shown has only three layers, which is for illustrative purposes only and not a limitation. Rather, it should be understood that the ANN model 502 can have any number of layers.

[0066] Each of these layers (e.g., 504, 506, or 508) may include a set of nodes. Nodes in input layer 504 may be fully connected to hidden layer 506, which may also be fully connected to output layer 508. That is, input layer 504 may be fully connected to every node in hidden layer 506. Each of these nodes performs a computation and generates an output, which may be fed forward to nodes in subsequent layers, which perform corresponding computations and similarly provide outputs to another subsequent layer, until an inference is determined in the output layer. Computations at one or more hidden layers (e.g., 506) between the input layer (e.g., 504) and the output layer (e.g., 508) may be considered intermediate operations.

[0067] ANN model 502 can be configured to receive input including (but not limited to) sensor data 510 (e.g., gyroscope data or Global Positioning Satellite (GPS) data) or other sensor data (such as images, audio, or speech input via a microphone). The input (e.g., 510) may be provided via a sensor hub or other sensors (e.g., 114). For example, the input (e.g., 510) may be supplied in a fixed-point format (such as a four-bit integer (INT4), an eight-bit integer (INT8) format, or other fixed-point formats). In some aspects, the sensor data 510 may be supplied in a floating-point format (such as a sixteen-bit floating-point (FP16) or a thirty-two-bit floating-point (FP32) format, or other floating-point formats).

[0068] Furthermore, the ANN model 502 can be configured to perform computations on inputs (e.g., 510) in a fixed-point format (e.g., INT8). In some aspects, the ANN model 502 can be configured to receive inputs (e.g., 510) in a floating-point format (e.g., FP32) and can perform quantization operations to convert the floating-point format inputs (e.g., FP32) to a fixed-point format (e.g., INT8). The ANN model 502 can be configured to process the inputs (e.g., 510) through layers (e.g., 504, 506, 508) of the ANN model 502 to perform inference tasks. For example, the ANN model 502 can be configured to perform tasks such as image classification, object detection, location and mapping, speech recognition, or other inference tasks.

[0069] Example architecture 500 can provide a unified inference-training pipeline including mixed-precision forward-backward propagation. For example, in forward propagation, computations for inference of ANN model 502 can be performed on a fixed-point processor 514 in INT8 or other fixed-point formats. Fixed-point processor 514 may include (but is not limited to) a DSP (e.g., 106) or a CPU (e.g., 102).

[0070] On the other hand, the ANN model 502 can be selectively trained during backpropagation using either a floating-point processor 516 or a fixed-point processor 514. That is, the computation of weight updates can be performed using either a floating-point processor 516 in floating-point format (e.g., FP16 or FP32 format) or a fixed-point processor 514, which may include (but is not limited to) a GPU (e.g., 104) or an NPU (e.g., 108).

[0071] In various ways, the computation of ANN model 502 can compute the model loss during forward propagation (e.g., concurrently with inference computation). For example, in one example, the output at output layer 508 can be compared with a ground truth example to determine the difference, which can be considered the model loss. Based on this loss, example architecture 500 can choose whether to train ANN model 502 using fixed-point processor 514 in fixed-point format or using floating-point processor 516 in floating-point format. For example, example architecture 500 can compare the model loss with one or more predefined thresholds.

[0072] In one example, if the model loss is greater than a predefined threshold (e.g., 0.10 or greater), the ANN model 502 can be trained in floating-point format using a floating-point processor 516. Weights and hidden layer tensors, inputs, or outputs (or subsets thereof) can be stored in memory for backpropagation. Weights and hidden layer tensors, inputs, or outputs can be in fixed-point format (e.g., eight-bit integers (INT8) or four-bit integers (INT4)). The floating-point processor 516 (e.g., GPU 104) can receive the stored forward propagation data (e.g., weights, hidden layer tensors, inputs, outputs, and loss) in fixed-point format, convert the forward propagation data to floating-point format (e.g., FP32), and perform backpropagation computation. In some aspects, the floating-point processor 516 can convert updated weights to fixed-point format. The floating-point processor 516 can transfer the weights back to the fixed-point processor 514, which can update the ANN model weights accordingly.

[0073] In this scenario, the model accuracy may be low enough that the accuracy improvement achieved by training the ANN model 502 with floating-point values ​​using a floating-point processor 516 can outweigh the expenditure of computing resources (e.g., processing, memory consumption, and power dissipation).

[0074] On the other hand, if the model loss is less than a predefined threshold (e.g., 0.01 or less than 0.05), the fixed-point processor 514 can be used to train the model. In this scenario, the model accuracy may be high enough that the marginal improvement in accuracy may not outweigh the expenditure of computational resources (e.g., processing, memory consumption, and power dissipation) used to train the ANN model 502 in floating-point format using the floating-point processor 516.

[0075] In some respects, the computations for inference performed on the ANN model 502 using a fixed-point processor 514 can be performed concurrently with computations performed using a floating-point processor 516 to determine weight updates in backpropagation.

[0076] In some respects, the loss of an ANN model can be conditionally calculated based on factors such as computational workload, memory capacity, or power budget (e.g., battery power). For example, on resource-constrained edge devices such as smartphones, users may be performing multiple tasks concurrently. Users may be playing online video games or virtual reality games or environments, streaming video, or interacting in social media scenarios. Therefore, the computational workload of other applications running on the smartphone may be high enough that the additional computation of the model loss of the ANN model 502 could hinder the performance of other applications and reduce user satisfaction. Therefore, if the workload is higher than a predefined threshold (or the available memory capacity and / or power budget is lower than a predefined threshold), the loss may not be calculated. In some respects, the determination of whether to calculate the loss can be based on a combination of such factors. Similarly, forward and backward propagation can be performed sequentially in batches based on the power budget or the availability of computational resources (e.g., workload and / or available memory capacity).

[0077] In some respects, the frequency of loss calculations may be determined based on, for example, device performance (e.g., the accuracy of ANN model 502) or power budget. Additionally, in various respects, loss calculations may be performed after the current processing workload of the fixed-point processor 514 has been completed. In some respects, the amount of forward propagation data stored may also be determined based on, for example, one or more of available memory, power budget, or workload.

[0078] Therefore, ANN model 502 can flexibly compute inferences and be trained while making trade-offs between computing environment (e.g., edge devices or servers), model accuracy, and resource constraints. In doing so, the stability of training complex ANN models can be increased.

[0079] Furthermore, the forward propagation quantization error can be considered training-aware. That is, unlike conventional methods where forward and backward propagation are independent processes, in various aspects of this disclosure, forward propagation can be considered a sub-module of training or as training forward propagation. The inferred forward propagation quantization error and the training forward propagation quantization error can be substantially the same, or can be completely identical in some respects. Thus, the inferred forward propagation quantization error may be in fixed-point (e.g., in a fixed-point processor such as a DSP). Therefore, during training, data from the DSP can be copied and converted to FP32 or FP16 in a floating-point processor (e.g., a GPU). In doing so, the quantization error can also be converted from fixed-point to floating-point, but the error can be preserved.

[0080] Figure 6This is a flowchart illustrating an example processor implementation of method 600 for mixed-precision inference and training according to various aspects of this disclosure. For example, the processor-implemented method 600 may be executed by one or more processors, such as CPUs (e.g., 102, 422), GPUs (e.g., 104, 426), DSPs (e.g., 106, 424), and / or NPUs (e.g., 108, 428)).

[0081] like Figure 6 As shown, the ANN model receives input at box 602. (See reference...) Figure 5 As described, the ANN model 502 can be configured to receive input including (but not limited to) sensor data 510 (e.g., gyroscope data or Global Positioning Satellite (GPS) data) or other sensor data (such as images, audio, or speech input via a microphone). The input (e.g., 510) may be provided via a sensor hub or other sensors (e.g., 114). For example, the input (e.g., 510) may be supplied in a fixed-point format (such as INT4 or INT8 or other fixed-point formats). In some aspects, the sensor data 510 may be supplied in a floating-point format (such as FP16 or FP32 or other floating-point formats).

[0082] At box 604, the ANN model uses a fixed-point processor to process the input. This processing of the input is performed in a fixed-point format to compute inferences. For example, see reference... Figure 5 As described, the ANN model 502 can be configured to perform computations on inputs (e.g., 510) in a fixed-point format (e.g., INT8). The ANN model 502 can be configured to process the inputs (e.g., 510) through layers (e.g., 504, 506, 508) of the ANN model 502 to perform inference tasks. For example, the ANN model 502 can be configured to perform tasks such as image classification, object detection, location and mapping, speech recognition, or other inference tasks.

[0083] At box 606, a model update for the ANN model is selectively computed using either a floating-point processor or a fixed-point processor based on the loss of the ANN model. This model update is computed in either floating-point or fixed-point format. For example, see reference... Figure 5 As described, the ANN model 502 can be selectively trained during backpropagation using either a floating-point processor 516 or a fixed-point processor 514. That is, the computation of weight updates can be performed using either a floating-point processor 516 in floating-point format (e.g., FP16 or FP32 format) or a fixed-point processor 514, which may include (but is not limited to) a GPU (e.g., 104) or an NPU (e.g., 108).

[0084] The computation of ANN model 502 can be used to compute the model loss during forward propagation (e.g., concurrently with inference computation). For example, in one example, the output at output layer 508 can be compared with a ground truth example to determine the difference, which can be considered the model loss. Based on this loss, example architecture 500 can choose to train ANN model 502 using either fixed-point processor 514 in fixed-point format or floating-point processor 516 in floating-point format. For example, example architecture 500 can compare the model loss with one or more predefined thresholds. For example, if the model loss is greater than a predefined threshold (e.g., 0.10 or greater), ANN model 502 can be trained using floating-point processor 516 in floating-point format. On the other hand, if the model loss is less than a predefined threshold (e.g., 0.01 or less than 0.05), ANN model 502 can be trained using fixed-point processor 514. In some aspects, the computation for inference of ANN model 502 using fixed-point processor 514 can be performed concurrently with the computation using floating-point processor 516 to determine weight updates during backpropagation.

[0085] Specific implementation examples are described in the following numbered clauses: 1. A processor-implemented method executed by at least one processor, the processor-implemented method comprising: Input is received through an artificial neural network (ANN) model; The input is processed using a fixed-point processor through the ANN model, and the processing of the input is performed in a fixed-point format to compute inference; and Based on the loss of the ANN model, a floating-point processor or the fixed-point processor is used to selectively compute model updates for the ANN model, the model updates being computed in floating-point or fixed-point format.

[0086] 2. The processor-implemented method according to Clause 1, further comprising concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

[0087] 3. The processor-implemented method according to clause 1 or 2, further comprising: The weights of the hidden layers and at least a subset of the intermediate outputs of the ANN model are stored in memory for computation of the inference; and The memory is accessed to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, and the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

[0088] 4. The method implemented by the processor according to any one of Clauses 1 to 3, wherein the ANN model is implemented on an edge device.

[0089] 5. A processor-implemented method according to any one of Clauses 1 to 4, wherein the loss is selectively calculated based on one or more of the computing workload of the edge device, available memory capacity, or power budget.

[0090] 6. A method implemented by a processor according to any one of Clauses 1 to 5, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

[0091] 7. A method implemented by a processor according to any one of Clauses 1 to 6, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

[0092] 8. An apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, is configured to: Input is received through an artificial neural network (ANN) model; The input is processed using a fixed-point processor through the ANN model, and the processing of the input is performed in a fixed-point format to compute inference; and Based on the loss of the ANN model, a floating-point processor or the fixed-point processor is used to selectively compute model updates for the ANN model, the model updates being computed in floating-point or fixed-point format.

[0093] 9. The apparatus according to Clause 8, wherein the at least one processor is further configured to: concurrently use the fixed-point processor to compute the inference and use the floating-point processor to compute the model update.

[0094] 10. The apparatus according to clause 8 or 9, wherein the at least one processor is further configured to: The weights of the hidden layers and at least a subset of the intermediate outputs of the ANN model are stored in memory for computation of the inference; and The memory is accessed to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, and the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

[0095] 11. The apparatus according to any one of clauses 8 to 10, wherein the ANN model is implemented on an edge device.

[0096] 12. The apparatus according to any one of clauses 8 to 11, wherein the at least one processor is further configured to selectively calculate the loss based on one or more of the edge device's computing workload, available memory capacity, or power budget.

[0097] 13. The apparatus according to any one of clauses 8 to 12, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

[0098] 14. The apparatus according to any one of clauses 8 to 13, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

[0099] 15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by one or more processors and comprising: Program code used to receive input through an artificial neural network (ANN) model; Program code for processing the input using a fixed-point processor through the ANN model, wherein the processing of the input is performed in a fixed-point format to compute inference; and Program code for selectively computing model updates for the ANN model using a floating-point processor or the fixed-point processor based on the loss of the ANN model, the model updates being computed in floating-point or fixed-point format.

[0100] 16. The non-transitory computer-readable medium according to Clause 15, wherein the program code includes program code for including concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

[0101] 17. The non-transitory computer-readable medium according to Clause 15 or 16, wherein the program code comprises: Program code for storing at least a subset of the weights of the hidden layers and intermediate outputs of the ANN model in memory for computation of the inference; and Program code for accessing the memory to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, wherein the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

[0102] 18. A non-transitory computer-readable medium according to any one of Clauses 15 to 17, wherein the ANN model is implemented on an edge device.

[0103] 19. A nontransitory computer-readable medium according to any one of Clauses 15 to 18, wherein the program code includes program code for selectively calculating the loss based on one or more of the computing workload of the edge device, available memory capacity, or power budget.

[0104] 20. The non-transitory computer-readable medium according to any one of Clauses 15 to 19, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

[0105] 21. The non-transitory computer-readable medium according to any one of Clauses 15 to 20, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

[0106] 22. An apparatus comprising: A component used to receive input through an artificial neural network (ANN) model; A component for processing the input using a fixed-point processor through the ANN model, wherein the processing of the input is performed in a fixed-point format to compute inference; and A component for selectively computing model updates for the ANN model using a floating-point processor or a fixed-point processor based on the loss of the ANN model, the model updates being computed in floating-point or fixed-point format.

[0107] 23. The apparatus according to Clause 22, further comprising components for concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

[0108] 24. The apparatus according to clause 22 or 23, further comprising: The component used to store at least a subset of the weights of the hidden layers and intermediate outputs of the ANN model in memory for computation of the inference; and A component for accessing the memory to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, wherein the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of the intermediate outputs of the ANN model.

[0109] 25. The apparatus according to any one of clauses 22 to 24, wherein the ANN model is implemented on an edge device.

[0110] 26. The apparatus according to any one of clauses 22 to 25, the apparatus further comprising a component for selectively calculating the loss based on one or more of the computing workload of the edge device, available memory capacity, or power budget.

[0111] 27. The apparatus according to any one of clauses 22 to 26, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

[0112] 28. The apparatus according to any one of clauses 22 to 27, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

[0113] In some aspects, the receiving component, processing component, and / or selective computing component may be a DSP 106, a GPU 104, a program memory associated with the GPU 104, a fully connected layer 362, an NPU 428, and / or a routing connection processing unit 216 configured to perform the described functions. In another configuration, the aforementioned components may be any module or any device configured to perform the functions described by the aforementioned components.

[0114] The various operations of the methods described above can be performed by any suitable component capable of performing the corresponding function. These components may include various hardware and / or software components and / or modules, including but not limited to circuits, application-specific integrated circuits (ASICs), or processors. Generally, in the cases where operations are illustrated in the accompanying drawings, these operations may have corresponding paired components with similar numbering plus functional components.

[0115] As used, the term "determine" encompasses a wide variety of actions. For example, "determine" can include calculation, computation, processing, derivation, research, searching (e.g., looking in a table, database, or other data structure), assertion, etc. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in memory), etc. Furthermore, "determine" can include parsing, selecting, choosing, building, etc.

[0116] As used, the phrase "at least one of the items in the list" refers to any combination of these items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc.

[0117] The various exemplary logic blocks, modules, and circuits described in this disclosure can be implemented or executed using a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic component, discrete hardware component, or any combination thereof designed to perform the described functions. While the general-purpose processor may be a microprocessor, in alternative embodiments, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0118] The steps or algorithms of the methods described in this disclosure may be directly embodied in hardware, a software module executed by a processor, or a combination of both. The software module may reside in any form of storage medium known in the art. Some examples of usable storage media include random access memory (RAM), read-only memory (ROM), flash memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software module may include a single instruction or multiple instructions and may be distributed across several different code segments, across different programs, and across multiple storage media. The storage medium may be coupled to the processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor.

[0119] The disclosed method includes one or more steps or actions for implementing the described method. The steps and / or actions of the method may be interchanged without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of a particular step and / or action may be modified without departing from the scope of the claims.

[0120] The described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system within the device. This processing system may utilize a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnect buses and bridges. The bus can link various circuits together, including processors, machine-readable media, and bus interfaces. The bus interface can be used to connect network adapters, etc., to the processing system via the bus. The network adapter can be used to implement signal processing functions. In some respects, user interfaces (e.g., keypads, displays, mice, joysticks, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, etc., which are well known in the art and will not be described further.

[0121] A processor may be responsible for managing the bus and general-purpose processing, including executing software stored on a machine-readable medium. A processor may be implemented using one or more general-purpose processors and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry that can execute software. Software should be interpreted broadly as instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, a machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, disks, optical disks, hard disks, or any other suitable storage medium, or any combination thereof. A machine-readable medium may be embodied as a computer program product. A computer program product may include packaging material.

[0122] In a hardware implementation, machine-readable media can be part of a processing system separate from the processor. However, as those skilled in the art will readily understand, machine-readable media, or any portion thereof, can be external to the processing system. By way of example, machine-readable media may include transmit lines, carrier waves modulated by data, and / or computer components separate from the device, all accessible to the processor via a bus interface. Alternatively or additionally, machine-readable media, or any portion thereof, may be integrated into the processor, such as in the case of a cache and / or a general-purpose register file. Although the various components discussed may be described as having a specific location, such as local components, they may also be configured in various ways, such as certain components being configured as part of a distributed computing system.

[0123] The processing system may be configured as a general-purpose processing system having one or more microprocessors providing processor functionality and external memory providing at least a portion of machine-readable medium, all of which are linked together with other supporting circuitry via an external bus architecture. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron and nervous system models. As another alternative, the processing system may be implemented using an application-specific integrated circuit (ASIC) having a processor, bus interface, user interface, supporting circuitry, and at least a portion of machine-readable medium integrated on a single chip, or using one or more field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic components, discrete hardware components, or any other suitable circuitry, or any combination of circuitry capable of performing the various functionalities described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality of the processing system depends on the specific application and the overall design constraints imposed on the system as a whole.

[0124] Machine-readable media may include multiple software modules. These software modules include instructions that, when executed by a processor, enable the processing system to perform various functions. Software modules may include send and receive modules. Each software module may reside in a single storage device or be distributed across multiple storage devices. For example, when a triggering event occurs, a software module may be loaded from a hard disk drive into RAM. During the execution of a software module, the processor may load some of the instructions into a cache to improve access speed. One or more cache lines may then be loaded into a general-purpose register file for processor execution. When the functionality of a software module is referred to below, it will be understood that such functionality is implemented by the processor when executing the instructions from that software module. Furthermore, it should be understood that aspects of this disclosure result in improvements to the functionality of a processor, computer, machine, or other system implementing such aspects.

[0125] If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted through a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one location to another. A storage medium can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and is accessible to a computer. Additionally, any connection is also appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, optical fiber, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared (IR), radio, and microwave, then such coaxial cable, optical fiber, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. The disks and optical discs used include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs. ® Optical discs, where magnetic disks typically reproduce data magnetically, and optical discs reproduce data optically using lasers. Therefore, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Furthermore, in other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.

[0126] Therefore, certain aspects may include a computer program product for performing the presented operations. For example, such a computer program product may include a computer-readable medium on which instructions are stored (and / or encoded) that can be executed by one or more processors to perform the described operations. In some aspects, the computer program product may include packaging material.

[0127] Furthermore, it should be understood that modules and / or other suitable components for performing the described methods and techniques may be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such devices can be coupled to a server to facilitate the delivery of components for performing the described methods. Alternatively, the various methods described can be provided via storage components (e.g., RAM, ROM, physical storage media such as CDs or floppy disks) so that the user terminal and / or base station can obtain the various methods once the storage component is coupled to or provided to the device. In addition, any other suitable techniques suitable for providing the described methods and techniques to the device may be utilized.

[0128] It should be understood that the claims are not limited to the precise configurations and components illustrated above. Various modifications, variations, and alterations may be made to the arrangement, operation, and details of the methods and apparatus described above without departing from the scope of the claims.

Claims

1. A processor-implemented method executed by at least one processor, the processor-implemented method comprising: Input is received through an artificial neural network (ANN) model; The input is processed using a fixed-point processor through the ANN model, and the processing of the input is performed in a fixed-point format to compute inferences; as well as Based on the loss of the ANN model, a floating-point processor or the fixed-point processor is used to selectively compute model updates for the ANN model, the model updates being computed in floating-point or fixed-point format.

2. The processor-implemented method of claim 1, further comprising concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

3. The processor-implemented method according to claim 1, further comprising: The weights of the hidden layers and at least a subset of the intermediate outputs of the ANN model are stored in memory for computation of the inference; as well as The memory is accessed to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, and the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

4. The processor-implemented method according to claim 1, wherein the ANN model is implemented on an edge device.

5. The processor-implemented method of claim 1, wherein the loss is selectively calculated based on one or more of the edge device's computational workload, available memory capacity, or power budget.

6. The method implemented by the processor according to claim 1, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

7. The processor-implemented method according to claim 1, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

8. An apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, is configured to: Input is received through an artificial neural network (ANN) model; The input is processed using a fixed-point processor through the ANN model, and the processing of the input is performed in a fixed-point format to compute inferences; as well as Based on the loss of the ANN model, a floating-point processor or the fixed-point processor is used to selectively compute model updates for the ANN model, the model updates being computed in floating-point or fixed-point format.

9. The apparatus of claim 8, wherein the at least one processor is further configured to: concurrently use the fixed-point processor to compute the inference and use the floating-point processor to compute the model update.

10. The apparatus of claim 8, wherein the at least one processor is further configured to: The weights of the hidden layers and at least a subset of the intermediate outputs of the ANN model are stored in memory for computation of the inference; and The memory is accessed to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, and the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

11. The apparatus of claim 8, wherein the ANN model is implemented on an edge device.

12. The apparatus of claim 8, wherein the at least one processor is further configured to selectively calculate the loss based on one or more of the edge device's computing workload, available memory capacity, or power budget.

13. The apparatus of claim 8, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

14. The apparatus of claim 8, wherein the floating-point format includes a 32-bit floating-point (FP32) or a 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or a 4-bit integer (INT4) format.

15. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by one or more processors and comprising: Program code used to receive input through an artificial neural network (ANN) model; Program code for processing the input using a fixed-point processor through the ANN model, wherein the processing of the input is performed in a fixed-point format to compute inference; and Program code for selectively computing model updates for the ANN model using a floating-point processor or the fixed-point processor based on the loss of the ANN model, the model updates being computed in floating-point or fixed-point format.

16. The non-transitory computer-readable medium of claim 15, wherein the program code includes program code for concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

17. The non-transitory computer-readable medium of claim 15, wherein the program code comprises: Program code for storing at least a subset of the weights of the hidden layers and intermediate outputs of the ANN model in memory for computing the inference; and Program code for accessing the memory to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, wherein the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of intermediate outputs of the ANN model.

18. The non-transitory computer-readable medium of claim 15, wherein the ANN model is implemented on an edge device.

19. The non-transitory computer-readable medium of claim 15, wherein the program code includes program code for selectively calculating the loss based on one or more of the computing workload of the edge device, available memory capacity, or power budget.

20. The non-transitory computer-readable medium of claim 15, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

21. The non-transitory computer-readable medium of claim 15, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.

22. An apparatus comprising: A component used to receive input through an artificial neural network (ANN) model; A component for processing the input using a fixed-point processor through the ANN model, wherein the processing of the input is performed in a fixed-point format to compute inference; and A component for selectively computing model updates for the ANN model using a floating-point processor or a fixed-point processor based on the loss of the ANN model, the model updates being computed in floating-point or fixed-point format.

23. The apparatus of claim 22, further comprising means for concurrently using the fixed-point processor to compute the inference and using the floating-point processor to compute the model update.

24. The apparatus of claim 22, further comprising: A component for storing at least a subset of the weights of the hidden layers and intermediate outputs of the ANN model in memory for computation of the inference; and A component for accessing the memory to retrieve at least one subset of the weights of the hidden layer and intermediate outputs of the ANN model, wherein the floating-point processor computes the model update based on the weights of the hidden layer and at least one subset of the intermediate outputs of the ANN model.

25. The apparatus of claim 22, wherein the ANN model is implemented on an edge device.

26. The apparatus of claim 22, further comprising a component for selectively calculating the loss based on one or more of the edge device's computing workload, available memory capacity, or power budget.

27. The apparatus of claim 22, wherein the floating-point processor includes a graphics processing unit or a neural processing unit, and the fixed-point processor includes a central processing unit or a digital signal processor.

28. The apparatus of claim 22, wherein the floating-point format includes a 32-bit floating-point (FP32) or 16-bit floating-point (FP16) format, and the fixed-point format includes an 8-bit integer (INT8) or 4-bit integer (INT4) format.