Enhancing quantization accuracy with domain-specific priors

WO2026206372A1PCT designated stage Publication Date: 2026-10-01QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040983
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-08-06
Publication Date
2026-10-01

Smart Images

  • Figure US2025040983_01102026_PF_FP_ABST
    Figure US2025040983_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and techniques are described herein for improved quantization of machine learning parameters for use in systems (e.g., Advanced Driver-Assistance Systems (ADAS) and / or other types of systems). For example, a computing device can process first data using a first layer of a machine learning model to obtain output activations. The first layer includes quantized weights. The computing device can determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values. The computing device can scale, based on the usable range of input values, the output activations to obtain scaled output activations. The computing device can quantize the scaled output activations to obtain quantized activations.
Need to check novelty before this filing date? Find Prior Art

Description

PATENTQualcomm Docket No 2503139WO1ENHANCING QUANTIZATION ACCURACY WITH DOMAIN-SPECIFIC PRIORSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims the benefit of Indian Patent Application No.202541029922, filed on March 28, 2025, the contents of which is incorporated herein by reference in its entirety for all purposes.FIELD

[0002] The disclosure relates generally to systems and techniques for improved quantization of machine learning parameters for use in systems, such as for use in Advanced Driver- Assistance Systems (ADAS).BACKGROUND

[0003] Machine learning systems (or models), such as neural networks (e.g., deep neural networks) are widely used for numerous applications, such as generative operations (e.g., to generate images, language / text outputs, etc.), object detection, object classification, object tracking, big data analysis, among others. For example, convolutional neural networks (CNNs) are able to extract high-level features, such as facial shapes, from an input image, and use these high-level features to output a probability that, for example, an input image includes a particular object.

[0004] A generative machine learning system can process data to generate desired output content from an input (e.g., a natural language input, input image(s) or video(s), a noise input such as for diffusion models, etc.). For instance, a language-based generative machine learning model (e.g., a large language model (LLM)) can generate natural language responses from natural language inputs and can incorporate various forms of data, such as audio, images and text. In some cases, a generative machine learning system can include an encoder that processes input features to generate output features (also referred to as embeddings or encodings) and can use the output features to generate relevant output content in a given form (e.g.. in a natural language form, as generated images and / or video, as audio, explanation of the content, etc.).Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO2

[0005] A generative machine learning system can perform a wide range of tasks, such as answering questions, providing explanations, generating creative content, assisting with coding, and offering recommendations. Various tools may be used in conjunction with the generative ML system to allow interaction with external systems, such as browsing the Internet, generating images, executing code, etc. Generative ML systems are designed to assist users in solving problems, learning new information, and enhancing productivity.SUMMARY

[0006] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summan be considered to identity7key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.

[0007] In some aspects, an apparatus for generating content is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: process first data using a first layer of a machine learning model to obtain output activations, wherein the first layer includes quantized weights; determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scale, based on the usable range of input values, the output activations to obtain scaled output activations; and quantize the scaled output activations to obtain quantized activations.

[0008] In some aspects, a method of generating content is provided. The method includes: processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer includes quantized weights; determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scaling, based on the usable range of input values, the output activations to obtain scaled output activations; and quantizing the scaled output activations to obtain quantized activations.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO3

[0009] In some aspects, a non-transitory computer-readable medium is provided having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to: process first data using a first layer of a machine learning model to obtain output activations, wherein the first layer includes quantized weights; determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scale, based on the usable range of input values, the output activations to obtain scaled output activations; and quantize the scaled output activations to obtain quantized activations.

[0010] In some aspects, an apparatus for generating content is provided. The apparatus includes: means for processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer includes quantized weights; means for determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values; means for scaling, based on the usable range of input values, the output activations to obtain scaled output activations; and means for quantizing the scaled output activations to obtain quantized activations.

[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0012] The foregoing, together with other features and embodiments, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Illustrative embodiments of the present application are described in detail below with reference to the following figures:

[0014] FIG. 1 illustrates an example implementation of a system-on-a-chip (SOC), in accordance with some examples;Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO4

[0015] FIG. 2 is a block diagram illustrating an example of a deep learning neural network, according to some aspects of the present disclosure;

[0016] FIG. 3 is a block diagram illustrating an example of a convolutional neural network (CNN), according to various aspects of the present disclosure;

[0017] FIG. 4 is a block diagram illustrating an example of a deep convolutional network, in accordance with aspects of the present disclosure;

[0018] FIG. 5 illustrates a technique for generating BEV features, in accordance with aspects of the present disclosure;

[0019] FIG. 6 illustrates a method for quantizing machine learning model parameters, in accordance with aspects of the present disclosure;

[0020] FIG. 7 illustrates a sigmoid input for a machine learning model, in accordance with aspects of the present disclosure;

[0021] FIG. 8 illustrates a range of a sigmoid function, in accordance with aspects of the present disclosure;

[0022] FIG. 9 illustrates an example of logarithmic input clamping, in accordance with aspects of the present disclosure;

[0023] FIG. 10 illustrates a range of an inverse sigmoid function, in accordance with aspects of the present disclosure;

[0024] FIG. 11 illustrates an example of grid-sample input range clamping, in accordance with aspects of the present disclosure;

[0025] FIG. 12 illustrates an example of epsilon handling for a pin-hole camera projection, in accordance with aspects of the present disclosure:

[0026] FIG. 13 is a flow diagram illustrating an example of a process for quantization of parameters of machine learning models, in accordance with aspects of the disclosure;Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO5

[0027] FIG. 14 is a block diagram of an example transformer in accordance with aspects of the present disclosure; and

[0028] FIG. 15 illustrates an example computing device architecture of an example computing device which can implement the various techniques described herein.DETAILED DESCRIPTION

[0029] Certain aspects and embodiments of this disclosure are provided below. Some of these aspects and embodiments may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth to provide a thorough understanding of embodiments of the application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and description are not intended to be restrictive.

[0030] The ensuing description provides example embodiments only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example embodiments will provide those skilled in the art with an enabling description for implementing an example embodiment. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0031] Machine learning models can be trained to perform various functions and / or provide various types of outputs. For instance, some generative machine learning models can provide a conversational interface that uses natural language prompts as inputs, such as text or voice. In some examples, a user can provide an input prompt in natural language to the generative machine learning model, and the generative machine learning model can provide a response in natural language form. The input prompt and the output response can optionally be combined with one or more other ty pes of information or data, such as images or files.

[0032] While machine learning models (e.g., neural networks) are powerful architectures capable of a wide range of useful tasks, such as recognizing objects in image data, they are likewise highly resource dependent. For example, neural networks may require Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO6significant compute, memory, power, and / or time resources for training and / or for inferencing. These resource requirements may significantly limit the ability to train and deploy neural networks to certain types of devices and for certain use cases. For instance, training of machine learning models may be a computationally intensive process that can take a relatively long time, a large quantity of training data, and many operations. As such, in some cases, quantization may be used to help reduce the computation that is required.

[0033] Quantization is a method of mapping continuous values to a smaller set of discrete finite values. For example, quantization approximates real-world values (e.g., floating point values) with representative values (e.g., integer values) that limit the precision and range of the original input. When applied to machine learning, such as to a neural network, quantization may significantly reduce resource usage for both training and inferencing. For example, performing massive numbers of integer operations during training of or inferencing with a quantized (e.g., reduced-precision) neural network may be significantly more efficient in terms of resource usage as compared to performing floating point operations with an unquantized (e.g., full-precision) neural network processing the same input data.

[0034] Quantization of trained machine learning models allows quantized trained machine learning models to be efficiently deployed on various devices. Quantization may allow a quantized trained machine learning model to perform operations (e.g., at the inference phase of operation) on a device relatively quickly. Quantization may include changing a format of parameters (e.g., weights and activations) of a machine learning model from the format in which the machine learning model was trained to a different format.

[0035] Post-training quantization (PTQ) is a process by which a trained machine learning model is quantized. For example, during training, a machine learning model may store parameters (e.g., weights) according to a first format that may allow a high degree of precision. For example, during training, the model may store parameters as floating-point numbers, such as 16-bit floating point numbers, which may be referred to as floatl6 or FP16, or 32-bit floating point numbers, which may be referred to as float32 or FP32. Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO7According to PTQ, after training is performed to generate the trained machine learning model and before the trained machine learning model is deployed onto a device for use by a device, the model may be quantized by, in part, changing the format used to store the parameters of the model from the first format to a second format (e.g., an integer format where the parameters are represented as integer numbers instead of floating point numbers). The second format may use less memory and / or be less computational expensive to use. For example, after quantization, the model may store parameters as integer numbers such as 16-bit integer numbers (which may be referred to as Inti 6), 8-bit integer numbers (which may be referred to as Int8), 4-bit integer numbers (which may be referred to as Int4), etc. It may be less computationally expensive to store and / or operate using integer numbers than floating point numbers. Thus, a device may conserve power and / or processing time when using a quantized trained machine learning model as compared with using an unquantized trained machine learning model. Accordingly, it may be advantageous to quantize machine learning models.

[0036] In various aspects of the present disclosure, the term "‘quantize.” “quantizing,” and like terms may be used as a verb and may be applied to machine learning models, and / or layers of a trained machine learning model. In such cases the term “quantize,” “quantizing,” and like terms may refer to quantizing parameters (e.g., weights and / or activations) as stored in the of the respective machine learning models and / or layers. For example, atained machine learning model may include weights (e.g., numerical values). During training, and after training, the weights may be stored in the trained machine learning model in a first format (e.g., float8, floatl6, float32, floatl6). Quantizing the trained machine learning model may include changing the weights from the first format to a second format (e.g., Int8).

[0037] Further, machine learning models may include activations. Activations can refer to values used as inputs and / or outputs of various functions, operations, and / or layers (or blocks or modules) of the machine learning model. The activations may have a certain format. For example, a function (of a particular layer) may expect to receive values formatted according to the certain format. For instance, the function may read the values from memory according to the certain format. Further, the function may output other values (e.g., based on processing the values read from memory by the function) according Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO8to a particular format (which may or may not be the same as the certain format). For instance, the function may write the other values to memory' in the particular format. For example, an unquantized trained machine learning model may use floatl6 to pass values between functions (e.g., outputting values from one function and reading the values by another function). In some cases, the activations of the machine learning model can also be quantized.

[0038] Quantizing numbers can result in a loss of precision. Quantizing parameters or weights of a trained machine learning model can result in degradation of the model. In the present disclosure, the term “error” is used to describe a degree of difference between outputs of an unquantized trained machine learning model and the trained machine learning model after being quantized (given the same inputs).

[0039] Not all elements of a neural network are equally resilient to quantization. That is, certain elements of a neural network may be more sensitive to quantization than others. Consequently, quantizing an entire neural network to a uniform bit-width (referred to a fixed-precision quantization (FPQ)) may result in reduced resource usage, but may also reduce task performance to an unacceptable level. To resolve this issue, mixed-precision quantization (MPQ) seeks to quantize different elements of a neural network at vary ing quantization rates, which results in reduced resource usage while maintaining task performance. The various elements of a neural network that may be quantized with MPQ include, for example, individual layers, groups of layers, sub-layers, weight channels, and others. Generally, MPQ can achieve higher accuracy for the same computational budget because it allocates higher precision to elements (e g., layers) that are more sensitive to quantization and reduces bitwidth for elements more robust to quantization.

[0040] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for improved quantization relating to machine learning models (also referred to as machine learning systems). According to some aspects, the systems and techniques can perform selective tuning of ranges of inputs (e.g., activations output from previous outputs) for various layers. For instance, some systems or layers may suffer significant accuracy loss when parameters are quantized with a certain precision (e.g., 8-bit precision Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO9in INT8). With selective tuning of ranges of inputs for such layers, the systems and techniques can mitigate against such a loss in accuracy and can thus enhance performance.

[0041] The systems and techniques described herein may be employed in a variety of devices or systems. Examples of devices and systems include, but are not limited to, vehicles or devices or systems of vehicles (e.g., Advanced Driver-Assistance Systems (ADAS), etc.), mobile robots, mobile devices such as mobile phones, extended reality (XR) devices (e.g.. augmented reality (AR) devices, virtual reality (VR) devices, and / or mixed reality (MR) devices), or other devices or systems. Such devices may include multiple sensors to gather information about an environment, as well as processing systems to process the sensor information for various purposes, such as route planning, navigation, collision avoidance, etc. In some aspects, the systems and techniques may be applied to generate a three-dimensional (3D) bird’s eye view (BEV) (e.g., a top-down view) multimodal feature maps of an environment.

[0042] Various aspects of the disclosure are discussed in detail below. While specific implementations are discussed, it should be understood that such implementations are described for illustration purposes only. A person skilled in the relevant art will recognize that other components and configurations may be used without parting from the spirit and scope of the disclosure.

[0043] FIG. 1 illustrates an example implementation of a system-on-a-chip (SOC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computational device (e.g., neural network with weights), delays, frequency bin information, task information, among other information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with a CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO10

[0044] The SOC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity. USB connectivity. Bluetooth connectivity, and the like, and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, theNPU is implemented in the CPU 102, DSP 106, and / or GPU 104. The SOC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and / or navigation module 120, which may include a global positioning system.

[0045] The SOC 100 may be based on an ARM instruction set. SOC 100 and / or components thereof may be configured to perform segmentation mask extrapolation. For example, the CPU 102, DSP 106, and / or GPU 104 may be configured to perform object detection using a visual language model via latent feature adaptation with synthetic data.

[0046] In some cases, the SOC 100 may process data using neural networks and / or machine learning (ML) systems. A neural network is an example of an ML system, and a neural network can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes of the input layer, processing is performed by hidden nodes of the one or more hidden layers, and an output is produced through output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of the neural network can include feature maps or activation maps that can include artificial neurons (or nodes). A feature map can include a filter, a kernel, or the like. The nodes can include one or more weights used to indicate an importance of the nodes of one or more of the layers. In some cases, a deep learning network can have a series of many hidden layers, with early layers being used to determine simple and low-level characteristics of an input, and later layers building up a hierarchy of more complex and abstract characteristics.

[0047] FIG. 2 is an illustrative example of a neural network 200 (e.g., a deep-learning neural network) that can be used to implement machine-leaming-based image generation, feature segmentation, implicit-neural-representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO11and / or automation. The neural network 200 may be run, for example, on one or more processors, such as CPU 102 of FIG. 1, GPU 104 of FIG. 1, DSP 106 of FIG. 1,NPU 108 of FIG. 1, etc.

[0048] An input layer 202 includes input data. Neural network 200 includes multiple hidden layers hidden layers 206a, 206b, through 206n. The hidden layers 206a, 206b, through hidden layer 206n include “n"’ number of hidden layers, where “n"’ is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. Neural network 200 further includes an output layer 204 that provides an output resulting from the processing performed by the hidden layers 206a, 206b, through 206n.

[0049] Neural network 200 may be, or may include, a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, neural network 200 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some cases, neural network 200 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.

[0050] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of input layer 202 can activate a set of nodes in the first hidden layer 206a. For example, as shown, each of the input nodes of input layer 202 is connected to each of the nodes of the first hidden layer 206a. The nodes of first hidden layer 206a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 206b. which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 206b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 206n can activate one or more nodes of the output layer 204, at which an output is provided. In some cases, while nodes (e.g.. node 208) in Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO12neural network 200 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.

[0051] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of neural network 200. Once neural network 200 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing neural network 200 to be adaptive to inputs and able to learn as more and more data is processed.

[0052] Neural network 200 may be pre-trained to process the features from the data in the input layer 202 using the different hidden layers 206a, 206b, through 206n to provide the output through the output layer 204. In an example in which neural network 200 is used to identify features in images, neural network 200 can be trained using training data that includes both images and labels, as described above. For instance, training images can be input into the network, with each training image having a label indicating the features in the images (for the feature-segmentation machine-learning system) or a label indicating classes of an activity in each image. In one example using object classification for illustrative purposes, a training image can include an image of a number 2. in which case the label for the image can be [00 1 000 000 0] .

[0053] In some cases, neural network 200 can adjust the weights of the nodes using a training process called backpropagation. As noted above, a backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backw ard pass, and parameter update is performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until neural network 200 is trained well enough so that the weights of the layers are accurately tuned.

[0054] For the example of identitying objects in images, the forward pass can include passing a training image through neural network 200. The weights are initiallyPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO13randomized before neural network 200 is trained. As an illustrative example, an image can include an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like).

[0055] As noted above, for a first training iteration for neural network 200, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1). With the initial weights, neural network 200 is unable to determine low-level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a cross-entropy loss. Another example of a loss function includes the mean squared error (MSE), defined< value of Etotal.

[0056] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. Neural network 200 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW, where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted asw~Wi”where w denotes a weight, wi denotes the initial weight, and q denotes a learning rate. The learning rate can be set Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO14to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.

[0057] Neural network 200 can include any suitable deep network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. Neural network 200 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs), a Recurrent Neural Networks (RNNs), among others.

[0058] FIG. 3 is an illustrative example of a convolutional neural network (CNN) 300. The input layer 302 of the CNN 300 includes data representing an image or frame. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like). The image can be passed through a convolutional hidden layer 304, an optional nonlinear activation layer, a pooling hidden layer 306, and fully connected layer 308 (which fully connected layer 308 can be hidden) to get an output at the output layer 310. While only one of each hidden layer is shown in FIG. 3, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 300. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.

[0059] The first layer of the CNN 300 can be the convolutional hidden layer 304. The convolutional hidden layer 304 can analyze image data of the input layer 302. Each node of the convolutional hidden layer 304 is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 304 can be considered as one or more filters (each filter corresponding to a different activation or feature map), with each convolutional iteration of a filter being a node or neuron of the Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO15convolutional hidden layer 304. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28x28 array, and each filter (and corresponding receptive field) is a 5x5 array, then there will be 24x24 nodes in the convolutional hidden layer 304. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptive field in the input image. Each node of the convolutional hidden layer 304 will have the same weights and bias (called a shared weight and a shared bias). For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for an image frame example (according to three color components of the input image). An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.

[0060] The convolutional nature of the convolutional hidden layer 304 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 304 can begin in the top-left comer of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 304. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top-left comer of the input image array). The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 304. For example, a filter can be moved by a step amount (referred to as a stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 304.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO16

[0061] The mapping from the input layer to the convolutional hidden layer 304 is referred to as an activation map (or feature map). The activation map includes a value for each node representing the filter results at each location of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a stride of 1) of a 28 x 28 input image. The convolutional hidden layer 304 can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 3 includes three activation maps. Using three activation maps, the convolutional hidden layer 304 can detect three different kinds of features, with each feature being detectable across the entire image.

[0062] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 304. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f(x) = max(0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the non-linear properties of the CNN 300 without affecting the receptive fields of the convolutional hidden layer 304.

[0063] The pooling hidden layer 306 can be applied after the convolutional hidden layer 304 (and after the non-linear hidden layer when used). The pooling hidden layer 306 is used to simplify the information in the output from the convolutional hidden layer 304. For example, the pooling hidden layer 306 can take each activation map output from the convolutional hidden layer 304 and generates a condensed activation map (or feature map) using a pooling function. Max -pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 306, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 304. In the example shown in FIG. 3, three pooling filters are used for the three activation maps in the convolutional hidden layer 304.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO17

[0064] In some examples, max -pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a stride (e.g., equal to a dimension of the filter, such as a stride of 2) to an activation map output from the convolutional hidden layer 304. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2x2 nodes in the previous layer (with each node being a value in the activation map). For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max -pooling filter is applied to an activation filter from the convolutional hidden layer 304 having a dimension of 24x24 nodes, the output from the pooling hidden layer 306 will be an array of 12x12 nodes.

[0065] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2x2 region (or other suitable region) of an activation map (instead of computing the maximum values as is done in max-pooling) and using the computed values as an output.

[0066] The pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image and discards the exact positional information. This can be done without affecting results of the feature detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Maxpooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 300.

[0067] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 306 to every one of the output nodes in the output layer 310. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 304 includes 3 x24x24 hidden feature nodes based on application of a 5 x5 local receptive field Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO18(for the filters) to three activation maps, and the pooling hidden layer 306 includes a layer of 3x12x12 hidden feature nodes based on application of max-pooling filter to 2x2 regions across each of the three feature maps. Extending this example, the output layer 310 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 306 is connected to every’ node of the output layer 310.

[0068] The fully connected layer 308 can obtain the output of the previous pooling hidden layer 306 (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 308 can determine the high-level features that most strongly correlate to a particular class and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 308 and the pooling hidden layer 306 to obtain probabilities for the different classes. For example, if the CNN 300 is being used to predict that an object in an image is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the obj ect, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person).

[0069] In some examples, the output from the output layer 310 can include an M-dimensional vector (in the prior example, M=10). M indicates the number of classes that the CNN 300 has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the M-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that there is a 5% probability that the image is the third class of object (e.g., a dog), an 80% probability that the image is the fourth class of object (e.g., a human), and a 15% probability7that the image is the sixth class of object (e.g., a kangaroo). The probability for a class can be considered a confidence level that the object is part of that class.

[0070] FIG. 4 is a block diagram illustrating an example of a deep convolutional network 450. The deep convolutional network 450 may include multiple different ty pes of layers Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO19based on connectivity and weight sharing. As shown in FIG. 4, the deep convolutional network 450 includes the convolution blocks 454A, 454B. Each of the convolution blocks 454A, 454B may be configured with a convolution layer (CONV) 456, a normalization layer (LNorm) 458, and a max pooling layer (MAX POOL) 460. Of note, the layers illustrated with respect to convolution blocks 454A and 454B are examples of layers that may be included in a convolution layer and are not intended to be limiting and other types of layers may be included in any order.

[0071] The convolution layers 456 may include one or more convolutional filters, which may be applied to the input data 452 to generate a feature map. Although only two convolution blocks 454A, 454B are shown, the present disclosure is not so limiting, and instead, any number of convolution blocks (e.g., convolution blocks 454A, 454B) may be included in the deep convolutional network 450 according to design preference. The normalization layer 458 may normalize the output of the convolution filters. For example, the normalization layer 458 may provide whitening or lateral inhibition. The max pooling layer 460 may provide down sampling aggregation over space for local invariance and dimensionality reduction.

[0072] The parallel filter banks, for example, of a deep convolutional network may be loaded on a processor such as a CPU or GPU, or any other type of processor 1510 discussed with respect to the computing device architecture 1500 of FIG. 15 to achieve high performance and low' power consumption. In alternative aspects, the parallel filter banks may be loaded on a DSP or an ISP of the computing device architecture 1500 of FIG. 15. In addition, the deep convolutional network 450 may access other processing blocks that may be present on the computing device architecture 1500 of FIG. 10, such as sensor processor and navigation module, dedicated, respectively, to sensors and navigation.

[0073] The deep convolutional network 450 may also include one or more fully connected layers, such as layer 462A (labeled “FC1”) and layer 462B (labeled “FC2”). The deep convolutional network 450 may further include a logistic regression (LR) layer 464. Between each layer 456, 458, 460, 462A, 462B, 464 of the deep convolutional network 450 are weights (not shown) that are to be updated. The output of each of the Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO20layers (e.g., 456, 458, 460, 462A, 462B, 464) may serve as an input of a succeeding one of the layers (e.g., 456, 458, 460, 462A, 462B, 464) in the deep convolutional network 450 to learn hierarchical feature representations from input data 452 (e.g., images, audio, video, sensor data and / or other input data) supplied at the first of the convolution blocks 454A. The output of the deep convolutional network 450 is a classification score 466 for the input data 452. The classification score 466 may be a set of probabilities, where each probability is the probability of the input data including a feature from a set of features.

[0074] In some cases, one or more convolutional networks, such as a DCN, may be incorporated into more complex ML networks. As an example, as indicated above, the deep convolutional network 450 may output probabilities that an input data, such as an image, includes certain features. The deep convolutional network 450 may then be modified to extract (e.g., output) certain features. Additionally, DCNs may be added to extract other features as well. This set of DCNs may function as feature extractors to identify features in an image. In some cases, feature extractors may be used as a backbone for additionally ML network components to perform further operations, such as image segmentation.

[0075] In some cases, cross-attention blocks may be used for a variety of applications where complex, non-linear relationships may be learned, such as for language-based tasks. For example, cross-attention blocks may be used in large language models. A crossattention block may be used in a ML model that can leam context about features to try to detect relationships between sequences of in data to focus on relevant information of a sequence of data while processing another sequence of data. Cross-attention blocks may¬ be inserted into other ML models, such as deep convolutional network 450 of FIG. 4, CNN 300 of FIG. 3, neural network 200 of FIG. 2, etc.

[0076] In some cases, one or more convolutional networks, such as a DCN, may be incorporated into more complex ML networks. As an example, as indicated above, the deep convolutional network 450 may output probabilities that an input data, such as an image, includes certain features. The deep convolutional network 450 may then be modified to extract (e.g., output) certain features. Additionally, DCNs may be added to extract other features as well. This set of DCNs may function as feature extractors to Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO21identify features in an image. In some cases, feature extractors may be used as a backbone for additional ML network components to perform further operations, such as localization, image segmentation, object detection, etc.

[0077] In some cases, the extracted features and images may be used to construct a three-dimensional (3D) bird’s eye view (BEV) (e.g., a top-down view) multimodal feature map of an environment. For example, an XR device and / or ADAS system may include a suite of sensors that may sense the environment using different techniques. Multimodal features may be generated based on data from multiple different types of sensors, such as an image sensor along with at least one other type of sensor, such as a LIDAR, RADAR, SODAR, SONAR, etc. Using different sensor types helps provide a more holistic understanding of the environment, increases robustness against failure and / or noise from a single sensor modality, and may help overcome occlusions. In some cases, a sensor type of a sensor may be based on how the sensor senses the environment. For example, two sensors which sense different parts of the electromagnetic spectrum may have different sensor types. Similarly, a sensor which senses reflection / refraction of projected light may have a different sensor type from another sensor which senses natural reflected / refracted light. In some cases, the different sensors may sense the environment in three dimensions. The multimodal features may be transformed into 3D BEV features to help provide a viewpoint invariant representation that encodes semantic information about the environment. In some cases, the 3D BEV features may be normalized based on sensor configuration to help enable generalizability of the multimodal 3D BEV features across systems with different sensors.

[0078] FIG. 5 illustrates a technique 500 for generating BEV features, in accordance with aspects of the present disclosure. In technique 500, input camera data 502 (e.g., images captured by cameras) may be input to a camera data encoder 504. The camera data encoder 504 may include one or feature extractors. These feature extractors may be ML based and be used to identify certain features in the camera data. As an example, the feature extractors may include one or more layers or transformer blocks which may include feature maps for recognizing certain features. The camera data encoder 504 may output the identified features as intermediate camera features 506. Of note, the input camera data 502 and camera data encoder 504 may operate in a 2D space (e.g., on a height Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO22and width axes with respect to the camera). A perspective transformation may be applied to the output intermediate camera features 508 which converts the intermediate camera features from, for example, a frontal view of an environment from a vehicle to BEV projected camera features 508 as if features were generated based on a camera positioned above the vehicle. In some cases, the perspective transformation may be ML based.

[0079] In some cases, the intermediate camera features 506 may be input to a 2D object detector 520 for object detection. In some cases, the 2D object detector 520 may be ML based and the 2D object detector may locate and identify objects that are imaged in the input camera data 502. For example, the 2D obj ect detector 520 may detect and / or identify7objects and place a bounding box around detected objects. The 2D object detector 520 may output bounding box location information indicating where the bounding boxes for detected objects are located in the images. Any 2D object detection model by the 2D object detector 520 may be used and examples of 2D object detectors 520 may include scale invariant feature transform (SIFT), you only look once (YOLO), single shot multibox detector (SSD), and the like. Of note, while shown using intermediate camera features 506 as input, it should be understood that the 2D object detector 520 may be trained to detect objects using, as input, the BEV projected camera features 508. In some cases, image segmentation may also be used.

[0080] In some cases, lidar data 510 may be received, for example as a lidar point cloud, captured by a lidar. Lidar may transmit a beam of ultraviolet, visible, or near infrared light into an environment and detects reflections of the beam from objects in the environment. Based on an amount of time needed for the reflections to be detected, distances to objects in the environment may be determined and lidar points may be described based on the point’s location on a width, height, and depth axes with respect to the lidar. Thus, the lidar data is three-dimensional data. The lidar data 510 may be input to a lidar data encoder 512. The lidar data encoder 512 may be similar to the camera data encoder 504, but configured (e.g., trained) to operate in a 3D space to identify features in the lidar data and output the identified features as intermediate lidar features 514. The intermediate lidar features 514 may then be flattened 516 to BEV projected lidar features, for example, by removing or averaging the height information (e.g., height axes, height channel, height dimension).Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO23

[0081] In some cases, the intermediate lidar features 514 (e.g., 3D features) may be input to a 3D object detector 522 for object detection. In some cases, the 3D object detector 522 may be ML based on the 3D object detector 522 may locate and identify objects using the intermediate lidar features 514. Any 3D object detector 522 may be used and examples of 3D object detectors 522 may include PointVoxel-RCNN (PV-RCNN), VoxelNet, PointNet, PointNet++, etc. Of note, while shown using intermediate lidar features 514 as input, the 3D object detector 522 may be configured to detect objects using, as input, the flattened 516 BEV projected lidar features. For example, the 3D object detector 522 may detect and / or identify objects based on the intermediate lidar features 514 and place a bounding box around detected objects along with label(s) identifying the detected object. The 3D object detector 522 may output 3D bounding box location information indicating where the bounding boxes for detected objects are located in the 3D space described by the lidar data 510 along with the label(s).

[0082] In some cases, the output 3D bounding box location information and labels may be input to a 2D to 3D projector 524. In some cases, the intermediate lidar features 514 may also be input to the 2D to 3D projector 524, along with sensor calibration information. The 2D to 3D projector 524 may project the 3D bounding box(es) for detected object(s) into 2D from 3D to generate projected 2D bounding box information for output. The projection may be based on the sensor calibration information. For example, the sensor calibration information may include information such as camera intrinsic information (e.g., focal length, principal point, etc.), pose information for the sensors, and the like. The 2D to 3D projector 524 may determine transformation parameters, such as a rotation, translation, etc. based on the pose information (e.g., difference between a pose of a camera and a lidar) and along with the camera intrinsic information, and information about the lidar sensor to project the 3D bounding box location from 3D to 2D. In some cases, this projected 2D bounding box information may include location information indicating where the projected 2D bounding boxes for detected objects in the intermediate lidar features 514 would be located relative to an image of the input camera data 502.

[0083] In some cases, the projected 2D bounding box(es) from the projected 2D bounding box information may be aligned 526 with the bounding boxes for detected objects from Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO24the 2D object detector 520. Aligning the bounding boxes from 2D object detector 520 and 3D object detector 522 may be useful to assure accuracy in perception and decision making. In some cases, reliable alignments may rely on precise sensor calibrations (e.g., in the sensor calibration information) to ensure that data from the different sensors line up. In some cases, calibrating the sensors may be performed as a part of manufacturing a device.

[0084] In real world applications, maintaining a precise sensor calibration for multiple sensors on a device can be challenging. For example, devices with sensors may be vulnerable to sensor / calibration drift, sensor inaccuracy, noise, physically relocated sensors, and the like, which may result in unreliable 3D to 2D bounding box alignments between a projected 2D bounding box and a bounding box from 2D object detection, especially in complex environments. Thus, techniques for dynamically updating calibrations as between sensors may be useful.

[0085] In some cases, particle filtering for dynamic camera calibration optimization and variance-aware cross-attention may be useful for handling uncertainty when performing 3D to 2D bounding box alignments.

[0086] As indicated above, 3D bounding boxes may be transformed to projected 2D bounding boxes using a set of transformation parameters (e.g., rotation , translation / . etc.) and intrinsic parameters K (e.g., focal length principal point ex, cy, etc ). The transformation parameters T may represent a transformation that may be applied to a 3D bounding box (e.g., from the lidar data) to generate a corresponding projected 2D bounding box such that T= [7?,t], where R may be rotation matrix, such as a 3x3 rotation matrix, and t may be a translation vector, such as a 3x1 translation vector. In some cases, the intrinsic parameters may be defined as K = fx, fy, cx, cy.

[0087] In some cases, a set of A particles may be initialized. In some cases, each particle of the set of A particles may be associated with a set of transformation parameters T and intrinsic parameters K (e.g., set of projection parameters). The set of projection parameters for each particle may be initialized with random parameter values. In some cases, the initialized random parameter values may be within certain reasonable bounds.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO25Weights may be assigned to each particle. In some cases, equal weights, such as a default weight) may be assigned to each particle.

[0088] As noted previously, the systems and techniques described herein relate to improved approaches for performing quantization of machine learning model parameters, such as activations. Such techniques reduce a need for computational resources, while minimizing performance degradation due to error or noise resulting from the quantization.

[0089] Quantization is a technique used to reduce the precision of the parameters in artificial intelligence models, such as weights and activations, from high-precision (e.g., 32-bit floating point) to lower-precision (e.g., 8-bit integer) formats. This reduction in precision significantly decreases the model size, leading to faster inference times, reduced memory usage, and lower energy consumption, which is crucial for deploying Al models on edge devices with limited resources. Techniques like mixed precision quantization, where some layers use higher precision while others use lower precision, ensure minimal accuracy drop while retaining most of the latency reduction gains. This approach is particularly beneficial for complex models like vision transformers and deformable DEtection TRansformer (DETR), which present numerous challenges for quantization.

[0090] When a scaling prior to quantization is too high, the quantized data may become very sparse and lot of noise is introduced. In the case of a Softmax function, a large scale may be mathematically redundant. Accordingly, using the mathematical property of Sigmoid, disclosed solutions restrict the scale of the input (by limiting input range) so that the representation of quantized data is much richer and is not sparse. Richer data means better representation, therefore better accuracy. Further, in general, the activations of a model are more challenging to quantize in a network than the weights. Weights are naturally limited in their range, and do not generally present many problems in quantization as they are well-behaved.

[0091] FIG. 6 illustrates a method 600 for quantizing machine learning model parameters, in accordance with aspects of the present disclosure. Method 600 depicts an approach to quantization, specifically, a selective tuning of ranges of inputs prior toPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO26quantization. Method 600 includes various blocks 602-620, each of which represents one or more operations.

[0092] At block 602, a decision is made as to whether to optimize (e.g., tune a range of) parameters of a machine learning model. If the decision is no C'N”), then method 600 proceeds to block 606, and operations occur in 32-bit floating point. If the decision is yes (“Y”), then method 600 proceeds to block 604.

[0093] At block 604, a decision is made as to whether to limit precision to 16-bit floating point numbers. If the decision is yes (“Y”), then method 600 proceeds to block 608, and quantization to 16-bit floating point is applied. If the decision is no (“N"’), then method 600 proceeds to block 610.

[0094] At block 610, a decision is made as to whether a representative dataset is available. A representative dataset is used to perform forward pass of all the operators in the model to collect data ranges of all tensors. Multiple algorithms are available to identify this data range, such as a histogram and a minimum / maximum across all inferences.

[0095] If such a dataset is not available (“N”), then method 600 proceeds to block 616. At block 616, quantization is applied. Then, method 600 proceeds to block 620. At block 620, a mixture of fixed point and floating point is used.

[0096] By contrast, at block 610, if a representative dataset is available (“Y”), then method 600 proceeds to block 612. At block 612, the data range for each activation is calibrated. Method 600 proceeds to block 614, where parameters and activations are quantized. Method 600 then proceeds to block 618, at which fixed point values are used in the model.

[0097] As discussed, quantization provides various advantages. For instance, some deep learning hardware accelerators (DL HWA) are optimized for 8-bit integer (INT8) computation. These accelerators can perform INT8 operations much faster and more efficiently compared to 16-bit floating point (FP16) or 16-bit integer (INTI 6) operations. This makes INT8 quantization particularly advantageous for deploying vision models onPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO27such hardware, as it maximizes the computational efficiency and throughput of the accelerators.

[0098] During quantization, floating point values are mapped to an N-bit (4 / 8 / 16) quantization space of the form:val_fp32 = scale * (y al quantized — zero_point), where scale is a positive real number used to map the floating-point numbers to a quantization space.

[0099] Scale may be calculated as follows:scale = (data_range_max — data_range_miri) I (quantization j~ange jnax — quantizations angejnin), where zero _point represents zero in the quantization space.

[0100] As discussed below the resulting quantized floating point zero value should be representable in a quantization space. This is because zero padding is used in many CNNs. If it is not possible to represent zero uniquely after quantization accuracy may be reduced.

[0101] FIG. 7 illustrates a sigmoid input for a machine learning model, in accordance with aspects of the present disclosure. Sigmoid input 700 includes addition operator 702 and sigmoid function 704. The range of the sigmoid function is illustrated further with respect to FIG. 8.

[0102] The sigmoid function 704 may be used to adjust a range of the data prior to quantization. As discussed, data range computation is helpful for quantizing activations in model inferencing. As depicted, the addition operator 702 receives an input and provides an output value to the sigmoid function 704, which provides an adjusted value for use in quantization.

[0103] Actual tensor statistics-based range selection works for most of the ML operators like convolution, matrix multiplications etc. But for some specific operators, these algorithms may fail and result in low er accuracy. For example, if an actual data range ofPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO28addition operator 702 activation in the system illustrated in FIG. 7 is [-512, 512], a scale of -4.00 would be needed to represent this tensor in quantized 8-bit integer format: scale = (data_range_max — data_range_min) I quantization_range _max — quantization_range _min)

[0104] In an example, scale is calculated as follows: scale = (512- (-512)) / (127 -(-128)) = 4.01. In this scale example, a result from the sigmoid function 704 would have only five unique values in the output and everything else would be quantized to one of these values having huge quantization loss as follows:-16 -> 0, -8 -> 0, -4 -> 0.017, 0 -> 0.5, 4 -> 0.98, 8 -> 1.0, 16-> 1.0

[0105] To compensate for this issue, some algorithms switch to 16-bit quantization for such operators.

[0106] FIG. 8 illustrates a range 800 of a sigmoid function, in accordance with aspects of the present disclosure. FIG. 8 depicts the sigmoid function referred to in FIG. 7. As depicted, the range 800 of the sigmoid function includes only an effective range [-6, 6] of values in each input tensor (float) will produce unique outputs (integer) within the available output range. As can be seen, beyond this range, all inputs will result in saturated outputs.

[0107] Accordingly, instead of selecting a range based on actual activation data range, disclosed techniques clamp the input of sigmoid. This would result utilizing complete tensor range in sigmoid output as follows:data_range_max = min(6, data_range_max) data_range_min = max(— 6, data_range_min) scale = (data_range_max — data_range_min) / (quantization_range_max — quantization_range _min)scale = (6 - (-6)) / (127 - (-128)) = 0.0470Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO29

[0108] With this scale, sigmoid would have all 255 unique values in the output for the given input range, resulting in the lowest possible quantization error. This form of logical range clamping can be applied to many machine learning operators. Examples of such operators include, but are not limited to: the sigmoid operator, the log operator with an inverse sigmoid operator, a grid-sample operator, and a pin-hole camera projection.

[0109] FIG. 9 illustrates an example 900 of logarithmic input clamping, in accordance with aspects of the present disclosure. Example 900 illustrates the Inverse sigmoid representation with various functions and operators such as clipping function 902, subtraction 904, clipping function 906, clipping function 908, division operator 910, and log operator 912. Other functions are possible.

[0110] The output of sigmoid is between [0.0, 1.0], While the sigmoid operator is defined as a single operator in the open neural network exchange (ONNX) specification, but the inverse sigmoid is not.

[0111] Clipping function 902 clips the range of input to [0.0, 1.0], Clipping function 906 operator is required for handling divide-by-zero case and right clip is needed to handle the case of log(0) = -inf. The actual data range to division operator 910 would be:[0, l / le-5], that is [0, 100000]

[0112] As the range of the output of log operator 912 is going to be [-6, 6], the input which is output of division operator 910 can be clamped between [exp(-6), exp(6)], or approximately [0.00247. 403.428793493],Sigmoid y = 1 / (1 + exp(-x))Inverse sigmoid x = In (y / (l — y))

[0113] As the output of sigmoid is between [0.0, 1.0], clipping function 902 clips the range of input to [0.0, 1.0], Clipping function 906 is required for handling divide-by-zero case and clipping function 908 is needed to handle the log(0) = -inf case.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO30

[0114] An actual data range to the output of division operator 910 would be [0, l / le-5], that is [0, 100000], As the range of the output of log operator 912 will to be [-6, 6], the input which is output of division operator 910 can be clamped between [exp(-6), exp(6)], or approximately [0.00247, 403.428793493],

[0115] With respect to the division operator 910, an artificially low epsilon value of le-5 may only be used for mathematical stability to avoid div-by-zero errors. However, because the quantization algorithm would not have any such knowledge, the quantization algorithm would operate as if the epsilon value of le-5 is the real minimum of the data and may attempt to account for the value in quantization. Because l / le-5 is 105, it is very7large and may compromise accuracy severely. In this case, keeping it to a modest value of 1.0 (where div-by-1 does not change the value of the input) may keep the data range to the actual range of the input, which can be much more friendly for 8-bit quantization.

[0116] FIG. 10 illustrates a range 1000 of an inverse sigmoid function, in accordance with aspects of the present disclosure. Range 1000 illustrates the inverse sigmoid representation. As can be seen, an output of the inverse sigmoid function is between 0 and 1.

[0117] FIG. 11 illustrates an example 1100 of grid-sample input range clamping, in accordance with aspects of the present disclosure. Example 1100 includes a depiction of the GridSample operator 1102, input 1104, configuration 1106, and output 1108. The input 1104, configuration 1106, and output 1108 are user interface components of a system used to model the GridSampel operator.

[0118] The GridSample operator 1102 has two inputs 1104: the activation tensor that needs to be sampled and the grid tensor that offsets into the first input. A definition of the GridSample operator 1102 is shown as follows:• Given an input and a flow-field grid, the operator computes the output using input values and pixel locations from grid.• The grid specifies the sampling pixel locations normalized by the input spatial dimensions. Therefore, it should have most values in the range of [-1, 1]. ForPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO31example, values x = -1, y = -1 is the left-top pixel of input, and values x = 1, y = 1 is the right-bottom pixel of input.• If the grid has values outside the range of [-1, 1], the corresponding outputs are handled as defined by padding mode. Example - padding_mode=" zeros": use 0 for out-of-bound grid locations.

[0119] As with previous cases, inputs to the Gridsample operator 1102 are expected to be within [-1, 1] (or realistically slightly wider like [-1.25, +1.25]) range. However, an offset producer layer would not have any knowledge about this constraint. Rather, such a layer would simply produce their outputs in their natural ranges ([-135, +135] in the example). Hence, disclosed solutions apply smart range clipping of [-1.25, +1.25] on the sampling tensor (which contains offsets) to keep the output quantized data rich in representation instead of being very' sparse.

[0120] As described above, the [-1 ,-l] and [1, 1] in the grid tensor represents the top left and bottom right comers of the first input. This grid tensor could be dynamically computed in the graph. As shown in the graph snippet in FIG. 11, the grid tensor in many of the grid-sample operators can have arbitrarily large ranges ([-140, 140] in the example). If these actual data ranges are considered for quantization, then the relevant grid offsets [-1, 1] would be always quantized [0, 0], resulting in zero accuracy with 8-bit computation for these operators. As the padded regions are also needed for maintaining the functionality of the model, clamping the tensor between [-1.25 to 1.25] would result in best accuracy at 8-bit inference.

[0121] FIG. 12 illustrates an example 1200 of epsilon handling for a pin-hole camera projection, in accordance with aspects of the present disclosure. Example 1200 includes mathematical functions such as max function 1202, slice function 1204, and div function 1206. In an example, slice function 1204 can extract a portion of a number. Example 1200 further depicts user interface panels, specifically, name 1208, category 1210, floating point selection 1212, shape 1214, value 1216, which in this example, may be used to generate values.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO32

[0122] As depicted, a first input is provided to max function 1202 and a second input is provided to slice function 1204. The outputs of max function 1202 and slice function 1204 are provided as inputs to div function 1206, which outputs a result.

[0123] In an example, a pinhole camera projection refers to a projection onto an image plane of an ideal pinhole camera, where the camera aperture is described as a point and no lenses are used to focus light.

[0124] In a bird’s eye view (BEV)-based model there is a need for doing an operation to project 3D co-ordinates to camera plane using pin-hole camera projection. An example can be represented as follows:x = f * X / Z, y = f * Y IZ, where f is the focal point.

[0125] The computation involves a ratio of 3D coordinates (X. Y) by the depth coordinate (Z) using a divide operation. Implementations often employ divide-by-zero avoidance techniques such as using a small epsilon in the denominator using floating point computations.

[0126] For fixed-point implementation, however, using very small epsilon values (such as le-5 as depicted in FIG. 12 can increase the input range (1 / epsilon) to the actual Div operator artificially to account for the outliers. This introduces huge quantization noise in the output of the Div operator, compromising accuracy in further quantized nodes.

[0127] Disclosed solutions boost this epsilon to a more reasonable value to 1.0 (instead of le-5) knowing that it does not affect the computation of x and y in any meaningful way. Note that if Z is zero (which is virtual condition anyway) it does not hurt to consider as Z =1.0.

[0128] FIG. 13 is a flow diagram illustrating an example of a process for quantization of parameters of machine learning models, in accordance with aspects of the disclosure. The process 1300 may be performed by a computing device (or apparatus) or a component (e.g.. a chipset. codec, etc.) of the computing device. The operations of the process 1300 may be implemented as software components that are executed and run on one or morePolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO33processors (such as CPU 102, GPU 104, DSP 106, NPU 108 of FIG. 1, processor 1510 of FIG. 15, etc.).

[0129] At block 1302. the computing device (or component thereof) may process first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights. Examples of machine learning models are depicted in FIGs. 2, 3, 4, and 14.

[0130] At block 1304, the computing device (or component thereof) may determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values. For instance, FIG. 8 depicts a usable range values associated with a sigmoid function, whereas FIG. 10 depicts the usable range of values associated with an inverse sigmoid function.

[0131] At block 1306, the computing device (or component thereof) may scale, based on the usable range of input values, the output activations to obtain scaled output activations.

[0132] At block 1308, the computing device (or component thereof) may quantize the scaled output activations to obtain quantized activations.

[0133] In some aspects, the computing device (or component thereof) may process second data using the second layer of the machine learning model and the quantized activations to generate output content. In an example, to scale the output activations to obtain the scaled output activations, computing device (or component thereof) is configured to clip a range of the output activations based on the usable range of input values.

[0134] In an example, the usable range of input values is based on a sigmoid function. In an example, the usable range of input values is based on a logarithmic operator of an inverse sigmoid function. In an example, the usable range of input values includes sampling offsets based on a grid sample operator. In an example, the usable range of input values is based on pin-hole camera projection. In an example, to scale the output activations to obtain the scaled output activations, the computing device (or componentPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO34thereof) is further configured to clip a lower range of the output activations. In an example, the computing device (or component thereof) is configured to process second data using the second layer of the machine learning model and the quantized activations to generate output content.

[0135] In some examples, the techniques or processes described herein may be performed by a computing device, an apparatus, and / or any other computing device. In some cases, the computing device or apparatus may include a processor, microprocessor, microcomputer, or other component of a device that is configured to carry out the steps of processes described herein.

[0136] The processes described herein can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, sand the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0137] In some cases, the devices or apparatuses configured to perform the operations of the process 1300 and / or other processes described herein may include a processor, microprocessor, micro-computer, or other component of a device that is configured to carry out the steps of the process 1300 and / or other process. In some examples, such devices or apparatuses may include one or more sensors configured to capture image data and / or other sensor measurements. In some examples, such computing device or apparatus may include one or more sensors and / or a camera configured to capture one or more images or videos. In some cases, such device or apparatus may include a display for displaying images. In some examples, the one or more sensors and / or camera are separate from the device or apparatus, in which case the device or apparatus receives the sensed data. Such device or apparatus may further include a network interface configured to communicate data.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO35

[0138] The components of the device or apparatus configured to carry out one or more operations of the process 1300 and / or other processes described herein can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The computing device may further include a display (as an example of the output device or in addition to the output device), a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.

[0139] The process 1300 is illustrated as a logical flow diagram, the operations of which represent sequences of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0140] Additionally, the processes described herein (e.g., the process 1300 and / or other processes) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program including a plurality ofPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO36instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0141] Additionally, the processes described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0142] FIG. 14 is a block diagram of an example transformer in accordance with some aspects of the disclosure. In a convolutional neural network (CNN) model, the number of operations required to relate signals from two arbitrary input or output positions grows in the distance between positions, which makes learning dependencies at different distant positions challenging for a CNN model. The transformer 1400 reduces the operations of learning dependencies by using an encoder 1410 and a decoder 1430 that implements an attention mechanism at different positions of a single sequence to compute a representation of that sequence. An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors. The output is computed as a weighted sum of the values, where the weight assigned to each value is computed by a compatibility function of the query with the corresponding key.

[0143] In some examples of a transformer, the encoder 1410 is composed of a stack of six identical layers and each layer has two sub-layers. The first sub-layer is a multi-head self-attention engine 1412, and the second sub-layer is a fully connected feed-forward network 1414. A residual connection (not shown) connects around each of the sub-layers followed by normalization.

[0144] In this example of a transformer 1400, the decoder 1430 is also composed of a stack of six identical layers. The decoder also includes a masked multi-head self-attentionPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO37engine 1432, a multi-head atention engine 1434 over the output of encoder 1410, and a fully connected feed-forward network 1426. Each layer includes a residual connection (not shown) around the layer, which is followed by layer normalization. The masked multi -head self-atention engine 1432 is masked to prevent positions from atending to subsequent positions and ensures that the predictions at position i can depend only on the known outputs at positions less than i (e.g., auto-regression).

[0145] In the transformer 1400, the queries, keys, and values are linearly projected by a multi -head attention engine into learned linear projects, and then atention is performed in parallel on each of the learned linear projects, which are concatenated and then projected into final values.

[0146] The transformer also includes a positional encoder 1440 to encode positions because the model does not contain recurrence and convolution and relative or absolute position of the tokens is needed. For example, the positional encodings are added to the input embeddings at the botom layer of the encoder 1410 and the decoder 1430. The positional encodings are summed with the embeddings because the positional encodings and embeddings have the same dimensions. A corresponding position decoder 1450 is configured to decode the positions of the embeddings for the decoder 1430.

[0147] In some aspects, the transformer 1400 uses self-atention mechanisms to selectively weigh the importance of different parts of an input sequence during processing and allow s the model to atend to different parts of the input sequence w hile generating the output. The input sequence is first embedded into vectors and then passed through multiple layers of self-atention and feed-forward networks. The transformer 1400 can process input sequences of variable length, making it well-suited for natural language processing tasks where input lengths can vary greatly. Additionally, the self-atention mechanism allows the transformer 1400 to capture long-range dependencies between words in the input sequence, which is difficult for RNNs and CNNs. The transformer with self-atention has achieved results in several natural language processing tasks that are beyond the capabilities of other neural networks and has become a popular choice for language and text applications. For example, the various large language models, such asPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO38a generative pretrained transformer (e g., ChatGPT, etc.) and other current models are ty pes of transformer networks.

[0148] FIG. 15 illustrates an example computing device architecture 1500 of an example computing device which can implement the various techniques described herein. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or computing device of a vehicle), or other device. The components of computing device architecture 1500 are shown in electrical communication with each other using connection 1505, such as a bus. The example computing device architecture 1500 includes a processing unit (CPU or processor) 1510 and computing device connection 1505 that couples various computing device components including computing device memory 1515, such as read only memory (ROM) 1520 and random access memory7(RAM) 1525, to processor 1510.

[0149] Computing device architecture 1500 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1510. Computing device architecture 1500 can copy data from memory71515 and / or the storage device 1530 to cache 1512 for quick access by processor 1510. In this w ay, the cache can provide a performance boost that avoids processor 1510 delays while waiting for data. These and other modules can control or be configured to control processor 1510 to perform various actions. Other computing device memory71515 may be available for use as well. Memory71515 can include multiple different types of memory7with different performance characteristics. Processor 1510 can include any general purpose processor and a hardware or software service, such as service 1 1532, service 2 1534, and service 3 1536 stored in storage device 1530, configured to control processor 1510 as well as a special -purpose processor where software instructions are incorporated into the processor design. Processor 1510 may be a self-contained system, containing multiple cores or processors, a bus. memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO39

[0150] To enable user interaction with the computing device architecture 1500, input device 1545 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. Output device 1535 can also be one or more of a number of output mechanisms known to those of skill in the art, such as a display, projector, television, speaker device, etc. In some instances, multimodal computing devices can enable a user to provide multiple types of input to communicate with computing device architecture 1500. Communication interface 1540 can generally govern and manage the user input and computing device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.

[0151] Storage device 1530 is a non-volatile memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs) 1525, read only memory (ROM) 1520, and hybrids thereof. Storage device 1530 can include services 1532, 1534.1536 for controlling processor 1510. Other hardware or software modules are contemplated. Storage device 1530 can be connected to the computing device connection 1505. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1510, connection 1505, output device 1535, and so forth, to carry out the function.

[0152] Aspects of the present disclosure are applicable to any suitable electronic device (such as security systems, smartphones, tablets, laptop computers, vehicles, drones, or other devices) including or coupled to one or more active depth sensing systems. While described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors, and are therefore not limited to specific devices.

[0153] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, one processing system and so on). As used Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO40herein, a device may be any electronic device with one or more parts that may implement at least some portions of this disclosure. While the below description and examples use the term ‘‘device'’ to describe various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or specific embodiments. For example, a system may be implemented on one or more printed circuit boards or other substrates, and may have movable or static components. While the below description and examples use the term “system” to describe various aspects of this disclosure, the term “system” is not limited to a specific configuration, type, or number of objects.

[0154] Specific details are provided in the description above to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. For clarity of explanation, in some instances the present technology7may be presented as including individual functional blocks including functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary’ detail in order to avoid obscuring the embodiments.

[0155] Individual embodiments may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO41

[0156] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.

[0157] The term “computer-readable medium’’ includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruct on(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitor ' medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as flash memory, memory or memory devices, magnetic or optical disks, flash memory, USB devices provided with non-volatile memory; networked storage devices, compact disk (CD) or digital versatile disk (DVD), any suitable combination thereof, among others. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0158] In some embodiments, the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy; carrier signals, electromagnetic waves, and signals per se.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO42

[0159] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary7tasks (e g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality7described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.

[0160] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.

[0161] In the foregoing description, aspects of the application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative embodiments of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate embodiments, the methods may be performed in a different order than that described.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO43

[0162] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein can be replaced with less than or equal to (“<”) and greater than or equal to (“>”) symbols, respectively, without departing from the scope of this description.

[0163] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operation, or any combination thereof.

[0164] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.

[0165] Claim language or other language reciting “at least one of’ a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B. C and C. A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of’ a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.

[0166] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor orPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO44multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y. and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z’' can mean that any single processor may only perform at least a subset of operations X, Y, and Z.

[0167] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.

[0168] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO45one component (e g., different components may perform different sub-functions of a function).

[0169] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0170] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory' (SDRAM), read-only memory' (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the formPolsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO46of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0171] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.

[0172] Illustrative aspects of the disclosure include:

[0173] Aspect 1. An apparatus for generating content, comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: process first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights; determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scale, based on the usable range of input values, the output activations to obtain scaled output activations; and quantize the scaled output activations to obtain quantized activations.

[0174] Aspect 2. The apparatus of Aspect 1, wherein, to scale the output activations to obtain the scaled output activations, the at least one processor is configured to clip a range of the output activations based on the usable range of input values.

[0175] Aspect 3. The apparatus of any of Aspects 1-2, wherein the usable range of input values is based on a sigmoid function.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO47

[0176] Aspect 4. The apparatus of any of Aspects 1-3, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

[0177] Aspect 5. The apparatus of any of Aspects 1-4, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.

[0178] Aspect 6. The apparatus of any of Aspects 1-5, wherein the usable range of input values is based on pin-hole camera projection and wherein to scale the output activations to obtain the scaled output activations, the at least one processor is further configured to clip a lower range of the output activations.

[0179] Aspect 7. The apparatus of any of Aspects 1-6, wherein the at least one processor is configured to process second data using the second layer of the machine learning model and the quantized activations to generate output content.

[0180] Aspect 8. A method of generating content, comprising: processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights; determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scaling, based on the usable range of input values, the output activations to obtain scaled output activations; and quantizing the scaled output activations to obtain quantized activations.

[0181] Aspect 9. The method of Aspect 8, wherein scaling the output activations to obtain the scaled output activations comprises clipping a range of the output activations based on the usable range of input values.

[0182] Aspect 10. The method of any of Aspects 8-9, wherein the usable range of input values is based on a sigmoid function.

[0183] Aspect 11. The method of any of Aspects 8-10, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

[0184] Aspect 12. The method of any of Aspects 8-11, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO48

[0185] Aspect 13. The method of any of Aspects 8-12, wherein the usable range of input values is based on pin-hole camera projection and wherein to scale the output activations to obtain the scaled output activations, the method further comprising clipping a lower range of the output activations.

[0186] Aspect 14. The method of any of Aspects 8-13, further comprising processing second data using the second layer of the machine learning model and the quantized activations to generate output content.

[0187] Aspect 15. Anon-transitory' computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights; determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values; scaling, based on the usable range of input values, the output activations to obtain scaled output activations; and quantizing the scaled output activations to obtain quantized activations.

[0188] Aspect 16. The non-transitory computer-readable medium of Aspect 15, wherein the usable range of input values is based on a sigmoid function.

[0189] Aspect 17. The non-transitory computer-readable medium of any of Aspects 15- 16, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

[0190] Aspect 18. The non-transitory computer-readable medium of any of Aspects 15- 17, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.

[0191] Aspect 19. The non-transitory computer-readable medium of any of Aspects 15- 18, wherein the usable range of input values is based on pin-hole camera projection and wherein scaling the output activations to obtain the scaled output activations comprises clipping a lower range of the output activations.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO49

[0192] Aspect 20. The non-transitory computer-readable medium of any of Aspects 15-19, wherein the instructions, when executed by the one or more processors, cause the one or more processors to process second data using the second layer of the machine learning model and the quantized activations to generate output content.

[0193] Aspect 21. An apparatus for generating content, the apparatus including one or more means for performing operations according to any of Aspects 8 to 14.Polsinelli Docket No. 094922-842957

Claims

PATENTQualcomm Docket No 2503139WO50CLAIMS WHAT IS CLAIMED IS:

1. An apparatus for generating content, comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to: process first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights;determine, based on a characteristic of a second layer of the machine learning model, a usable range of input values;scale, based on the usable range of input values, the output activations to obtain scaled output activations; andquantize the scaled output activations to obtain quantized activations.

2. The apparatus of claim 1, wherein, to scale the output activations to obtain the scaled output activations, the at least one processor is configured to clip a range of the output activations based on the usable range of input values.

3. The apparatus of claim 1, wherein the usable range of input values is based on a sigmoid function.

4. The apparatus of claim 1, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

5. The apparatus of claim 1, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.

6. The apparatus of claim 1, wherein the usable range of input values is based on pin-hole camera projection and wherein to scale the output activations to obtain the scaled output activations, the at least one processor is further configured to clip a lower range of the output activations.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO517. The apparatus of claim 1. wherein the at least one processor is configured to process second data using the second layer of the machine learning model and the quantized activations to generate output content.

8. A method of generating content, comprising:processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights;determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values;scaling, based on the usable range of input values, the output activations to obtain scaled output activations; andquantizing the scaled output activations to obtain quantized activations.

9. The method of claim 8, wherein scaling the output activations to obtain the scaled output activations comprises clipping a range of the output activations based on the usable range of input values.

10. The method of claim 8, wherein the usable range of input values is based on a sigmoid function.

11. The method of claim 8, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

12. The method of claim 8, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.

13. The method of claim 8, wherein the usable range of input values is based on pinhole camera projection and wherein to scale the output activations to obtain the scaled output activations, the method further comprising clipping a lower range of the output activations.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO5214. The method of claim 8. further comprising processing second data using the second layer of the machine learning model and the quantized activations to generate output content.

15. A non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to:processing first data using a first layer of a machine learning model to obtain output activations, wherein the first layer comprises quantized weights;determining, based on a characteristic of a second layer of the machine learning model, a usable range of input values;scaling, based on the usable range of input values, the output activations to obtain scaled output activations; andquantizing the scaled output activations to obtain quantized activations.

16. The non-transitory computer-readable medium of claim 15, wherein the usable range of input values is based on a sigmoid function.

17. The non-transitory computer-readable medium of claim 15, wherein the usable range of input values is based on a logarithmic operator of an inverse sigmoid function.

18. The non-transitory computer-readable medium of claim 15, wherein the usable range of input values comprises sampling offsets based on a grid sample operator.

19. The non-transitory computer-readable medium of claim 15, wherein the usable range of input values is based on pin-hole camera projection and wherein scaling the output activations to obtain the scaled output activations comprises clipping a lower range of the output activations.Polsinelli Docket No. 094922-842957PATENTQualcomm Docket No 2503139WO5320. The non-transitory computer-readable medium of claim 15, wherein the instructions, when executed by one or more processors, cause the one or more processors to process second data using the second layer of the machine learning model and the quantized activations to generate output content.Polsinelli Docket No. 094922-842957