Adaptive hybrid resolution processing

By generating input tokens of different resolutions and processing these tokens using neural networks and transformer models, dynamically determining the scale of the image area is solved, and the deep learning model has high computational complexity and memory requirements in image processing is achieved, achieving more efficient training and processing.

CN120202491APending Publication Date: 2025-06-24QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076479.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-29
Filing Date
2023-10-02
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Deep learning machine learning models, especially transformer models, lead to high computational complexity, high memory requirements and increased training difficulty due to the tokenization process when processing image data.

Method used

Optimize the token number and calculate cost by generating input tokens at different resolutions and processing these tokens using neural network models and transformer models to dynamically determine the scale of the image area.

Benefits of technology

It reduces the number of tokens during the transformer model processing, reduces the computational complexity and memory requirements, and improves the training efficiency and model processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202491A_ABST
    Figure CN120202491A_ABST
Patent Text Reader

Abstract

Systems and techniques for adaptive hybrid resolution processing are described. According to some aspects, a device may divide an input image into a first token having a first resolution and a second token having a second resolution. The device may generate a first token representation for token (s) from a first token corresponding to a first region of the input image; and generating a second token representation for token (s) from the second token corresponding to the first region of the input image. The device may process the first token representation and the second token representation using a neural network model to determine either the first resolution or the second resolution as a scale of the first region of the input image. The device may process a first region of the input image according to a scale of the first region using a transformer neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 424,761, filed on Nov. 11, 2022, which is hereby incorporated by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure generally relates to using one or more machine - learning systems to process image data. Aspects of the present disclosure relate to providing adaptive hybrid - resolution processing for generating input tokens at different resolutions for input into a machine - learning system or model (e.g., a transformer neural - network system or model). Background Art

[0004] Deep - learning machine - learning models (e.g., neural networks) can be used to perform various tasks, such as detection and / or recognition (e.g., scene or object detection and / or recognition), depth estimation, pose estimation, image reconstruction, classification, three - dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, image processing, etc. Deep - learning machine - learning models can be general - purpose and can achieve high - quality results in various tasks. However, although deep - learning machine - learning models can be general - purpose and accurate, these models can be large and slow, and generally have high memory requirements and computational costs. In many cases, the computational complexity of the model can be very high, and the model can be difficult to train.

[0005] In some cases, a machine - learning model can utilize one or more transformers. Tokens are used by the transformer as the basic unit of inference. For example, an input image can be partitioned into several tokens, which can be input into the transformer for processing. However, the tokenization process can be sub - optimal because the tokens do not carry semantic meaning, and the number of tokens increases quadratically with the image size, which can lead to an inefficient machine - learning model.

[0006] Overview

[0007] Systems and techniques are described for processing data (e.g., one or more images or video frames) by generating input tokens of the data at different resolutions for input into a machine - learning system or model, such as a transformer.

[0008] In some aspects, a method for processing image data is provided. The method includes: dividing an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generating a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generating a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; processing the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and processing the first region using a transformer neural network model according to the scale of the first region of the input image.

[0009] In some aspects, an apparatus for processing image data is provided. The apparatus includes at least one memory and at least one processor (e.g., configured in a circuit system), the at least one processor being coupled to the at least one memory and configured to: divide an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generate a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generate a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; process the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and process the first region using a transformer neural network model according to the scale of the first region of the input image.

[0010] In some aspects, a non-transitory computer-readable medium storing instructions is provided, the instructions when executed by one or more processors (e.g., configured in a circuit system) cause the one or more processors to: divide an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generate a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generate a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; process the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and process the first region using a transformer neural network model according to the scale of the first region of the input image.

[0011] In some aspects, a device for processing image data is provided. The device includes: means for dividing an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; means for generating a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; means for generating a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; means for using a neural network model to process the first set of token representations and the second set of token representations to determine the first resolution or the second resolution as the scale of the first region of the input image; and means for using a transformer neural network model to process the first region according to the scale of the first region of the input image.

[0012] In some aspects, one or more of the devices described herein are, are part of, and / or include one or more of the following: an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device, a wireless communication device, a camera, a personal computer, a laptop computer, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device acting as a server device (such as a mobile phone), an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-described device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyroscopic testers, one or more accelerometers, any combination thereof, and / or other sensors).

[0013] This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0014] The foregoing and other features and aspects will become more apparent when reference is made to the following specification, claims, and appended drawings. Brief Description of the Drawings

[0016] Exemplary embodiments of the present application are described in detail below with reference to the following drawings:

[0017] Figure 1 is a diagram illustrating an example of a convolutional neural network (CNN) according to aspects of the present disclosure;

[0018] Figure 2 is a diagram illustrating an example of a classification neural network system according to aspects of the present disclosure;

[0019] Figure 3 is a diagram illustrating an example of an existing tokenization process according to aspects of the present disclosure;

[0020] Figure 4 is a diagram illustrating an example of a system including a preprocessing engine that generates mixed-scale data for a transformer according to aspects of the present disclosure.

[0021] Figure 5 is a diagram illustrating a high-level overview of generating mixed-scale tokens according to aspects of the present disclosure;

[0022] Figure 6 is a diagram illustrating the differences between a previous transformer process and the mixed-scale method disclosed herein according to aspects of the present disclosure;

[0023] Figure 7 is a diagram illustrating an example of an existing merging and pruning process compared to the dynamic method disclosed herein according to aspects of the present disclosure;

[0024] Figure 8 is a diagram illustrating a comparison between fixed-scale tokens and dynamic mixed-scale patterns according to aspects of the present disclosure;

[0025] Figure 9 is a diagram illustrating a binary gate decision process for each spatial location and each input image according to aspects of the present disclosure;

[0026] Figure 10 is a diagram illustrating an example of a method for performing object detection according to aspects of the present disclosure;

[0027] Figure 11 is a diagram illustrating an example of a computing system according to aspects of the present disclosure.

[0028] Detailed Description

[0029] Certain aspects of the present disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination, which will be obvious to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of the aspects of the present application. However, it is obvious that the aspects may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0030] The following description only provides example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the example aspects will provide those skilled in the art with an enabling description for implementing the example aspects. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application set forth in the appended claims.

[0031] As previously mentioned, some deep learning machine learning models (e.g., neural networks) can utilize one or more transformers. A transformer is a specific type of neural network. An example of a system using transformers is the MobileViT system described in MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Mehta, Rastegari, ICLR, 2022 (MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Mehta, Rastegari, ICLR, 2022), which is incorporated herein by reference. When used for certain tasks (e.g., image classification and object detection), transformers perform well, but require a large amount of computation to perform image classification and object detection tasks.

[0032] Transformers typically use tokens as the basic units for processing or inference. For example, an input image can be divided into N tokens or patches (e.g., square tokens or patches), all of which have the same size. However, the tokenization process may be suboptimal because the tokens do not carry semantic meaning, and the number of tokens increases quadratically with the image size, which may lead to an inefficient model. Transformers incur a high cost in terms of multiply-accumulate computations (MACs) (which may be required to vary with the number of input tokens). For example, for an input size of 384 pixels, a patch size or scale value of 8 can produce 2048 tokens. In another example, a patch size or scale value of 32 can produce 100 tokens. In these two examples, there is a significant difference in the number of tokens used and the number of MACs required, especially at the lower end of the number of tokens (e.g., 0 - 300 or other values). In many cases, current transformer-based models can utilize many more patches than actually needed, such as because the patches do not carry any semantic information. Thus, generally speaking, it is desirable to construct signal processing in a way that reduces the number of patches and thus the number of tokens.

[0033] In addition, a transformer can include a global attention block and an independent token refinement block, which are affected by the number of input tokens at each layer. For example, in the global attention block of a transformer, each token of the input image provides an update to every other token during processing. The processing is performed in parallel. The updates can be weighted by the degree of similarity between two tokens. The cost of the global attention process can be quadratic in nature, such as O(N 2 ) computations in the presence of N tokens. In addition, in the individual feed-forward network (FFN) of a transformer, each token is updated independently, and the parameters of the FFN are shared across all tokens. The processing of the FFN can also be performed in parallel. The cost of the FFN scales linearly with the FFN cost as O(N cost FFN ).

[0034] This document describes systems, apparatuses (or devices), methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for generating input tokens at different resolutions or scales for input into a machine learning system or model, such as a transformer-based neural network system or model. The systems and techniques can address the problem of transformers having a large cost in processing input images due to the number of tokens required for processing.

[0035] According to some aspects, the systems and techniques can predict the tokenization scale for each region of an image as a preprocessing step before the transformer (e.g., performed by a preprocessing engine). Intuitively, uninformative image regions (such as the background) can be processed at a coarser scale than the foreground without losing information, resulting in a smaller total number of tokens.

[0036] To capture such behavior, a conditional gating engine (e.g., a neural network layer) can be trained to select the best tokenization scale for each coarse local region within the input image. In some cases, the gating engine can include a lightweight multi-layer perceptron (MLP) that takes the local coarse region of the image as input and predicts the tokenization scale for that region of the image, resulting in a dynamic number of tokens per image. Since the gating engine operates at the input level, the gating engine is agnostic to the choice of the transformer backbone.

[0037] To avoid potential issues with learning such scale selection engines (e.g., training with additional parameters for each scale or having a cumbersome training pipeline with multiple stages), a unified single-stage model can be trained by maximizing parameter sharing across scales. Additionally, to avoid the gating engine getting stuck in poor local minima (e.g., always outputting the same trivial static pattern), a training loss can be used that enables finer control over the learned gating distribution, thus enhancing the dynamic behavior of the hybrid-scale tokenization. To reduce the training cost, an adaptive pruning strategy can be performed during training, which can rely on the underlying mapping between the coarse tokens and the fine tokens.

[0038] The dynamic scale (or hybrid-scale) selection gating mechanism acts as a preprocessing stage, agnostic to the choice of the transformer backbone, and can be jointly trained with the transformer in a single stage with hybrid-scale tokens as input. Additionally, when training the dynamic gates, batch shaping generalization can be used to better handle multi-dimensional distributions. The resulting loss provides better control over the learned scale distribution and allows for easier and better initialization of the gates. By locally defining the gates only at the coarse token level and adopting an adaptive pruning strategy during training, the training overhead incurred by handling the token sets for each scale can also be reduced.

[0039] As previously mentioned, the systems and techniques introduce hybrid-scale (or hybrid-resolution) transformers that can be configured as unified models for handling (e.g., in parallel) input tokens from multiple scales. As described above, a preprocessing engine can be applied to the input before the transformer to generate tokens with hybrid resolution or scale. For example, the preprocessing engine can perform minimum token selection based on the gating decisions made by the gating engine and can mask the input patches (e.g., of one or more images) after the gating decisions. For example, based on the mask output by the gating decisions, the input image can be partitioned into tokens of different resolutions. Thus, the tokens fed to the transformer can be configured with multiple different tokens having different scales.

[0040] In some cases, the preprocessing engine can be the first or early in terms of processing the layers of the input image. The preprocessing engine can dynamically generate hybrid-scale tokens that cover the entire image. If necessary, the transformer can be modified to handle the hybrid-scale tokens. Subsequently, the transformer can process the entire image in a more efficient manner compared to the prior art and utilize the hybrid information from multiple image scales. In an illustrative example, processing an input image using the prior art can result in 2000 tokens being processed by the transformer, resulting in 2000^2 computations. However, using the systems and techniques described herein, the number of tokens can be reduced to 250 or less, which results in the required computations being reduced to 250^2 or less.

[0041] Hybridizing information from multiple image scales (e.g., by dividing the image into tokens of different scales) can result in computational efficiency without sacrificing performance, at least in part because not all regions of the image contain fine details. For example, some regions of the image (e.g., the blue sky) have little variation in texture (e.g., color, edges, etc.), and other regions (e.g., buildings, people, vehicles, etc.) may have detailed textures. In some cases, the gating decisions of the preprocessing engine can be based on learning dynamic hybrid-scale patterns early in the network (e.g., in one or more initial layers of a neural network model that includes a transformer), such as learning the background of a scene in the image relative to one or more salient objects in the scene. In some cases, some images may have different image complexities, where one image may be a less complex sunset, while another image may be a cluttered room with many different objects. For example, the preprocessing engine can learn a finer scale for salient objects (or in some cases for all foreground objects) or for more complex parts of the input image. The preprocessing engine can learn a coarser scale for small patches corresponding to the background in the input image or for simpler parts of the input image. The preprocessing engine performs the above-mentioned gating decisions based on the importance or complexity of each region of the input image. The output of the gating decision (e.g., a mask) can indicate which image regions the model will place higher focus on by selecting a finer resolution. Thus, the systems and techniques can adapt to spend more time on complex regions of the image (or more complex images overall) and save more computations on simple parts of the image (or simple images overall) to optimize the average computational cost.

[0042] Early learning of hybrid-scale patterns based on the scene in the image can improve the efficiency of the system across all layers. Leveraging the property that different regions of the image contain different levels of information, the systems and techniques achieve a hybrid image scale based on the characteristics of each part of the image.

[0043] The systems and techniques described herein can be used to enhance machine learning systems (e.g., transformer-based models of neural networks) for performing any task of processing image or video frames. One benefit of the hybrid scale or resolution tokenization process is that it can improve the efficiency-accuracy tradeoff of vision transformers configured to process one or more image or video frames. For example, the systems and techniques can be helpful when processing sparse images with small objects (such as images for aerial detection, where the image is captured overhead and at a distance).

[0044] Aspects of the present application will be described with reference to the accompanying drawings.

[0045] Figure 1 is a diagram illustrating an example of a CNN 100. The input layer 102 of the CNN 100 includes data representing an image. For example, the data can include a digital array representing image pixels, where each number in the array includes a value from 0 to 255, which describes the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28×28×3 digital array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.). The image can pass through convolutional hidden layers 104, an optional non-linear activation layer, pooling hidden layers 106, and fully connected hidden layers 108 to obtain an output at the output layer 110. Although only one of each hidden layer is shown in Figure 1 only one of each hidden layer is shown, one of ordinary skill in the art will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 100. As previously described, the output can indicate an object of a single class or can include probabilities of the classes that best describe the objects in the image.

[0046] The first layer of CNN 100 is the convolutional hidden layer 104. The convolutional hidden layer 104 analyzes the image data of the input layer 102. Each node of the convolutional hidden layer 104 is connected to a region of nodes (pixels) in the input image, which is called the receptive field. The convolutional hidden layer 104 can be considered as one or more filters (each filter corresponding to a different activation map or feature map), where each convolutional iteration of the filter is a node or neuron of the convolutional hidden layer 104. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of that filter. In an illustrative example, if the input image includes a 28×28 array and each filter (and the corresponding receptive field) is a 5×5 array, there will be 24×24 nodes in the convolutional hidden layer 104. Each connection between a node and the receptive field of that node learns a weight, and in some cases, a global bias is learned so that each node learns to analyze its specific local receptive field in the input image. Each node of the hidden layer 104 will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has an array of weights (numbers) and the same depth as the input. For a video frame example, the filter will have a depth of 3 (according to the three color components of the input image). An illustrative example of the size of the filter array is 5×5×3, which corresponds to the size of the receptive field of the node.

[0047] The convolutional nature of the convolutional hidden layer 104 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, the filter of the convolutional hidden layer 104 can start from the upper left corner of the input image array and can be convolved around the input image. As mentioned above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 104. At each convolutional iteration, the values of the filter are multiplied by the corresponding number of raw pixel values of the image (e.g., the 5×5 filter array is multiplied by the 5×5 array of input pixel values at the upper left corner of the input image array). The multiplications from each convolutional iteration can be added together to obtain a sum for that iteration or node. Next, the process continues at the next position in the input image according to the receptive field of the next node in the convolutional hidden layer 104.

[0048] For example, the filter can move one step amount to the next receptive field. The step amount can be set to 1 or other appropriate amount. For example, if the step amount is set to 1, the filter will move 1 pixel to the right at each convolutional iteration. Processing the filter at each unique position of the input amount produces a number representing the filter result for that position, thus resulting in the determination of a sum value for each node of the convolutional hidden layer 104.

[0049] The mapping from the input layer to the convolutional hidden layer 104 is referred to as an activation map (or feature map). The activation map includes the value of each node, which represents the filtering result at each position of the input quantity. The activation map can include an array that includes various sum values generated by each iteration of the filter over the input quantity. For example, if a 5x5 filter is applied to each pixel of a 28x28 input image (with a stride amount of 1), the activation map will include a 24x24 array. The convolutional hidden layer 104 can include a number of activation maps in order to identify multiple features in the image. Figure 1 The example shown in Figure 1 includes three activation maps. By using three activation maps, the convolutional hidden layer 104 can detect three different types of features, where each feature is detectable across the entire image.

[0050] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 104. The non-linear layer can be used to introduce non-linearity into a system that has been computing linear operations. An illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. The ReLU layer can apply the function f(x) = max(0, x) to all values in the input quantity, which changes all negative activations to 0. Thus, the ReLU can increase the non-linear properties of the CNN 100 without affecting the receptive field of the convolutional hidden layer 104.

[0051] The pooling hidden layer 106 can be applied after the convolutional hidden layer 104 (and after the non-linear hidden layer when in use). The pooling hidden layer 106 is used to simplify the information in the output from the convolutional hidden layer 104. For example, the pooling hidden layer 106 can take each activation map output from the convolutional hidden layer 104 and use a pooling function to generate a compressed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. Other forms of pooling functions (such as average pooling, L2 norm pooling, or other suitable pooling functions) are used by the pooling hidden layer 104. The pooling function (e.g., max pooling filter, L2 norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 104. In Figure 1 the example shown in Figure 1 , three pooling filters are used for the three activation maps in the convolutional hidden layer 104.

[0052] In some examples, max pooling can be used by applying a max pooling filter with a stride amount (e.g., equal to the dimension of the filter, such as a stride amount of 2) and a size of, e.g., 2x2, to the activation map output from the convolutional hidden layer 104. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2×2 nodes in the previous layer (where each node is a value in the activation map). For example, four values (nodes) in the activation map will be analyzed by the 2x2 max pooling filter in each iteration of the filter, and the maximum value among these four values is output as the "max" value. If such a max pooling filter is applied to the activation filter from the convolutional hidden layer 104 with a dimension of 24×24 nodes, the output from the pooling hidden layer 106 will be an array of 12×12 nodes.

[0053] In some examples, an L2 norm pooling filter can also be used. The L2 norm pooling filter includes: calculating the square root of the sum of the squares of the values in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as in max pooling), and using the calculated value as the output.

[0054] Intuitively, the pooling function (e.g., max pooling, L2 norm pooling, or other pooling functions) determines whether a given feature is found anywhere in a region of the image and discards the exact location information. This can be done without affecting the feature detection result because once a feature is found, its exact location is less important than its approximate location relative to other features. The benefit of max pooling (and other pooling methods) is that the pooled features are much fewer, thus reducing the number of parameters required in the subsequent layers of the CNN100.

[0055] The last layer of connections in the network is the fully connected layer, which connects each node in the pooling hidden layer 106 to each output node in the output layer 110. Using the above example, the input layer includes 28×28 nodes, which encode the pixel intensities of the input image; the convolutional hidden layer 104 includes 3×24×24 hidden feature nodes, which are based on applying a 5×5 local receptive field (for the filter) to three activation maps; and the pooling hidden layer 106 includes a layer of 3×12×12 hidden feature nodes, which are based on applying a max pooling filter to a 2×2 region across each of the three feature maps. Expanding this example, the output layer 110 can include ten output nodes. In such examples, each node in the 3x12x12 pooling hidden layer 106 is connected to each node in the output layer 110.

[0056] The fully connected layer 108 can obtain the output of the previous pooling layer 106 (which should represent the activation map of high-level features) and determine the features most relevant to a specific category. For example, the fully connected layer 108 can determine the high-level features most strongly related to a specific category and can include the weights (nodes) of the high-level features. The product between the weights of the fully connected layer 108 and the pooling hidden layer 106 can be calculated to obtain the probabilities of different categories. For example, if the CNN 100 is used to predict that the object in the video frame is a person, higher values representing high-level features of a person (e.g., the presence of two legs, the presence of a face at the top of the object, the presence of two eyes in the upper left and upper right of the face, the presence of a nose in the middle of the face, the presence of a mouth at the bottom of the face, and / or other features common to people) will appear in the activation map.

[0057] In some examples, the output from the output layer 110 can include an M-dimensional vector (in existing examples, M = 10), where M can include the number of categories from which the program must choose when classifying an object in an image. Other example outputs can also be provided. Each number in the N-dimensional vector can represent the probability that the object belongs to a specific category. In an illustrative example, if the 10-dimensional output vector representing objects of ten different categories is [0 0 0.05 0.8 0 0.15 0 0 0 0], then this vector indicates that there is a 5% probability that the image is an object of the third category (e.g., a dog), an 80% probability that the image is an object of the fourth category (e.g., a person), and a 15% probability that the image is an object of the sixth category (e.g., a kangaroo). The probability for a category can be considered as the confidence level that the object is part of that category.

[0058] One problem with using convolutional layers to extract the backbone of high-level features (such as Figure 1 the CNN 100 therein) is that the receptive field is limited by the size of the convolutional kernel. For example, convolution cannot extract global information. As mentioned above, transformers can be used to extract global information, but they are computationally expensive and thus add significant latency when performing object classification and / or detection tasks.

[0059] Figure 2A classification neural network system 200 including a Mobile Vision Transformer (MobileViT) block 226 is described. The MobileViT block 226 builds on the initial application of transformers in language processing and applies that technology to images. Transformers for image processing measure the relationships between input tokens or pixels as the basic unit of analysis. However, computing the relationships for every pixel pair in a typical image is costly in terms of memory and computation. MobileViT 2-2 computes the relationships between pixels in various small parts of an image (e.g., 16x16 pixels) at a significantly reduced cost. These parts (along with positional embeddings) are placed in a sequence. The embeddings are learnable vectors. Each part is arranged in a linear sequence and multiplied by an embedding matrix. The result, along with the positional embeddings, is fed into a transformer 208.

[0060] An input image 214 having a height H, a width W, and a number of channels (e.g., H*W pixels with 3 channels corresponding to red, green, and blue components) is provided to a convolutional block 216. The convolutional block 216 applies a 3x3 convolutional kernel (with a stride amount or stride of 2) to the input image 214. The output of the convolutional block 216 passes through a number of MobileNetv2 (MV2) blocks 218, 220, 222, 224 to generate a downsampled output feature set 204 having a height H, a width W, and a dimension C. Each MobileNetv2 218, 220, 222, 224 is a feature extractor for extracting features from the output of the previous layer. In some cases, other feature extractors other than MobileNetv2 blocks may be used. The blocks that perform downsampling are marked with ↓2 (corresponding to a stride amount or stride of 2).

[0061] As Figure 2 shown, the MobileViT block 226 is described in more detail compared to other MobileViT blocks 230 and 234 of the classification neural network system 200. As shown, the MobileViT block 226 can use a convolutional layer to process the features 204 to generate a local representation 206. A convolutional transformer can produce a global representation by unfolding the data or feature set (having dimensions of H, W, d), performing a transformation using the transformer, and folding the data back again to produce another feature set (having dimensions of H, W, d) which is then output to a fusion layer 210. The MobileViT block replaces the local processing of convolutional operations with global processing using a transformer.

[0062] The fusion layer 210 can fuse data or compare it with the original data 204 to generate the output feature Y 212. The feature Y 212 output from the MobileViT block 226 can be processed by the MobileNetv2 block 228, followed by another application of the MobileViT block 230, followed by another application of the MobileNetv2 block 232 and another application of the MobileViT block 234. Subsequently, a 1x1 convolutional layer 236 (e.g., a fully connected layer) can be applied to generate the output global pool 238. Note that the downsampling of data across various block operations results in an image size of 128x128 being obtained at block 216 and a 1x1 output being produced at block 238, which includes the global pool of linear data.

[0063] Figure 3 FIG. 300 is a diagram illustrating an example of an existing tokenization process. In one example, the input image 214 can have an image size of 384 pixels. The image 214 can be partitioned into patches or tokens of different scales. In one example, the patch size 302 is 8, which results in generating 2048 tokens to cover the entire image. The patch size can refer to the number of pixels in the corresponding patch. In one aspect, the patch size can refer to one side of a square patch (e.g., for a patch with a resolution of 16x16 pixels, the patch size = 16). Representation 304 is not to scale. In another representation 304, the patch size or scale can be 32, which can result in generating 100 tokens. Typically, a transformer or component divides the image 214 into square patches or tokens of the same size, which are subsequently used as the basic units for inference or determining what objects are in the image 214. The transformer performs multiply-accumulate calculations (MACs) that are dependent on the number of input tokens, and thus the more input tokens, the higher the cost. As previously mentioned, the number of tokens can increase quadratically with the image size, which results in an inefficient model.

[0064] Figure 4 FIG. is an example of a system 400 including a preprocessing engine 402 for generating hybrid-scale (or hybrid-resolution) tokens for a transformer 226 according to the systems and techniques described herein. In some cases, Figure 4 One or more components described in Figure 2 are similar to those shown in Figure 4 The preprocessing engine 402 of Figure 2 can replace the components 216, 218, 220, 222, 224 shown in Figure 4 or can appear in series with these components at any position. In the discussion of Figure 5 .

[0065] Figure 5It is a diagram showing an example of a high-level overview 500 of generating mixed-scale tokens. Generally speaking, it can be beneficial to position the preprocessing engine 402 early in the processing of the input image 214. Positioning the preprocessing engine 402 "early" can mean that the first layer of the network is the preprocessing engine 402, or the preprocessing engine 402 can be used as one of the first set of input layers of the network. In some cases, the preprocessing engine 402 may include a combination of neural network layers (such as convolutional layers, linear layers, non-linear layers, max pooling layers, and softmax layers). Within the preprocessing engine 402, the type and number of layers can vary. The overall processing is described, but it can also be implemented through various different hierarchical structures.

[0066] The preprocessing engine 402 may include a first engine 404 that divides the input image 214 into at least two small blocks or token sets at multiple resolutions. For example, the first engine 404 can divide the input image 214 into a first token set at a first resolution, which can be a coarse resolution with relatively large small blocks or tokens (such as the small block 502 shown in Figure 5 ). The first engine 404 can divide the input image 214 into a second token set at a second resolution, which can be a fine resolution with relatively small small blocks or tokens (such as the small block 504 shown in Figure 5 ). In an illustrative example, the input image 214 can be chunked into coarse image regions of size Sc x Sc. Although two different resolutions are shown, additional small block sets can also be created at finer or coarser resolutions. Note that the definition of fine or coarse can be relative to each other or other resolutions.

[0067] Each coarse region can be processed by a gate 408 (e.g., shown as gate 508 in Figure 5 ), which can be labeled as gate g. In some cases, the gates 408, 508 may include a 4-layer multi-layer perceptron (MLP). The gates 408, 508 can output a binary decision on whether the region is to be further processed at a coarse scale or a fine scale (e.g., in the form of a mask m, which can be generated by a mask engine 410 in some cases). The resulting mask m defines a set of mixed-scale tokens for the input image. When needed, the corresponding mixed-scale position encoding can be obtained by linearly interpolating the fine-scale position encoding to the coarse scale. The tokens can then be sent to a transformer 226, which can be a standard transformer backbone T. The transformer 226 can output task-related predictions.

[0068] The embedding engine 406 can perform an operation of embedding each token that covers a specific region of the input image 214. For example, the embedding engine 406 can process the tokens in each region of the input image 214 (e.g., using one or more convolutional layers, activation layers, pooling layers, etc.) to generate a feature vector for each token. Coarse tokens can be grouped with finer-scale tokens corresponding to a common region of the input image 214. For example, as Figure 5 shown, a first coarse token 506 corresponding to a region of the input image 214 is grouped with four fine tokens 507 that also correspond to the same region of the input image 214. For example, the portion of the input image 214 covered by the coarse token 506 is collocated with the portion of the input image 214 covered by the four fine tokens 507. In some cases, tokens of different resolutions may not be perfectly collocated. For example, compared to the region of the input image 214 covered by the coarse token 506, the four fine tokens 507 can cover a slightly larger region of the input image 214 (referred to as "overflow"). The presence or absence of overflow can be based on several different factors, such as the characteristics of the image portion (smooth and consistent or complex), the desired level of accuracy, the desired amount of computation, and / or other factors.

[0069] In an illustrative example, as Figure 5 shown, a factor of 4 can be used for the transition from a coarse resolution (one token for a portion of the image) to a fine resolution (e.g., four tokens for the same portion of the image or at least a partially overlapping portion of the image). Other factors can also be used. Additionally, while the example shape of a token is square, other shapes are also contemplated. In one example, the shape for a resolution (or different shapes for resolutions) can be selected based on a similar model such as a gate 408 (shown as gate 508 in Figure 5 ), which can also be referred to as a gating engine that can be trained on how to output a token shape based on the characteristics of the data.

[0070] The embedded tokens 510, 512 can be combined (e.g., concatenated) and input as an input vector (or other representation) into the gate 408. The gate 408 can perform a binary operation to determine whether to use a fine scale or a coarse scale for a particular region or portion of the input image 214. In one example, the gate 408 can process an image portion related to a coarse resolution and determine whether that image portion is simple, complex, or has some other characteristic that leads to a binary output decision to represent that portion at a coarse resolution or a fine resolution. For example, if the portion of the input image 214 covered by the concatenated tokens is consistent across that portion (e.g., includes pixels associated with a blue sky), a coarse token can be assigned to represent that portion of the image. On the other hand, if the region is detailed and has different colors or shapes (e.g., includes pixels corresponding to one or more persons), the gate 408 can determine that the region should be represented by a set of fine tokens 512.

[0071] In some cases, the gate 408 can include a softmax layer that can output a softmax distribution at each scale (e.g., coarse scale and fine scale). In one aspect, the gate 408 can include multiple layers of the network, such as one or more linear layers, followed by a softmax layer that evaluates the input vector (including the embedded coarse token 510 and the embedded fine token 512) to determine which token or set of tokens at which resolution should represent a particular portion or region of the input image 214. More flexible training constraints can be applied to the gate 408, which enhances more diverse dynamic patterns, as described in more detail below with respect to Figure 9 In some cases, the gate 408 can represent multiple trained gates, where each corresponding gate operates on a spatial location within the input image. In some cases, the gate 408 can also be trained to select from more than two resolution options. Embeddings for the entire image can be provided to the gate 408 via the input vector such that information about how the particular portion being processed relates to the entire image is available for determining the resolution selection.

[0072] In some aspects, the gate 408 can receive other information as input. For example, additional information can be combined (e.g., concatenated) with the embedded vector of the token to generate an input vector provided to the gate 408. For example, the two-dimensional position of the token in the input image 214 (or the position of a portion of the input image 214 that includes the token) can also be provided to the gate 408 in the input vector such that the portion can be evaluated in the context of the position in the larger image. In an illustrative example, when the portion being processed in the input image is in the middle, or at a corner or edge, then the corresponding position can affect the output resolution decision. In some cases, it may be more efficient to use a finer resolution for the portion in the middle of the image 214 and a coarser resolution for the edge or corner of the image 314. The gate 408 can be trained using one or more of various types of input data for its evaluation.

[0073] In some cases, various designs can be used for the gate 408. In one aspect, the gate 408 can receive the concatenated embeddings of tokens at all scales in a region, which can be useful for lossless token reduction. In such aspects, the preprocessing engine 402 can embed each token covering the region and can feed the concatenated embeddings to the gate 408, which can output the softmax distribution at each scale. As a result, each pixel in the image is represented by one token at the scale selected by the gate and thus no information is lost. In another aspect, the gate 408 can take individual tokens as input, where the input to the gate 408 can be individual tokens provided serially rather than a set of concatenated tokens. Such a solution may not be lossless but can allow for the generation of a unified model covering both token pruning and dynamic scale selection. Such a solution can be useful in global tasks such as image classification. In such cases, the gate 408 can receive a coarse-scale token and determine whether to discard or retain the token (as a binary decision) rather than receiving the concatenated embedding (with the coarse-scale token plus the associated fine-scale tokens). Subsequently, the gate 408 can receive a fine-scale token and determine whether to discard or retain the token (as a binary decision).

[0074] In some aspects, the Transformer 226 model is fully shared across various scales. In some examples, the system can use random crop data augmentation, which can help the initial token linear embedding to be robust to scale changes. In some cases, position encoding indicating the two-dimensional spatial position information of each token can be added as input to the Transformer 226. In some cases, to share the position encoding across scales, the system can linearly interpolate the two-dimensional positions to match each given scale, allowing the position encoding parameters to be effectively shared. For example, the system can linearly interpolate the position encodings of four fine-scale tokens to generate a coarse-scale position encoding for the corresponding covering coarse-scale token.

[0075] The output of gate 408 is then provided to a masking engine 410 which masks the input patches or tokens based on the data received from gate 408. As Figure 5 shown, the masking provides a coarse token 514, 516 that covers the corresponding region for each region of the input image 214 or a set of fine tokens 518, 520 that cover the corresponding region. In Figure 5 the example, the set of masked tokens can cover all parts of the input image 214 using the set of fine tokens or the coarse token. The set of masked tokens is provided to the transformer 226 which can further process the image.

[0076] Network 500 can generally be defined as two different blocks, including a preprocessing engine 402 as a first block and a transformer 226 as a second block. In some cases, the input layer can be considered as a unified model embodied as the preprocessing engine 402 and before the transformer 226. The preprocessing engine 402 enables the transformer 226 to be the second layer capable of handling input tokens from multiple scales at once. The masking engine 410 feeds the masked tokens to the transformer 226. Using such a structure, a traditional transformer architecture (without any changes) can be used for the transformer 226.

[0077] Figure 6 FIG. 600 is a diagram that illustrates the differences between the prior transformer processes 602, 608 and the hybrid scale method 612 disclosed herein. In the standard vision transformer method 602, there is a fixed scale for all images and tokens 604. The embedding of the patches involves providing positional encoding to each patch or token and providing the result to the vision transformer 226. In the modified vision transformer method 608, a per-image fixed scale is used for all tokens plus one model per scale. Thus, the system can select a first set of tokens 604 at a first resolution, embed 606 the positional encoding of the tokens 604, and use a first vision transformer 226 to process the tokens 604. Another set of tokens 610 at a different resolution can be embedded 606 with positional encoding and provided to a second vision transformer 226 to process the set of tokens 610 at the second resolution.

[0078] In contrast, the method 612 disclosed herein performs per-image and per-token scaling using a single model, where starting from an initial set of tokens 604 at a first resolution, another set of tokens at co-located positions and different scales can be used to generate a set of masked tokens, where some tokens 614 are at the first resolution and another set of tokens 616 are at a second (possibly finer) resolution. Such a unified hybrid-scale model can have the benefits of being parameter-efficient in terms of memory savings and data-efficient in terms of being easier to train. In this case, the embedding process 606 includes linear interpolation of two-dimensional positions as part of the positional encoding (not provided in previous methods), such that the positional encoding is shared across different scales. This method enables the unified hybrid-scale model 226, which is a vision transformer, to handle hybrid-scale inputs 612 more parameter-efficiently and data-efficiently.

[0079] In one scenario, the preprocessing engine 402 can select a scale, learn the positional encoding for one scale (e.g., the fine scale), and subsequently obtain the other scale (e.g., the coarse scale) by linear interpolation from one scale to the other. Subsequently, the result can be fed into the transformer 226. The transformer 226 is aware of each scale for each token. The transformer 226 will also be aware of the two-dimensional position and scale for each token. The transformer 226 can be aware of how to accommodate different resolutions of tokens.

[0080] The linear interpolation can operate as follows. If the two-dimensional position of a fine-scale token (such as token 616) is known, but the resolution selected for that part of the image is the coarse resolution, and thus there is a coarse-resolution token at that position instead of four fine-resolution tokens, the system can take the average of their positions or the linear interpolation of the two-dimensional positions of the four fine-resolution tokens, and this can represent the position of the coarse token at that place. This can occur in a situation where the system has learned a vector for each position of the fine-resolution tokens. There is a vector for each of those cells. To obtain the coarse-token orientation, the system can average the four corresponding vectors that have been learned for the four associated fine-resolution tokens. The new identifier of the coarse token inherently has some information associated with the coarse resolution. It can provide more information in an integrated manner.

[0081] In one example, only the fine-scale resolution tokens are embedded with positional vectors, and when the gate 408 selects the coarse resolution for a part of the image, the system will always use the fine-scale resolution token positional vectors to interpolate the position of the coarse-scale tokens.

[0082] This method is applicable to an example structure where four fine-resolution tokens 616 can be co-located on a coarse-resolution token. If the different-resolution tokens are not co-located, or if the relationship between different resolutions is more complex than the example framework, other methods can be implemented to obtain the appropriate position of the coarse-level token based on two or more fine-resolution tokens having some association with the coarse-resolution token (such as at least partial overlap of a part of the input image).

[0083] Figure 7 FIG. 700 is a diagram illustrating an example of the existing merging 702 and pruning 704 processes compared to the dynamic method 706 disclosed herein. In the token merging method 702, a unique token downsampling projection is performed using a fixed number of tokens. These methods are problematic, and the benefit of the disclosed method lies in the unique token downsampling layer 706. As Figure 7 shown, in the merging method 702, the transformer 708 can output M learned tokens based on N input tokens. This method can be performed by known software tools such as PathChMerger, Perceiver, and TokenLearner. It is typically located in the middle of the network architecture. For example, for PatChmerger and TokenLearner, it is typically in the middle of the architecture and has a very small number of output tokens M (e.g., M = 16). For Perceiver, it is typically at the beginning of the architecture (and sometimes even repeated in the middle of the architecture), but the actually learned number of tokens M is quite large (e.g., M = 512), so Perceiver is not very competitive in terms of efficiency compared to other efficient transformers.

[0084] The first half of the network remains inefficient because it is processing N tokens until the token merging 702 occurs to reduce the number to M tokens. The transformer 708 outputs M tokens as the learned projection directions, and the number is less than the N input tokens. Note that the transformation or merging of tokens occurs approximately in the middle of the processing by the transformer 708, as indicated by its shape.

[0085] The token pruning method 704 is also shown, where slow iterative token reduction is performed with a fixed token pruning ratio. For example, the tokens can be ranked by (decreasing) category attention (attention regarding special category tokens), and then the transformer 710 can periodically prune a certain percentage of the low-ranked tokens, as indicated by the shape of the transformer 710. This method typically prunes a fixed x percentage of the tokens in each operation of the transformer 710 and is thus not very flexible.

[0086] This method 706 includes a unique token downsampling layer that utilizes a dynamic number of tokens. The preprocessing engine 402 can utilize a lightweight preprocessing engine to select tokens of a dynamic scale for each spatial location. Note the shape of the rectangle 712 associated with the transformer 226. In this case, there are a small number of tokens throughout the process (the number of tokens does not change, or is pruned or merged). The image 716 therein can be simpler and does not require as many tokens to process. The shape of the rectangle 714 associated with the transformer 226 represents the concept that more tokens are needed for a more complex image 718. The number of tokens remains the same throughout the transformer processing (which is also represented by the shape of the rectangles 712, 714).

[0087] The transformer 226 can be a lightweight preprocessing module. The transformer 226 can be a unified model because it can handle tokens of different scales in the same input set as shown. This process selects tokens of a dynamic scale for each spatial location. Benefits of this approach include early token reduction, plus task-agnostic gates can be easily transferred to only one routing decision for any transformer model. In some cases, the preprocessing engine 402 is configured to process the input image as one of the first layers or the first layer in the overall model. Thus, there is no need to prune or merge tokens at the various stages of processing in the network, and the number of tokens can remain the same throughout the network.

[0088] Figure 8 FIG. 800 is a diagram illustrating a comparison of a fixed-scale token and a dynamic hybrid-scale pattern in accordance with aspects of the present disclosure. For example, in the fixed-scale method 802, there is a fixed number of tokens for all images. In the example fixed-scale method 802, there are 64 tokens for representing an image. In the present method 804, the scale associated with the tokens for each region is determined. In the example 804, some tokens (such as token 806) are larger due to the characteristics of the image in that region, while other tokens (such as token 808) are relatively smaller. In the image 810, the token 812 is inherently larger or coarser, while the token 814 is relatively smaller and covers a fine resolution.

[0089] Figure 9FIG. 0 is a diagram illustrating a binary gating decision process for each spatial location and each input image in accordance with aspects of the present disclosure. The transparent region 902 represents the coarse-scale resolution of the region, while the filled region 904 represents the fine-scale of the region. The horizontal axis represents the spatial location of the region, and the vertical axis represents the input images, where each row is associated with a separate corresponding input image. Token selection by the gate 408 occurs across these two dimensions. For each corresponding image and each corresponding region of that image, a binary decision is made. The distribution of the binary decisions across the diagram 900 can inform the model how to condition. The idea is to ensure that different patterns exist across different images. If the patterns across different images are the same, the system does not introduce efficiency because the system can simply prune some images to reduce computation. Generally, enhanced conditioning methods act on both the input image dimension and the spatial location dimension such that each image has its own gating behavior response and each token also has its own gating behavior response. Other training loss methods are inferior to this novel method. For example, an average-constrained training loss (L0 loss) (which sets a value such as expecting half of the tokens to be at the coarse scale and the other half of the tokens to be at the fine scale) can result in unconditional target sparsity. This method only constrains the average of the distribution. A batch-shaping loss across the input image dimension is also not desirable because each spatial location has the same sparsity pattern. There is per-image conditioning, but in the batch-shaping loss method, it is only across one dimension (per image), rather than for each spatial location. Here, the system constrains the distribution along the samples to follow a beta prior with mean μ.

[0090] In one example, the mean μ can be a parameter learned per model independent of each spatial location, thus encouraging conditioning across spatial locations. In contrast, the batch-shaping loss would assume the mean μ is the same for all spatial locations.

[0091] Conditioning means that if two different images are input into the model, the model should output two different behaviors or masks conditioned on the characteristics of the corresponding images. Such enhanced conditioning is achieved in a method where the binary gating decision (training of the gates) is for each spatial location and for each input image. In one aspect, the disclosed method can be characterized as hyperprior training, where the system constrains each token-dependent distribution to follow a different learned prior. The learned parameters are controlled by the hyperprior (e.g., but not limited to a Gaussian N, whose flexibility depends on an additional variance hyperparameter σ). Given a data training batch with N samples and d input tokens, the aggregated output of all the gates can be labeled as G, a matrix of size (N, d). The gate 408 can be trained to match a certain target sparsity μ.

[0092] In one example, the spatial location 908 highlighted across images is always selected to be at a coarse scale. In this case, the entropy of this spatial location is low, and the model wastes training capacity, which may lead to lower accuracy. The system can determine that it does not need to learn the gate for this spatial location because the token is always off for this spatial location. The gate 408 should be trained for each image, and Figure 9 the use of the information in can help with how to train the model and provide enhanced conditionality to improve efficiency by possibly avoiding training the gate for spatial locations with low entropy across images.

[0093] Using the learned conditional gate 408, the model matches the static hybrid resolution accuracy but has the benefit of allowing a dynamic computational cost per image. This method can achieve a reduction in computational cost without sacrificing accuracy.

[0094] Figure 10 An example method 1000 for processing image data is illustrated (such as using a preprocessing engine 402 (e.g., as described with respect to Figure 4 ). At block 1002, the method 1000 may include partitioning an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution. In some cases, the first token representation set includes a single token representation according to the first resolution, and the second token representation set includes multiple token representations according to the second resolution. In some cases, the first resolution or the second resolution may be determined as the scale for a first region of the input image based on one or more characteristics of the input image. In some cases, the one or more characteristics of the input image may include a smoothness value associated with the first region of the input image, a complexity value associated with the first region of the input image, how many colors are associated with the input image, or a contrast value associated with the first region of the input image. In some cases, the input image includes image patches of the image.

[0095] At block 1004, the method 1000 may include generating a first set of token representations for one or more tokens from the first set of tokens corresponding to the first region of the input image. In some cases, generating the first set of token representations may include processing the first set of tokens using a linear neural network layer to generate a first set of embedding vectors.

[0096] At block 1006, the method 1000 may include generating a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image. In some cases, generating the second set of token representations may include processing the second set of tokens using a linear neural network layer to generate a second set of embedding vectors.

[0097] At block 1008, method 1000 may include processing a first set of token representations and a second set of token representations using a neural network model to determine a first resolution or a second resolution as the scale of a first region of the input image. In some cases, the neural network model is shared across regions of the input image. In some cases, the neural network model may include a Softmax layer configured to determine a distribution over the first resolution and the second resolution.

[0098] At block 1010, method 1000 may include processing the first region using a Transformer neural network model according to the scale of the first region of the input image. In some cases, the Transformer neural network model may be configured to adaptively blend resolution data based on masked processing.

[0099] In some cases, method 1000 may further concatenate the first set of token representations and the second set of token representations to generate a concatenated set of token representations. Processing of the first set of token representations and the second set of token representations may include processing the concatenated set of token representations using a neural network model to determine a first resolution or a second resolution as the scale of a first region of the input image.

[0100] In some cases, method 1000 may further include determining a respective scale for each respective region of the input image and / or determining a respective positional encoding for each region of the input image.

[0101] In some cases, for a region of the input image determined to have a scale corresponding to the second resolution, method 1000 may include: determining that the respective positional encoding includes determining a final positional encoding of the region as a linear interpolation of a plurality of initial positional encodings determined for the region.

[0102] In some cases, method 1000 may further include generating a mask for the input image, the mask indicating the respective scale determined to be the first resolution or the second resolution for each respective region of the input image.

[0103] In some examples, the processes described herein (e.g., method 1000 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, method 1000 may be performed by Figure 11 the computing system 1100 shown in

[0104] A computing device can include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or smartwatch, or other wearable device), a server computer, a computing device of an autonomous vehicle or an autonomous vehicle, a robotic device, a television, and / or any other computing device having the resource capabilities to perform the processes described herein (including method 500 and / or other processes described herein). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on Internet Protocol (IP) or other types of data.

[0105] The components of a computing device can be implemented in circuitry. For example, the components may include and / or may use electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits)), and / or may include and / or may use computer software, firmware, or any combination thereof, to perform the various operations described herein.

[0106] Method 1000 is illustrated as a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. In general, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined and / or performed in parallel in any order to implement the processes.

[0107] Additionally, method 1000 and / or other processes described herein can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, implemented by hardware, or a combination thereof. As mentioned above, the code can be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0108] The methods disclosed herein include combinations of features not found in other methods. The disclosed features can include one or more hybrid scales for a set of tokens and can include dynamic features where the number of tokens in a masked set of tokens can vary for an image. The method is lossless because each part of the image is represented by one or more tokens at a given scale. Thus, all parts of the image are represented and there is no loss. The method also provides an efficient feed-forward network (FFN) in a transformer. There is a reduced number of MAC operation counts in the FFN. No other method has all these features. Applications of this method can include, but are not limited to, image classification in this context. Any vision task can benefit from these concepts, especially intensive tasks such as segmentation. Any vision transformer can utilize this preprocessing engine 402 disclosed herein. Other uses can include real-time video processing, extended reality, or any other vision processing.

[0109] Figure 11 is a diagram illustrating an example of a system for implementing certain aspects of the techniques herein. Specifically, Figure 11 illustrates an example of a computing system 1100, which can be, for example, any computing device, remote computing system, camera, or any component thereof that constitutes an internal computing system, where the components of the system are in communication with each other using connection 1105. Connection 1105 can be a physical connection using a bus or a direct connection to a processor 1110 (such as in a chipset architecture). Connection 1105 can also be a virtual connection, a networking connection, or a logical connection.

[0110] In some aspects, computing system 1100 is a distributed system where the functions described in this disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, and so on. In some aspects, one or more of the described system components represent many such components, each component performing some or all of the functions described for that component. In some aspects, the components can be physical or virtual devices.

[0111] Example system 1100 includes at least one processing unit (CPU or processor) 1110 and a connection 1105 that couples various system components including system memory 1115 (such as read only memory (ROM) 1120 and random access memory (RAM) 1125) to the processor 1110. Computing system 1100 may include a cache 1111 of high-speed memory that is directly connected to, adjacent to, or integrated as part of the processor 1110.

[0112] The processor 1110 may include any general-purpose processor and hardware services or software services such as services 1132, 1134, and 1136 stored in the storage device 1130 and configured to control the processor 1110, as well as a dedicated processor, where software instructions are incorporated into the actual processor design. The processor 1110 may be substantially a fully self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. A multi-core processor may be symmetric or asymmetric.

[0113] To enable user interaction, computing system 1100 includes an input device 1145 that may represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, and so on. Computing system 1100 may also include an output device 1135, which may be one or more of several output mechanisms. In some instances, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 1100. Computing system 1100 may include a communication interface 1140, which generally may manage and control user input and system output.

[0114] The communication interface may execute or facilitate receiving and / or transmitting wired or wireless communications using a wired and / or wireless transceiver, including using an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, port / plug, an Ethernet port / plug, a fiber optic port / plug, a dedicated wired port / plug, wireless signal transmission, low energy (BLE) wireless signal transmission, Those communications that are wireless signal transmissions, radio frequency identification (RFID) wireless signal transmissions, near field communication (NFC) wireless signal transmissions, dedicated short range communication (DSRC) wireless signal transmissions, 802.11 Wi-Fi wireless signal transmissions, WLAN signal transmissions, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmissions, public switched telephone network (PSTN) signal transmissions, integrated services digital network (ISDN) signal transmissions, 3G / 4G / 5G / long term evolution (LTE) cellular data network wireless signal transmissions, ad hoc network signal transmissions, radio wave signal transmissions, microwave signal transmissions, infrared signal transmissions, visible light signal transmissions, ultraviolet light signal transmissions, wireless signal transmissions along the electromagnetic spectrum, or some combination thereof.

[0115] Communication interface 1140 may also include one or more GNSS receivers or transceivers that are used to determine the location of computing system 1100 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States-based Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0116] Storage device 1130 can be a non-volatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as cassette tapes, flash memory cards, solid state memory devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes / magnetic strips, any other magnetic storage medium, flash memory, memristor memory, any other solid state memory, compact disc read-only memory (CD-ROM) optical discs, rewritable compact discs (CD) optical discs, digital video discs (DVD) optical discs, Blu-ray discs (BDD) optical discs, holographic optical discs, another optical medium, secure digital (SD) cards, micro secure digital (microSD) cards, Cards, smart card chips, Europay, MasterCard and Visa (EMV) chips, subscriber identity module (SIM) cards, mini / micro / nano / pico SIM cards, other integrated circuit (IC) chips / cards, RAM, static RAM (SRAM), dynamic RAM (DRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), other memory chips or cartridges, and / or combinations thereof.

[0117] The storage device 1130 may include software services, servers, services, etc., which, when the code defining such software is executed by the processor 1110, cause the system to perform functions. In some aspects, the hardware services that perform specific functions may include software components stored in a computer-readable medium connected to the necessary hardware components (such as the processor 1110, connection 1105, output device 1135, etc.) to perform functions. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include a non-transitory medium in which data can be stored and which does not include carrier waves and / or transient electrical signals propagated wirelessly or through a wired connection.

[0118] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include a non-transitory medium in which data can be stored and which does not include carrier waves and / or transient electrical signals propagated wirelessly or through a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a procedure, function, subroutine, program, routine, subroutine, module, engine, software package, class, or any combination of instructions, data structures, or program statements. Code segments may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted by any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0119] In some aspects, a computer-readable storage device, medium, and memory may include a wire or wireless signal that includes a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0120] Specific details are provided in the above description to provide an exhaustive understanding of the aspects and examples provided herein. However, one of ordinary skill in the art will understand that these aspects may be practiced without these specific details. For clarity, in some instances, the techniques of the present invention may be presented as including various functional blocks that include devices, device components, steps or routines in a method implemented in software or a combination of hardware and software. Additional components other than those shown and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the aspects.

[0121] Individual aspects may be described above as a process or method, which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or concurrently. Additionally, the order of the operations may be rearranged. A process terminates when its operations are complete, but a process may have additional steps not included in the figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0122] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code. Examples of computer-readable media that may be used to store instructions, information used during the methods according to the described examples, and / or information created include magnetic or optical disks, flash memory, USB devices providing non-volatile memory, networked storage devices, etc.

[0123] Devices implementing the various processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take on any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may execute the necessary tasks. Typical examples of form factors include: laptop devices, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be implemented with peripheral devices or plug-in cards. As a further example, such functionality may also be implemented on a circuit board among different chips or different processes executing on a single device.

[0124] Instructions, a medium for conveying these instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0125] In the foregoing description, aspects of the present application are described with reference to their specific aspects, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the illustrative aspects of the present application have been described in detail herein, it is to be understood that the various inventive concepts may be implemented and employed in other various ways, and the appended claims are not to be construed as including such variations, unless limited by the prior art. The various features and aspects of the foregoing application may be used singly or in combination. In addition, the aspects may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the drawings are to be regarded as illustrative rather than limiting. For purposes of illustration, the methods are described in a particular order. It should be appreciated that in alternative aspects, the methods may be performed in a different order than that described.

[0126] Those of ordinary skill in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this specification.

[0127] In cases where components are described as “configured to” perform certain operations, such configuration may be implemented, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor, or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0128] The phrase "coupled to" means that any component is physically connected to another component directly or indirectly, and / or any component is in communication with another component directly or indirectly (e.g., connected to the other component through a wired or wireless connection and / or other suitable communication interface).

[0129] In the claims language or other language of the present disclosure that recites "at least one" in a set and / or "one or more" in a set, such language indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claims language that recites "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claims language that recites "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of "at least one" in a set and / or "one or more" in a set does not limit the set to the items listed in the set. For example, the claims language that recites "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0130] The various illustrative logical blocks, modules, engines, circuits, and algorithmic steps described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, engines, circuits, and steps are described above in terms of their functionality in a generalized form. Whether such functionality is implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of this application.

[0131] The techniques described herein may also be implemented using electronic hardware, computer software, firmware, or any combination thereof. These techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a handheld wireless communication device, or an integrated circuit device with multiple uses, including applications in handheld wireless communication devices and other devices. Any features described as modules, engines, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the above-described methods, algorithms, and / or operations. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. These techniques may additionally or alternatively be at least partially implemented by a computer-readable communication medium carrying or conveying program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0132] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, as used herein, the term “processor” may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0133] Exemplary aspects of the present disclosure include:

[0134] Aspect 1. A processor-implemented method for processing image data, the method comprising: dividing an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generating a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generating a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; processing the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and processing the first region using a transformer neural network model according to the scale of the first region of the input image.

[0135] Aspect 2. The processor-implemented method according to Aspect 1, wherein: generating the first set of token representations includes processing the first set of tokens using a linear neural network layer to generate a first set of embedding vectors; and generating the second set of token representations includes processing the second set of tokens using a linear neural network layer to generate a second set of embedding vectors.

[0136] Aspect 3. The processor-implemented method according to any one of Aspect 1 or 2, wherein the first set of token representations includes a single token representation according to the first resolution, and wherein the second set of token representations includes a plurality of token representations according to the second resolution.

[0137] Aspect 4. The processor-implemented method according to any one of Aspects 1 to 3, further comprising: concatenating the first set of token representations and the second set of token representations to generate a concatenated set of token representations; wherein processing the first set of token representations and the second set of token representations includes processing the concatenated set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image.

[0138] Aspect 5. The processor-implemented method according to any one of Aspects 1 to 4, further comprising: determining a corresponding scale for each corresponding region of the input image.

[0139] Aspect 6. The processor-implemented method according to any one of Aspects 1 to 5, further comprising: determining a corresponding position encoding for each region of the input image.

[0140] Aspect 7. The processor-implemented method according to any one of Aspects 1 to 6, wherein for a region of the input image determined to have a scale corresponding to the second resolution, determining the corresponding position encoding includes determining the final position encoding of the region as a linear interpolation of a plurality of initial position encodings determined for the region.

[0141] Aspect 8. The processor-implemented method as in any one of Aspects 1 to 7 further includes: generating a mask for the input image, the mask indicating the respective scales determined for each corresponding region of the input image as the first resolution or the second resolution.

[0142] Aspect 9. The processor-implemented method as in any one of Aspects 1 to 8, wherein the transformer neural network model is configured to process the adaptive hybrid resolution data based on the mask.

[0143] Aspect 10. The processor-implemented method as in any one of Aspects 1 to 9, wherein the neural network model is shared across regions of the input image.

[0144] Aspect 11. The processor-implemented method as in any one of Aspects 1 to 10, wherein the neural network model includes a Softmax layer configured to determine the distribution over the first resolution and the second resolution.

[0145] Aspect 12. The processor-implemented method as in any one of Aspects 1 to 11, wherein the first resolution or the second resolution is determined as the scale of the first region of the input image based on one or more characteristics of the input image.

[0146] Aspect 13. The processor-implemented method as in any one of Aspects 1 to 12, wherein the one or more characteristics of the input image include a smoothness value associated with the first region of the input image, a complexity value associated with the first region of the input image, how many colors are associated with the input image, or a contrast value associated with the first region of the input image.

[0147] Aspect 14. The processor-implemented method as in any one of Aspects 1 to 13, wherein the input image includes image patches of the image.

[0148] Aspect 15. An apparatus for processing image data includes: at least one memory; and at least one processor coupled to the at least one memory and configured to: partition an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generate a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generate a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; process the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and process the first region using a transformer neural network model according to the scale of the first region of the input image.

[0149] Aspect 16. The apparatus for processing image data as in aspect 15, wherein the at least one processor is further configured to: generate a first set of token representations including processing a first set of tokens using a linear neural network layer to generate a first set of embedding vectors; and generate a second set of token representations including processing a second set of tokens using a linear neural network layer to generate a second set of embedding vectors.

[0150] Aspect 17. The apparatus for processing image data as in aspect 15 or 16, wherein the first set of token representations includes a single token representation according to a first resolution, and wherein the second set of token representations includes a plurality of token representations according to a second resolution.

[0151] Aspect 18. The apparatus for processing image data as in any one of aspects 15 to 17, wherein the at least one processor is further configured to: concatenate the first set of token representations and the second set of token representations to generate a concatenated set of token representations; process the concatenated set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of a first region of the input image.

[0152] Aspect 19. The apparatus for processing image data as in any one of aspects 15 to 18, wherein the at least one processor is further configured to: determine a respective scale for each respective region of the input image.

[0153] Aspect 20. The apparatus for processing image data as in any one of aspects 15 to 19, wherein the at least one processor is further configured to: determine a respective position encoding for each region of the input image.

[0154] Aspect 21. The apparatus for processing image data as in any one of aspects 15 to 20, wherein the at least one processor is further configured to: for a region of the input image determined to have a scale corresponding to the second resolution, determine the respective position encoding by: determining the final position encoding of the region as a linear interpolation of a plurality of initial position encodings determined for the region.

[0155] Aspect 22. The apparatus for processing image data as in any one of aspects 15 to 21, wherein the at least one processor is further configured to: generate a mask for the input image, the mask indicating the respective scale determined to be the first resolution or the second resolution for each respective region of the input image.

[0156] Aspect 23. The apparatus for processing image data as in any one of aspects 15 to 22, wherein the transformer neural network model is configured to process adaptive mixed-resolution data based on the mask.

[0157] Aspect 24. An apparatus for processing image data as in any one of aspects 15 to 23, wherein a neural network model is shared across regions of the input image.

[0158] Aspect 25. An apparatus for processing image data as in any one of aspects 15 to 24, wherein the neural network model includes a Softmax layer configured to determine distributions at a first resolution and a second resolution.

[0159] Aspect 26. An apparatus for processing image data as in any one of aspects 15 to 25, wherein the first resolution or the second resolution is determined as a scale of a first region of the input image based on one or more characteristics of the input image.

[0160] Aspect 27. An apparatus for processing image data as in any one of aspects 15 to 26, wherein the one or more characteristics of the input image include a smoothness value associated with a first region of the input image, a complexity value associated with the first region of the input image, how many colors are associated with the input image, or a contrast value associated with the first region of the input image.

[0161] Aspect 28. An apparatus for processing image data as in any one of aspects 15 to 27, wherein the input image includes image patches of the image.

[0162] Aspect 29. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of aspects 1 to 14.

[0163] Aspect 30. A device for classifying image data on a mobile device, the device including one or more means for performing the operations according to any one of aspects 1 to 14.

Claims

1. A processor-implemented method for processing image data, the method comprising: Dividing an input image into a first set of tokens with a first resolution and a second set of tokens with a second resolution; Generating a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; Generating a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; Processing the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; And Processing the first region using a transformer neural network model according to the scale of the first region of the input image.

2. The processor-implemented method according to claim 1, wherein: Generating the first set of token representations includes processing the first set of tokens using a linear neural network layer to generate a first set of embedding vectors; and Generating the second set of token representations includes processing the second set of tokens using the linear neural network layer to generate a second set of embedding vectors.

3. The processor-implemented method according to claim 1, wherein the first set of token representations includes a single token representation according to the first resolution, and wherein the second set of token representations includes a plurality of token representations according to the second resolution.

4. The processor-implemented method according to claim 1, further comprising: Cascading the first set of token representations and the second set of token representations to generate a cascaded set of token representations; Wherein processing the first set of token representations and the second set of token representations includes processing the cascaded set of token representations using the neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image.

5. The processor-implemented method according to claim 1, further comprising: Determining a corresponding scale for each corresponding region of the input image.

6. The processor-implemented method according to claim 5, further comprising: Determining a corresponding position encoding for each region of the input image.

7. The processor-implemented method according to claim 6, wherein for a region of the input image determined to have a scale corresponding to the second resolution, determining the corresponding position encoding includes determining the final position encoding of the region as a linear interpolation of a plurality of initial position encodings determined for the region.

8. The processor-implemented method according to claim 1, further comprising: Generating a mask for the input image, the mask indicating the corresponding scale determined as the first resolution or the second resolution for each corresponding region of the input image.

9. The processor-implemented method according to claim 8, wherein the transformer neural network model is configured to process adaptive mixed-resolution data based on the mask.

10. The processor-implemented method according to claim 1, wherein the neural network model is shared across regions of the input image.

11. The processor-implemented method according to claim 1, wherein the neural network model includes a Softmax layer configured to determine distributions on the first resolution and the second resolution.

12. The processor-implemented method according to claim 1, wherein the first resolution or the second resolution is determined as the scale of the first region of the input image based on one or more characteristics of the input image.

13. The processor-implemented method according to claim 12, wherein the one or more characteristics of the input image include a smoothness value associated with the first region of the input image, a complexity value associated with the first region of the input image, how many colors are associated with the input image, or a contrast value associated with the first region of the input image.

14. The processor-implemented method according to claim 1, wherein the input image includes image patches of an image.

15. An apparatus for processing image data, comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: divide an input image into a first set of tokens having a first resolution and a second set of tokens having a second resolution; generate a first set of token representations for one or more tokens from the first set of tokens corresponding to a first region of the input image; generate a second set of token representations for one or more tokens from the second set of tokens corresponding to the first region of the input image; process the first set of token representations and the second set of token representations using a neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image; and process the first region using a transformer neural network model according to the scale of the first region of the input image.

16. The apparatus for processing image data according to claim 15, wherein the at least one processor is further configured to: generating the first set of token representations includes processing the first set of tokens using a linear neural network layer to generate a first set of embedding vectors; and generating the second set of token representations includes processing the second set of tokens using the linear neural network layer to generate a second set of embedding vectors.

17. The apparatus for processing image data according to claim 15, wherein the first set of token representations includes a single token representation according to the first resolution, and wherein the second set of token representations includes a plurality of token representations according to the second resolution.

18. The apparatus for processing image data according to claim 15, wherein the at least one processor is further configured to: concatenate the first set of token representations and the second set of token representations to generate a concatenated set of token representations; and Process the cascaded set of token representations using the neural network model to determine the first resolution or the second resolution as the scale of the first region of the input image.

19. The apparatus for processing image data according to claim 15, wherein the at least one processor is further configured to: Determine a corresponding scale for each respective region of the input image.

20. The apparatus for processing image data according to claim 19, wherein the at least one processor is further configured to: Determine a corresponding position encoding for each region of the input image.

21. The apparatus for processing image data according to claim 20, wherein the at least one processor is further configured to: For a region of the input image determined to have a scale corresponding to the second resolution, determine the corresponding position encoding by: determining the final position encoding of the region as a linear interpolation of a plurality of initial position encodings determined for the region.

22. The apparatus for processing image data according to claim 15, wherein the at least one processor is further configured to: Generate a mask for the input image, the mask indicating the corresponding scale determined as the first resolution or the second resolution for each respective region of the input image.

23. The apparatus for processing image data according to claim 22, wherein the transformer neural network model is configured to process adaptive mixed-resolution data based on the mask.

24. The apparatus for processing image data according to claim 15, wherein the neural network model is shared across regions of the input image.

25. The apparatus for processing image data according to claim 15, wherein the neural network model includes a Softmax layer, the Softmax layer being configured to determine a distribution over the first resolution and the second resolution.

26. The apparatus for processing image data according to claim 15, wherein the at least one processor is configured to: determine the first resolution or the second resolution as the scale of the first region of the input image based on one or more characteristics of the input image.

27. The apparatus for processing image data according to claim 26, wherein the one or more characteristics of the input image include a smoothness value associated with the first region of the input image, a complexity value associated with the first region of the input image, how many colors are associated with the input image, or a contrast value associated with the first region of the input image.

28. The apparatus for processing image data according to claim 15, wherein the input image includes image patches of an image.