Processing data using convolution as TRANSFORMER operation

CN119998816APending Publication Date: 2025-05-13QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069925.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-27
Filing Date
2023-09-28
Publication Date
2025-05-13

Smart Images

  • Figure CN119998816A_ABST
    Figure CN119998816A_ABST
Patent Text Reader

Abstract

Systems and techniques are described herein for processing data (e.g., image data) using convolution as a transformer (CAT) operation. The method includes: receiving, at a convolution engine of a machine learning system, a first set of features associated with an image and having a three-dimensional shape; applying, via the convolution engine, a depth-by-depth separable convolution filter to the first feature set to generate a first output; applying, via the convolution engine, a point-by-point convolution filter to the first output to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; modifying the second output to the three-dimensional shape to generate a second feature set; and combining the first set of features and the second set of features to generate an output set of features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure as a whole relates to processing data (e.g., image data) using convolution as a transformer (CAT) operation. In some aspects, the present disclosure relates to using a CAT engine to provide an approximation or estimation of one or more operations on data (e.g., processing image data for object detection, object classification, object recognition, map building operations such as simultaneous localization and mapping (SLAM), etc.). Additional or alternative aspects of the present disclosure relate to a self-attention as feature fusion (SAFF) engine for improving prediction results (e.g., object detection prediction). Background Art

[0002] Deep learning machine learning models (e.g., neural networks) can be used to perform a variety of tasks, such as detection and / or recognition (e.g., scene or object detection and / or recognition), map building (e.g., SLAM), depth estimation, pose estimation, image reconstruction, classification, three-dimensional (3D) modeling, dense regression tasks, data compression and / or decompression, image processing, and other tasks. Deep learning machine learning models can be general and can achieve high-quality results in a variety of tasks. However, although deep learning machine learning models can be general and accurate, the models are typically large and slow, and often have high memory requirements and computational costs. In many cases, the computational complexity of the model can be high, and the model can be difficult to train.

[0003] Some deep learning machine learning models (e.g., neural networks) may include lightweight architectures for performing image classification and / or object detection. Such lightweight models or systems include components for classifying or predicting objects in an image (e.g., a dog and its tail, head, legs, or paws). In some aspects, lightweight machine learning models can be useful when included as part of a system operating on a mobile device that does not have as much computing power or capabilities as other systems such as network servers or more powerful computing systems. In some cases, a machine learning model (e.g., a lightweight machine learning model) may utilize one or more transformers. However, due to the large amount of computation required for transformer operation, transformers can result in high latency. Summary of the invention

[0004] Systems and techniques are described for processing data using CAT operations, such as performing object detection, object classification, object recognition, map building operations (e.g., SLAM), and / or other operations on the data. In some cases, these CAT operations can be performed in lightweight scenarios (e.g., when performing object detection using a mobile device).

[0005] In some aspects, a method is provided, comprising: receiving a first feature set at a convolution engine of a machine learning system, the first feature set associated with an image, the first feature set associated with a three-dimensional shape; applying a depth-wise separable convolution filter to the first feature set via the convolution engine to generate a first output; applying a point-wise convolution filter to the first output via the convolution engine to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; modifying the second output to the three-dimensional shape to generate a second feature set; and combining the first feature set and the second feature set to generate an output feature set.

[0006] In some aspects, an apparatus is provided, comprising: at least one memory and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: receive a first feature set associated with an image at a convolution engine of a machine learning system, the first feature set associated with a three-dimensional shape; apply a depthwise separable convolution filter to the first feature set via the convolution engine to generate a first output; apply a pointwise convolution filter to the first output via the convolution engine to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; and modify the second output to the three-dimensional shape to generate a second feature set, and combine the first feature set and the second feature set to generate an output feature set.

[0007] In some aspects, a non-transitory computer-readable medium having instructions stored thereon is provided that, when executed by one or more processors, causes the one or more processors to: receive, at a convolution engine of a machine learning system, a first feature set associated with an image, the first feature set being associated with a three-dimensional shape; apply, via the convolution engine, a depthwise separable convolution filter to the first feature set to generate a first output; apply, via the convolution engine, a pointwise convolution filter to the first output to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; and modify the second output to the three-dimensional shape to generate a second feature set, and combine the first feature set and the second feature set to generate an output feature set.

[0008] In some aspects, an apparatus is provided. The apparatus includes: means for receiving a first feature set associated with an image at a convolution engine of a machine learning system, the first feature set associated with a three-dimensional shape; means for applying a depthwise separable convolution filter to the first feature set via the convolution engine to generate a first output; means for applying a pointwise convolution filter to the first output via the convolution engine to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; means for modifying the second output to the three-dimensional shape to generate a second feature set; and means for combining the first feature set and the second feature set to generate an output feature set.

[0009] In some aspects, the processes described herein (e.g., process 500 and / or other processes described herein) may be performed by a computing device or apparatus or a component or system of a computing device or apparatus (e.g., a chipset, one or more processors (e.g., CPU, GPU, NPU, DSP, etc.), an ML system such as a neural network model, etc.). In some aspects, one or more of the apparatuses described herein are and / or include and / or are part of an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device, a wireless communication device, a camera, a personal computer, a laptop, a vehicle, or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the apparatus includes a camera or multiple cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Illustrative examples of the present application are described in detail below with reference to the following drawings:

[0013] Figure 1 is a diagram illustrating an example of a convolutional neural network (CNN) according to aspects of the present disclosure;

[0014] Figure 2 is a diagram illustrating an example of a classification neural network system according to aspects of the present disclosure;

[0015] Figure 3 is a diagram illustrating an example of a convolution as a transformer (CAT) engine according to aspects of the present disclosure;

[0016] Figure 4 is a diagram illustrating an example of a Self-Attention as Feature Fusion (SAFF) engine according to aspects of the present disclosure;

[0017] Figure 5 is a diagram illustrating an example of a method for performing object detection according to aspects of the present disclosure;

[0018] Figure 6 is a diagram illustrating an example of a computing system according to aspects of the present disclosure. DETAILED DESCRIPTION

[0019] Certain aspects of the present disclosure are provided below. Some of these aspects may be applied independently, and some of them may be applied in combination, which is obvious to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it is apparent that various aspects may be implemented without these specific details. Each of the drawings and descriptions is not intended to be restrictive.

[0020] The following description provides only example aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the example aspects will provide a description that can be used to implement the example aspects to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the essence and scope of the present application as set forth in the appended claims.

[0021] As noted above, some deep learning machine learning models (e.g., neural networks), referred to as lightweight machine learning models (e.g., lightweight neural networks), may include fewer layers or components than machine learning models with complex architectures. For example, lightweight neural networks may be used to perform image classification and / or object detection. Such lightweight models or systems are intended for deployment on mobile devices or other devices that have limited computing power or capabilities like other systems.

[0022] In some cases, a system with a lightweight machine learning model may include the use of anchor boxes, backbone networks, and heads (e.g., prediction heads). For example, a lightweight system may divide an input image into multiple grids and may generate a set of anchor boxes for each grid. An anchor box may be a bounding rectangular box on one or more pixels in an image. The backbone of the system may include a convolutional neural network (CNN), one or more transformers, and / or a hybrid backbone structure for extracting features from an image. For example, a transformer or visual transformer may be used to divide an image into a sequence of non-overlapping patches, and then use multi-head self-attention in the transformer to learn inter-patch representations. The head of the lightweight system may be used to extract position-specific features from multiple output resolutions and predict offsets relative to a fixed anchor box. Using a head in a lightweight system may be problematic due to the lack of feature aggregation for prediction from multiple scales, in which case the amount of data available for prediction is limited.

[0023] A transformer is a specific type of neural network. An example of a system that uses a transformer is the MobileViT system described in MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer, Mehta, Rastegari, ICLR, 2022, which is incorporated herein by reference. Transformers perform well when used for image classification and object detection, but require a lot of computation to perform image classification and object detection tasks.

[0024] For example, suppose n vectors are input to a transformer, which may compute the dot product of each vector with every other vector and then apply a Softmax layer (or in some cases, a multilayer perceptron (MLP)). After the transformer applies the Softmax layer, the transformer computes a weighted combination of the outputs. The transformer performs a large amount of computation when performing such operations. In some cases, the large amount of computation may be attributed to the use of pairwise self-attention in addition to the Softmax function. In addition, the Softmax operation performed by the Softmax layer is known to be computationally expensive and slow. For example, the Softmax function converts a vector of K real numbers into a probability distribution of K possible results. In neural network applications, the number of possible results K may be large. For example, in the case of a neural language model that predicts the most likely result from a vocabulary, the possible results may contain millions of possible words. For image-based predictions, the results may be even higher. This large number of possible results may make the calculations for the Softmax layer computationally expensive. In addition, the gradient descent backpropagation method used to train such a neural network involves calculating the Softmax for each training example. The number of training examples may also become large.

[0025] Using transformers as part of the backbone structure can therefore result in high latency due to the large amount of computation required for transformer operations (e.g., as compared to systems using CNNs as the backbone). The latency is particularly noticeable on mobile devices with limited computing resources. Such computational effort is also a major limiting factor in developing more powerful object prediction models, especially on mobile devices with limited computing power.

[0026] Systems and techniques for improving machine learning systems (e.g., neural networks) are described herein. According to some aspects, systems and techniques provide a convolution as a transformer (CAT) engine that is used to perform one or more operations on data (e.g., sensor data, such as image data from one or more image sensors or cameras, radar data from one or more radar sensors, light detection and ranging (LIDAR) data from one or more LIDAR sensors, any combination thereof, and / or other data), such as object detection, object classification, object recognition, map building operations (e.g., SLAM), and / or other operations on data.

[0027] In some aspects, the CAT engine can be used in a backbone neural network architecture, such as for object classification or object detection in an image. The CAT engine can utilize convolution operations (e.g., to perform image classification, object detection, and / or operations), such as by maintaining a large global receptive field to approximate or perform operations equivalent to those of a transformer. Thus, the CAT engine can improve (e.g., reduce) latency compared to systems using transformers. For example, convolution operations can be used by the CAT engine instead of complex transformer encoder operations, such as (Q, K, V) encoding, Softmax, multi-layer perceptron (MLP), etc. In some aspects, the use of the CAT engine can reduce the theoretical complexity of an operation from O(n) to O(n). 2 *d) or O(n*log n *d) is reduced to O(n*d), where “O()” corresponds to the complexity of the function, n corresponds to the input sequence length (e.g., the number of pixels) and d is the per-pixel feature vector.

[0028] Additionally or alternatively, in some aspects, these systems and techniques provide a self-attention as feature fusion (SAFF) engine that can be used in a backbone neural network architecture, such as for object classification or object detection in an image. The SAFF engine can use a transformer (e.g., using self-attention) to fuse features extracted from the backbone at multiple scales or abstraction levels. For example, the SAFF engine can utilize a transformer (Q, K, V) encoding to project features from different scales or abstraction levels onto a common space for self-attention.

[0029] In some cases, a SAFF engine may be used in a neural network with or without a CAT engine. Similarly, in some cases, a CAT engine may be used in a neural network with or without a SAFF engine.

[0030] Details related to these systems and techniques are described below with respect to the various figures.

[0031] Figure 1is a diagram illustrating an example of a convolutional neural network (CNN) 100. An input layer 102 of the CNN 100 includes data representing an image. In some aspects, the data may include an array of numbers representing pixels of the image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at that location in the array. Using the previous example from above, the array may include a 28×28×3 array of numbers having 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, etc.). The image may be passed through a convolutional hidden layer 104, an optional non-linear activation layer, a pooling hidden layer 106, and a fully connected hidden layer 108 to obtain an output at an output layer 110. Although Figure 1 Only one hidden layer in each hidden layer is shown in , but one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in CNN 100. As previously described, the output may indicate a single category of an object, or may include a probability of a category that best describes an object in the image.

[0032] The first layer of CNN 100 is a convolutional hidden layer 104. Convolutional hidden layer 104 analyzes the image data of input layer 102. Each node of convolutional hidden layer 104 is connected to an area of ​​nodes (pixels) of the input image called a receptive field. Convolutional hidden layer 104 can be considered as one or more filters (each filter corresponds to a different activation or feature map), where each convolution iteration of the filter is a node or neuron of convolutional hidden layer 104. In some aspects, the area of ​​the input image covered by the filter at each convolution iteration will be the receptive field of the filter. In some cases, if the input image includes a 28×28 array and each filter (and corresponding receptive field) is a 5×5 array, there will be 24×24 nodes in convolutional hidden layer 104. Each connection between a node and the receptive field of that node learns a weight, and in some cases learns an overall bias, so that each node learns to analyze its specific local receptive field in the input image. Each node of hidden layer 104 will have the same weights and biases (referred to as shared weights and shared biases). In some cases, the filter has an array of weights (numbers) and the same depth as the input. For the video frame example, the filter will have a depth of 3 (according to the three color components of the input image). In some aspects, the size of the filter array can be 5×5×3, corresponding to the size of the receptive field of the node.

[0033] The convolutional nature of the convolutional hidden layer 104 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. In some aspects, the filter of the convolutional hidden layer 104 may start at the upper left corner of the input image array and may be convolved around the input image. As noted above, each convolution iteration of the filter may be considered as a node or neuron of the convolutional hidden layer 104. In each convolution iteration, the value of the filter is multiplied by the original pixel values ​​of the corresponding number of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the upper left corner of the input image array). The multiplications from each convolution iteration may be added together to obtain the sum of that iteration or node. Next, the process is continued at the next position in the input image according to the receptive field of the next node in the convolutional hidden layer 104.

[0034] In some aspects, the filter may be moved by a step amount (also referred to as a stride) to the next receptive field. The step amount may be set to 1 or another suitable amount. In some cases, if the step amount is set to 1, the filter will move 1 pixel to the right at each convolution iteration. Processing the filter at each unique position of the input volume produces a number representing the filter result at that position, thereby determining a sum value for each node of the convolution hidden layer 104.

[0035] The construction of the map from the input layer to the convolutional hidden layer 104 is called an activation map (or feature map). The activation map includes a value for each node that represents the filter result at each location of the input volume. The activation map may include an array that includes various summed values ​​produced by each iteration of the filter over the input volume. In some aspects, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map may include a 24×24 array. The convolutional hidden layer 104 may include several activation maps in order to identify multiple features in the image. Figure 1 The example shown in includes three activation maps. Using these three activation maps, the convolutional hidden layer 104 can detect three different kinds of features, each of which is detectable over the entire image.

[0036] In some exemplary aspects, a nonlinear hidden layer can be applied after the convolutional hidden layer 104. The nonlinear layer can be used to introduce nonlinearity into a system that already computes linear operations. In some aspects, the nonlinear layer can be a rectified linear unit (ReLU) layer. The ReLU layer can apply the function f(x)=max(0,x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, the ReLU can increase the nonlinear properties of the CNN 100 without affecting the receptive field of the convolutional hidden layer 104.

[0037] A pooling hidden layer 106 may be applied after the convolutional hidden layer 104 (and after the nonlinear hidden layer when used). The pooling hidden layer 106 may be used to simplify the information in the output from the convolutional hidden layer 104. In some aspects, the pooling hidden layer 106 may take each activation map output from the convolutional hidden layer 104 and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by a pooling hidden layer. The pooling hidden layer 104 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. A pooling function (e.g., a max pooling filter, an L2 norm filter, or other suitable pooling filter) may be applied to each activation map included in the convolutional hidden layer 104. In Figure 1 In the method shown in , three pooling filters are used for the three activation maps in the convolution hidden layer 104.

[0038] In some aspects, max pooling may be used by applying a max pooling filter (e.g., having a size of 2×2) with a stride (e.g., equal to the dimension of the filter, such as a stride of 2) to the activation map output from the convolutional hidden layer 104. The output from the max pooling filter may include the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer may summarize an area of ​​2×2 nodes in the previous layer (each node is a value in the activation map). In some aspects, the four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter having a dimension of 24×24 nodes from the convolutional hidden layer 104, the output from the pooling hidden layer 106 will be an array of 12×12 nodes.

[0039] In some aspects, an L2 norm pooling filter may also be used. The L2 norm pooling filter includes calculating the square root of the sum of the squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (rather than calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0040] Intuitively, a pooling function (e.g., max pooling, L2 norm pooling, or other pooling functions) determines whether a given feature is found anywhere in a region of an image, and discards the exact location information. The pooling function can be performed without affecting the results of feature detection, because once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max pooling (and other pooling methods) provides the following benefits: far fewer features are pooled, thereby reducing the number of parameters required in subsequent layers of CNN 100.

[0041] The last layer of connections in the network can be a fully connected layer that connects each node from the pooling hidden layer 106 to each of the output nodes in the output layer 110. Using the above example, the input layer includes 28×28 nodes that encode the pixel intensities of the input image, the convolutional hidden layer 104 includes 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for the filter) to three activation maps, and the pooling layer 106 includes a layer of 3×12×12 hidden feature nodes based on applying a maximum pooling filter to a 2×2 region on each of the three feature maps. As an extension of the above concept, the output layer 110 may include ten output nodes. In some aspects, each node of the 3×12×12 pooling hidden layer 106 may be connected to each node of the output layer 110.

[0042] The fully connected layer 108 can obtain the output of the previous pooling layer 106 (which should represent the activation map of the high-level features) and determine the features most relevant to a particular category. In some aspects, the fully connected layer 108 layer can determine the high-level features that are most strongly associated with a particular category, and can include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 108 and the pooling hidden layer 106 can be calculated to obtain probabilities for different categories. In some aspects, if the CNN 100 is used to predict that the object in the video frame is a person, there will be high values ​​in the activation map representing the high-level features of the person (e.g., there are two legs, there is a face at the top of the object, there are two eyes at the upper left and upper right of the face, there is a nose in the middle of the face, there is a mouth at the bottom of the face, and / or other features common to people).

[0043] In some aspects, the output from the output layer 110 may include an M-dimensional vector (e.g., the value of the vector may be M=10), where M may include the number of categories that the program must choose from when classifying an object in an image. Other example outputs may also be provided. Each number in the M-dimensional vector may represent the probability that an object belongs to a certain category. In some aspects, if the 10-dimensional output vector represents that objects of ten different categories are [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that the image is a 5% probability of an object of the third category (e.g., a dog), the image is an 80% probability of an object of the fourth category (e.g., a person), and the image is a 15% probability of an object of the sixth category (e.g., a kangaroo). The probability for a category may be considered as the confidence level that the object is part of that category.

[0044] Backbones such as using convolutional layers to extract high-level features Figure 1One problem with the CNN 100 in is that the receptive field is limited by the size of the convolution kernel. For example, convolution cannot extract global information. As noted above, transformers can be used to extract global information, but are computationally expensive and therefore add significant latency when performing object classification and / or detection tasks.

[0045] Figure 2 A classification neural network system 200 including a mobile visual transformer (MobileViT) block 226 is illustrated. The MobileViT block 226 relies on the initial application of transformers in language processing and applies the technique to image processing. Transformers for image processing measure the relationship between pairs of input labels or pixels as the basic unit of analysis. However, computing relationships for each pair of pixels in a typical image is prohibitive in terms of memory and computation. The MobileViT block 262 calculates the relationship between pixels in various small parts of the image (e.g., 16×16 pixels) at a significantly reduced cost. These parts (with positional embeddings) are placed in sequence. The embedding is a learnable vector. Each part is arranged into a linear sequence and multiplied by an embedding matrix. The result with the positional embedding is fed to the transformer 208.

[0046] like Figure 2 As shown in and according to some aspects, an input image 214 having a height H, a width W, and a number of channels (e.g., H*W pixels with 3 channels corresponding to red, green, and blue color components) may be provided to a convolution block 216. The convolution block 216 applies a 3×3 convolution kernel (with a stride or stride of 2) to the input image 214. The output of the convolution block 216 passes through a plurality of MobileNetv2 (MV2) blocks 218, 220, 222, 224 to generate a downsampled output feature set 204 having a height H, a width W, and a dimension C. Each MobileNetv2 218, 220, 222, 224 is a feature extractor for extracting features from the output of the previous layer. Other feature extractors besides the MobileNetv2 block may be used in some cases. The blocks that perform downsampling are marked with the symbol “↓2”, which corresponds to a stride or stride of 2.

[0047] like Figure 2As shown in , the MobileViT block 226 is illustrated in more detail than the other MobileViT blocks 230 and 234 of the classification neural network system 200. As shown, the MobileViT block 226 can use a convolution layer to process the feature set 204 to generate a local representation 206. The transformer as a convolution can produce a global representation by unfolding the data or feature set (having dimensions of H, W, d), performing the transformation using a transformer, and folding the data again to obtain another feature set (having dimensions of H, W, d) that is then output to the fusion layer 210. The MobileViT block 226 uses the transformer 208 to replace the local processing of the convolution operation with a global processing (e.g., using a global representation).

[0048] The fusion layer 210 may fuse or compare the data with the original feature set 204 to generate output features Y 212. The output features Y 212 output from the MobileViT block 226 may be processed by a MobileNetv2 block 228, followed by another application of a MobileViT block 230, followed by another application of a MobileNetv2 block 232, and yet another application of a MobileViT block 234. A 1×1 convolutional layer 236 (e.g., a fully connected layer) may then be applied to generate a global pooling 238 of the output. The downsampling of the data on the various block operations may result in taking an image size of 128×128 at the block 216 and generating a 1×1 output (shown as a global pooling 238) that includes a global pooling of linear data.

[0049] Figure 3 is a diagram illustrating an example of a convolution as a transformer (CAT) engine 300 in accordance with aspects of the systems and techniques described herein. The CAT engine 300 may be implemented as a machine learning system such as Figure 2 In some aspects, the CAT engine 300 may represent or approximate the operations or results of the transformer 208 of the classification neural network system 200 (e.g., approximate the self-attention and global feature extraction performed by the transformer 208) with much lower complexity and latency.

[0050] As shown, a feature set 302 is received at a CAT engine 300. In some aspects, the feature set 302 (also referred to as a feature representation) may include Figure 2The H×W×d local representation 206 of the original image 214. In some aspects, the feature set 302 may represent any extracted feature data associated with the original image 214. The feature set 302 may represent an intermediate feature map extracted from one or more convolutions of the original image 214. The feature set 302 may include one or more of a tensor, a vector, an array, a matrix, or other data structures including values ​​associated with features. The feature set 302 may have a shape that may be three-dimensional, including a height H, a width W, and a dimension D. In some aspects, the feature set may be associated with an image having H*W pixels, where each pixel has a color value (e.g., a red (R) value, a green (G) value, and a blue (B) value) as dimension D.

[0051] The CAT engine 300 performs a depthwise convolution on the feature set 302 via a depthwise convolution engine 303 to generate a first output feature set 304 that can represent global information. In some aspects, the CAT engine 300 can apply a depthwise separable convolution filter to the feature set 302 to generate the first output feature set 304. The feature set 302 can be referred to as an intermediate feature set (e.g., an intermediate tensor, vector, etc.). In an illustrative example, the depthwise convolution engine 303 can be a multi-layer perceptron (MLP) operating on the spatial domain of the feature set 302.

[0052] In some aspects, the operation of the depth-wise convolution engine 303 may include where an HxW kernel may be applied to each channel in the depth dimension D of the feature set 302, where H is the height dimension and W is the width dimension. In the depth-wise convolution engine 303, the kernel size may be the same as the feature set 302. In particular, the height H and width W of the kernel may be equal to the height H and width W of the feature set 302, respectively, resulting in the first output feature set 304 having a dimension of 1×1×D (one value (1,1) for each channel in the depth dimension D of the feature set 302). The values ​​in the H×W kernel may be multiplied by the feature values ​​at each corresponding position in the H×W feature set 302 in each channel, and the depth-wise convolution engine 303 may perform an operation to determine the value of each channel. In some aspects, the operation may include a multiply-accumulate (MAC) operation. The MAC operation may include applying the kernel to the first channel and performing an element-wise multiplication of the channel's features multiplied by the weights of the kernel and the sum. The depth-wise convolution engine 303 may apply the H×W kernel to each channel independently or to two or more of the channels (eg, all channels in some cases) in parallel.

[0053] Typically, the image size is larger than the kernel size. Keeping the kernel size the same as the feature set 302 enables the output to be an approximation, proxy, or estimate of the global information from the spatial data (channel locations) in the feature set 302. In some aspects, the first output feature set 304 may represent a single global vector that represents the global information extracted from the spatial domain of the feature set 302.

[0054] The first output feature set 304 may be provided to a point-by-point convolution engine 305. The point-by-point convolution engine 305 may apply a point-by-point convolution filter to the first output feature set 304 and extract global information from the channel dimension D. In some aspects, the combination of the depth-by-depth convolution engine 303 and the point-by-point convolution engine 305 may be referred to as a convolution engine. The point-by-point convolution engine 305 may apply a kernel with a kernel size of 1×1×D (e.g., the kernel size may have a depth equal to the depth D of the feature set 302, referred to as the number of channels), which may be a fully connected layer. The convolution engine 305 may iterate the 1×1 kernel through each point or value of the first output feature set 304, thereby processing the values ​​in all channels in the dimension D of the first output feature set 304. In some aspects, the first output feature set 304 may be considered to be a 1×D input vector, and the matrix associated with the point-by-point convolution may be a D×D matrix. The 1×D input vector or matrix multiplied by the D×D matrix produces a second output of a 1×D vector or matrix. The point-by-point convolution engine 305 can extract information into a single value by processing all channels of the first output feature set 304 and returning a single value. The point-by-point convolution engine 305 can therefore extract information from the channel dimension D to generate a feature set 306, which can be considered as global information such as a global vector. The feature set 306 can then be modified (e.g., by replicating the values ​​in each dimension D of the feature set 306 in the H and W dimensions) to obtain another feature set 308 with H×W×D dimensions. The feature set 308 can be considered as global information that has the same three-dimensional shape as the feature set 302 and approximates or estimates global information in both the spatial domain and the channel dimension. One method of implementing the modification by replicating the corresponding values ​​in each dimension D of the feature set 306 is as follows:

[0055] y=reshape(global_information,(H,W,D)).

[0056] Using this formula, the CAT engine 300 can distribute the values ​​of the vector in the H and W dimensions. The feature set 308 can be characterized as an output vector that represents an approximation or estimate of the global information from both the spatial dimension and the channel dimension. The CAT engine 300 can then perform an element-by-element product 310 between the feature set 308 (global information) and the feature set 302 to generate an output feature set 312 (e.g., as an output feature map). In some cases, the element-by-element product 310 can be performed as follows:

[0057] x=elementwise_product(x,y).

[0058] Element-wise product 310 determines how the local vectors in feature set 302 relate to the global information in feature set 306 or feature set 308, and thus extracts how each local vector in feature set 302 relates to feature set 306 or 308 (e.g., the feature set may be global information, such as a global vector, as noted above). If two corresponding comparison values ​​are similar, then the values ​​will be relatively larger. If two corresponding comparison values ​​are not similar, then the values ​​will be relatively larger, but in a negative direction. This feature of CAT engine 300 allows CAT engine 300 to simulate Figure 2 302. In transformer 208, the approach is to take the dot product between each corresponding pixel and every other pixel, which is computationally expensive. Instead of this process, CAT engine 300 uses the global information in feature set 308 as an approximation, and CAT engine 300 uses an element-wise product 310 between feature set 308 and the global information in feature set 302, rather than implementing a dot product. Output feature set 312 represents a feature set that identifies the correlation between feature set 302 and the global information or feature set 308.

[0059] In such cases, the CAT engine 300 may be applied iteratively (twice or more times). In such cases, a first convolution engine (which may include a depth-wise convolution engine 303 and a point-wise convolution engine 305) may process the input feature map to generate an intermediate feature map. A second convolution engine (e.g., a depth-wise convolution engine 303, a point-wise convolution engine 305) may process the intermediate feature map to generate an output feature map. Additional convolution engines (e.g., a depth-wise convolution engine 303, a point-wise convolution engine 305) may also be applied.

[0060] The benefit of using the CAT engine 300 is to address the high latency caused by the self-attention layer and the use of the Softmax function by the transformer 208. By using the CAT engine 300, the latency can be reduced by more than 50% relative to using the transformer 208. Using simpler functions than the self-attention layer in the transformer 208 provides lower latency and lower number of calculations, and thus the tradeoff can be useful for mobile device object prediction. Therefore, the method disclosed above provides an improvement in the performance of a device performing object classification or detection. In some aspects, the complexity is reduced from O(n 2 *d) is reduced to O(n*d). In one aspect, the method uses convolution operations instead of complex transformer encoder operations such as (Q,K,V) encoding, Softmax function and multi-layer perceptron (MLP).

[0061] Figure 4The self-attention as feature fusion (SAFF) engine 400 is illustrated as an example. The SAFF engine 400 can be implemented as a machine learning system such as Figure 2 A SAFF engine 400 is a neural network system 200 for classifying objects, an object detection system, or other system. The SAFF engine 400 can extract features from multiple resolution intermediate feature maps 402, 404, 406, and can use a self-attention layer 408 (also referred to as a self-attention engine) to fuse feature maps that can be used for prediction. Visual self-attention is a mechanism for correlating different positions of a single sequence to compute a representation of the same sequence. The SAFF engine 400 allows a neural network to focus on specific parts of a complex input (in the disclosed example, intermediate feature maps 402, 404, etc.) one by one until the entire data set can be classified.

[0062] The SAFF engine 400 or self-attention layer 408 aggregates features from multiple different resolutions or scales from different intermediate feature maps 402, 404, 406 and increases the accuracy of the final prediction. Each feature map from the intermediate feature maps 402, 404, 406 includes a set of features output by a convolutional layer of an underlying neural network (e.g., backbone), where later feature maps have lower dimensional data from the original image 214. The different sizes of the feature maps are derived from the input image 214 and the Figure 2 The output features of the image 214 are downsampled to a size of 128×128 at block 216, the output features of the image are downsampled to a size of 64×64 at block 218, the output features of the image are downsampled to a size of 32×32 at block 224, and so on. Each of the feature maps of different sizes may correspond to Figure 4 The intermediate feature maps 402, 404, 406 in , these intermediate feature maps can be referred to as intermediate feature maps 402, 404, 406.

[0063] Each intermediate feature map 402, 404, 406 is detected at a specific resolution associated with the anchor box or box rectangle discussed above. The use of three intermediate feature maps 402, 404, 406 is only an example. Any number of two or more feature maps can be used in conjunction with the concept of the SAFF engine 400.

[0064] The input image in a typical system is received at the backbone of the neural network system. The backbone can be something like Figure 2 The classification neural network system 200 in the embodiment of the present invention may include Figure 3 CAT 300 engine to replace Figure 2208 shown in FIG. Features are extracted from multiple layers. In some aspects, some of the layers of the backbone or new layers added are used to extract features from different resolutions.

[0065] The SAFF engine 400 may make predictions for each resolution or for each intermediate feature map. In some aspects, in a conventional approach, an intermediate feature map 402 may have a resolution of 19×19 values ​​or pixels, so the system will perform 19×19 predictions from that layer. The 19×19 values ​​may involve using 19×19 anchor boxes as discussed above. Another intermediate feature map 404 may have a resolution of 10×10 values, so the system will perform 10×10 predictions. The intermediate feature map 406 may have a resolution of 3×3 values, so the system will perform 3×3 predictions. The SAFF engine 400 looks at different feature maps and predicts objects in the image. The conventional approach described above is a single shot multi-box detector (SSD) architecture.

[0066] Figure 4 The improvements disclosed in involve extracting features from multiple resolution intermediate feature maps 402, 404, 406 and providing this data to a self-attention layer 408 which fuses the feature maps that can be used for prediction by a prediction layer 410.

[0067] Each intermediate feature map 402, 404, 406 may represent a corresponding continuous output of a convolution filter (or other method). Each corresponding intermediate feature map 402, 404, 406 may have a reduced size from the original image 214 or from a previous input to the corresponding convolution filter. Each intermediate feature map 402, 404, 406 may represent a different output and / or may be associated with performing detection at a specific resolution related to the applied anchor box. Therefore, if the intermediate feature map 402 has a 19×19 dimension, the intermediate feature map 402 involves a 19×19 anchor box. If the system predicts from an intermediate layer or map 402, 404, 406, a specific lower-level layer or intermediate feature map 402 can be used to detect, for example, the tail of a dog, and another intermediate feature map 404 can be used to detect the head of a dog. If the intermediate feature map 406 has a 1×1 dimension, the intermediate feature map 406 can be used to predict objects in the entire image with only one anchor box. In other words, if a picture of a dog covers the entire image 214, then it may be possible to predict the dog from the 1×1 dimensional intermediate map 406. The feature map 406 may be a single vector with all the information associated with the original image, which in some aspects may be 300×300 pixels. The 3×3 intermediate feature map (in one aspect, feature map 404) may attempt to predict nine objects within the image 214. Each intermediate feature map 402, 404, 406 may be used to predict an object at a specific scale.

[0068] A method or process of providing self-attention as feature fusion via the SAFF engine 400 may include obtaining a first feature from an intermediate feature map having a first resolution (e.g., the intermediate feature map 402) and extracting a second feature from an intermediate feature map having a second resolution (e.g., the intermediate feature map 404). Based on the first feature and the second feature (e.g., the features in the corresponding intermediate feature maps 402 and 404), the SAFF engine 400 may fuse the first feature and the second feature via the self-attention layer 408 to obtain a fused feature. The method or process may include predicting (using the prediction layer 410) an element in the image based on the fused feature.

[0069] In some cases, the self-attention layer 408 may perform a dot product of each of the first features of the first resolution intermediate feature map 402 with each other feature in the second resolution intermediate feature map 404 to generate the first feature, and perform a dot product of each of the second features of the second resolution intermediate feature map with each other feature in the first resolution intermediate feature map 402 to generate the second feature. The SAFF engine 400 (e.g., the self-attention layer 408) may apply a Softmax function to the first feature and the second feature, and may perform a weighted sum on the first feature and the second feature to fuse the first feature and the second feature from multiple resolutions for prediction by the prediction layer 410.

[0070] In some aspects, if there are more than two feature maps ( Figure 4 406), then for each pixel in the intermediate feature map 402, the self-attention layer 408 may apply self-attention for the corresponding pixel in the feature map 402 to every other pixel in the intermediate feature map 402 plus every other pixel of one or more other feature maps 404, 406. The result will be to obtain useful information for that pixel in one intermediate feature map 402 from those other feature maps 404, 406, and calculate a new feature vector based on the combined data. Thus, while the feature set previously used for each feature map would be feature 1, feature 2, and feature 3 for each of the intermediate feature maps 402, 404, 406, the SAFF engine 400 may generate feature vectors. 1' ,feature 2' and Features 3' The new value of . 1' ,feature 2' and Features 3' Each of the layers can be used to make separate predictions, and the SAFF engine 400 can fuse information from other layers to predict the 1' ,feature 2' and Features 3' Returns better predictions. These features, i.e., features 1' ,feature 2'and Features 3' More robust than feature 1, feature 2, and feature 3. For example, a smaller feature map such as feature map 406 may have views or information associated with other intermediate feature maps 402, 404, thereby making the prediction ability of the SAFF engine 400 more robust. The SAFF engine 400 introduces a more global view of the data by applying information from other feature maps to make predictions for any individual pixel in any of the intermediate feature maps 402, 404, 406.

[0071] The SoftMax function converts a vector of K real numbers into a probability distribution of K possible results. The SoftMax function can be used as the last activation function of a neural network to normalize the output of the network to a probability distribution on the predicted output category. In some aspects, before applying the Softmax function, some vector components may be negative, or greater than one; and may not sum to 1. After applying the Softmax function, each component will be in an interval and the components will add up to 1, so that they can be interpreted as probabilities. Functions other than the formal Softmax can also be applied, which perform conversions of real numbers to probability distributions.

[0072] The method may also include performing a first prediction based on the first feature, performing a second prediction based on the second feature, and performing a third prediction based on the fused feature. Although the above method only includes two intermediate feature maps, such as map construction 402, 404, the method may also cover three intermediate feature maps including map construction 406 or more intermediate feature maps.

[0073] In general, the use of the SAFF engine 400 and the self-attention layer 408 fuses predictions from two or more of the multiple scales, resolutions, or intermediate feature maps 402, 404, 406, rather than making separate predictions. The SAFF engine 400 (and / or prediction layer 410) still makes three predictions, but the SAFF engine 400 fuses information from adjacent intermediate feature maps 402, 404, 406 via the self-attention layer 408 to make the final prediction smarter by using the fused information.

[0074] Figure 5 Examples are provided for use, for example, with a CAT engine (e.g., as described with respect to Figure 3 ) and / or SAFF engine 400 (e.g., as described with respect to Figure 45. The process 500 of processing image data using a convolution engine of a machine learning system (described in detail in the foregoing) may include, at block 502, receiving a first feature set at a convolution engine of a machine learning system. In some aspects, the first feature set may be associated with an image (or other data). In some examples, the first feature set may be associated with a three-dimensional shape. In some cases, the first feature set may include a first tensor, a first vector, a first matrix, a first array, and / or other representations of values ​​that include the first feature set.

[0075] At block 504, the method may include applying, via a convolution engine of the machine learning system, a depthwise separable convolution filter to the first feature set to generate a first output. In some aspects, the convolution engine may include Figure 3 CAT engine 300. For example, the depth-wise convolution engine 303 of the CAT engine 300 may apply depth-wise separable convolution filters to the first feature set, as described herein. The convolution engine may be configured to perform transformer operations (e.g., the convolution engine approximates the operations of transformer 208). Additionally or alternatively, in some aspects, the convolution engine may be configured to perform pair-wise self-attention and global feature extraction operations (e.g., the convolution engine approximates pair-wise self-attention models and global feature extraction for image classification). In some aspects, the depth-wise separable convolution filters applied by the convolution engine (e.g., the depth-wise convolution engine 303) may include a spatial multilayer perceptron (MLP), a fully connected layer, or other layers that extract information from the spatial domain of the first feature set to generate a first output.

[0076] At block 506, process 500 may include applying a point-wise convolution filter to the first output via a convolution engine (e.g., CAT engine 300) to generate a second output based on (or approximating) global information from a spatial dimension and a channel dimension associated with the image. For example, point-wise convolution engine 305 of CAT engine 300 may apply a point-wise convolution filter to the first output to generate the second output, as described herein. In some aspects, the point-wise convolution filter applied by the convolution engine (e.g., point-wise convolution engine 305) may include a channel multilayer perceptron (MLP) that extracts information from the channel dimension of the first feature set to generate the second output.

[0077] At block 508, process 500 may include modifying the second output into a three-dimensional shape to generate a second feature set. In some cases, the first feature set may be associated with a local representation of the image, and the second feature set may be associated with a global representation of the image. In some cases, the second feature set may include a second tensor, a second vector, a second matrix, a second array, and / or other representations that include values ​​of the second feature set.

[0078] At box 510, process 500 may include combining the first feature set and the second feature set to generate an output feature set. In some aspects, the output feature set may include an output tensor, an output vector, an output matrix, an output array, and / or other representation of values ​​that include the output feature set. In some cases, process 500 may include performing image classification associated with the image based on the output feature set via a machine learning system (e.g., to classify one or more objects in the image). In some cases, process 500 may include performing object detection associated with the image based on the output feature set via a machine learning system (e.g., to determine the location and / or pose of one or more objects in the image, and in some cases generate a corresponding bounding box for each of the one or more objects). In some cases, combining the first feature set and the second feature set to generate the output feature set may include performing an element-wise product 310 between the first feature set 302 and the feature set 308, as described with respect to Figure 3 As described.

[0079] In some aspects, process 500 may also include receiving a third feature set at a second convolution engine of the machine learning system (e.g., an additional CAT engine of the machine learning system). The third feature set may be generated based on the output feature set. Process 500 may also include applying an additional depth-wise separable convolution filter to the third feature set via the second convolution engine to generate a third output. Process 500 may include applying an additional point-wise convolution filter to the third output via the second convolution engine to generate a fourth output based on global information from the spatial dimensions and channel dimensions associated with the image. Process 500 may also include modifying the fourth output into a three-dimensional shape to generate a fourth feature set. Process 500 may include combining the third feature set and the fourth feature set to generate an additional output feature set.

[0080] In some cases, process 500 may apply principles from SAFF engine 400 independently or in combination with the principles of the CAT engine described above. For example, process 500 may include obtaining a first feature from a first intermediate feature map based on the output feature set. The first intermediate feature map may have a first resolution. Process 500 may also include obtaining a second feature from a second intermediate feature map based on the output feature set. The second intermediate feature map may have a second resolution different from the first resolution. Process 500 may include combining the first feature and the second feature via a self-attention engine (e.g., SAFF engine 400) to generate a fused feature, and predicting the location of an object in the image via a prediction layer (e.g., prediction layer 410) based on the fused feature. In some aspects, the first intermediate feature map may include Figure 4 The intermediate feature map 402 of , and the second intermediate feature map may include the intermediate feature map 404.

[0081] In some cases, to combine the first feature and the second feature, the process 500 may include performing, via a self-attention engine (e.g., SAFF engine 400), a first dot product of each of the first features in the first intermediate feature map with each of the second features in the second intermediate feature map 404. The process 500 may include performing, via a self-attention layer 408, a second dot product of each of the second features in the second intermediate feature map with each of the first features in the first intermediate feature map. The process 500 may also include applying, via the self-attention layer 408, a Softmax function to an output of the first dot product and an output of the second dot product. The process 500 may include performing, via the self-attention layer 408, a weighted sum of the outputs of the Softmax function to combine the first feature and the second feature.

[0082] The process 500 may also include performing a first prediction based on the first feature via a prediction layer (e.g., prediction layer 410), and performing a second prediction based on the second feature via the prediction layer. The process 500 may include performing a third prediction based on the fused feature via the prediction layer 410. Although the above method includes references to two intermediate feature maps, more than two intermediate feature maps may be processed in a similar manner using the SAFF engine 400.

[0083] In some aspects, the processes described herein (e.g., process 500 and / or other processes described herein) may be performed by a computing device or apparatus or a component or system of a computing device or apparatus (e.g., a chipset, one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), etc.), an ML system such as a neural network model, etc.). The computing device or apparatus may be a vehicle or a component or system of a vehicle, a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, and / or a mixed reality (MR) device), or other type of computing device. In some cases, the computing device or apparatus may be Figure 6 The computing system 600, vehicle and / or other computing device or apparatus may be configured to:

[0084] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, an AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device having resource capabilities to perform the processes described herein, including process 500 and / or other processes described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some aspects, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.

[0085] The components of the computing device may be implemented in circuits. In some aspects, the components may include and / or may be implemented using electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits)), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0086] Process 500 is illustrated as a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0087] Additionally, process 500, methods, and / or other processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, through hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program that includes a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0088] Figure 6 is a diagram illustrating a system for implementing certain aspects of the present technology. In particular, Figure 6 Illustrated is a computing system 600, which may be any computing device constituting an internal computing system, a remote computing system, a camera, or any components thereof, wherein the components of the system communicate with each other using connection 605. Connection 605 may be a physical connection using a bus, or a direct connection into processor 610, such as in a chipset architecture. Connection 605 may also be a virtual connection, a networked connection, or a logical connection.

[0089] In some aspects, computing system 600 is a distributed system in which the functionality described in the present disclosure may be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a number of such components that each perform some or all of the functionality for which the component is described. In some aspects, each component may be a physical or virtual device.

[0090] System 600 includes at least one processing unit (CPU or processor) 610 and connections 605 that couple various system components including system memory 615 such as read only memory (ROM) 620 and random access memory (RAM) 625 to processor 610. Computing system 600 may include cache 611 of high-speed memory directly connected to, in close proximity to, or integrated as part of processor 610.

[0091] Processor 610 may include any general purpose processor and hardware or software services such as services 632, 634, and 636 stored in storage device 630 configured to control processor 610, as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 610 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0092] To enable user interaction, the computing system 600 includes an input device 645 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 600 may also include an output device 635, which may be one or more of a plurality of output mechanisms. In some examples, a multimodal system may enable a user to provide multiple types of input / output to communicate with the computing system 600. The computing system 600 may include a communication interface 640, which may generally govern and manage user input and system output.

[0093] The communication interface can perform or facilitate the use of wired and / or wireless transceivers to receive and / or send wired or wireless communications, including using audio jacks / plugs, microphone jacks / plugs, universal serial bus (USB) ports / plugs, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Dedicated wired ports / plugs, Wireless signal transmission, Low power (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, WLAN signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / Long Term Evolution (LTE) cellular data network wireless signal transmission, self-organizing network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.

[0094] The communication interface 640 may also include one or more GNSS receivers or transceivers for determining the location of the computing system 600 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the Global Positioning System (GPS) of the United States, the Global Navigation Satellite System (GLONASS) of Russia, the BeiDou Navigation Satellite System (BDS) of China, and the Galileo GNSS of Europe. There is no limitation to operating on any particular hardware arrangement, and thus the basic features herein may be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0095] The storage device 630 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as a cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / magnetic stripe, any other magnetic storage medium, flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, a Europay, MasterCard and Visa (EMV) chip, a Subscriber Identity Module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, RAM, static RAM (SRAM), dynamic RAM (DRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASH EPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0096] Storage device 630 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 610, causes the system to perform functions. In some aspects, hardware services that perform specific functions may include software components for implementing functions stored in a computer-readable medium connected to necessary hardware components such as processor 610, connection 605, output device 635, etc. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection.

[0097] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or transient electronic signals that are propagated wirelessly or on a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may store thereon code and / or machine-executable instructions that may represent any combination of a process, function, subroutine, program, routine, subroutine, module, engine, software package, class, or instruction, data structure, or program statement. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.

[0098] In some aspects, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer readable storage media specifically excludes media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0099] Specific details are provided in the above description to provide a thorough understanding of the aspects provided herein. However, it will be appreciated by those skilled in the art that these aspects can also be practiced without these specific details. For clarity of explanation, in some cases, the present technology may be presented as including a separate functional block, which includes a device, device component, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein may be used. In some respects, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid confusing these aspects in unnecessary details. In other instances, known circuits, processes, algorithms, structures and techniques may be shown without unnecessary details to avoid confusing various aspects.

[0100] Various aspects may be described above as a process or method, which is depicted as a flow chart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flow chart may describe an operation as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. The process is terminated when the operations of the process are completed, but the process may have additional steps not included in the accompanying drawings. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to a calling function or a main function.

[0101] The process and method according to the above-mentioned example can be realized using the computer executable instructions stored or the computer executable instructions otherwise obtained from the computer readable medium. In some respects, such instructions may include instructions and data that make a general-purpose computer, a special-purpose computer or a processing device perform a certain function or a functional group or otherwise configure a general-purpose computer, a special-purpose computer or a processing device to perform a certain function or a functional group. The part of the computer resources used can be accessed through the network. In some respects, the computer executable instructions can be binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer readable media that can be used to store instructions, information used and / or information created during the method according to the example include disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0102] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. According to some aspects, such functionality may also be implemented on circuit boards among different chips or different processes executed on a single device.

[0103] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0104] In the above description, various aspects of the present application are described with reference to their specific aspects, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary aspects of the present application have been described in detail herein, it is to be understood that each inventive concept can be implemented and adopted in various other ways, and the appended claims are not intended to be interpreted as including these variations, unless limited by the prior art. Various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, various aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader essence and scope of this specification. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For the purpose of illustration, each method is described in a specific order. It should be appreciated that, in alternative aspects, each method can be performed in a different order than described.

[0105] One of ordinary skill in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.

[0106] Where a component is described as being “configured to” perform certain operations, in some aspects such configuration may be achieved by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0107] The phrase "coupled to" means that any component is directly or indirectly physically connected to another component, and / or any component is directly or indirectly in communication with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0108] Claim language or other language in the present disclosure that states "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. In some aspects, claim language that states "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In some aspects, claim language that states "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. In some aspects, claim language that states "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0109] Claim language or other language that recites "at least one processor, the at least one processor configured to," "at least one processor configured to," "one or more processors, the one or more processors configured to," etc., indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language that recites "at least one processor, the at least one processor configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or multiple processors are each tasked with a particular subset of operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that recites "at least one processor, the at least one processor configured to: X, Y, and Z" can mean that any single processor can perform only at least a subset of operations X, Y, and Z.

[0110] Where one or more elements are referred to as performing a function (e.g., steps of a method), one element may perform all functions, or more than one element may perform the functions together. When more than one element performs the functions together, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed entirely by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where one or more elements are referred to as being configured to cause another element (e.g., a device) to perform a function, one element may be configured to cause another element to perform all functions, or more than one element may be configured together to cause another element to perform the functions.

[0111] In the case of referring to an entity (e.g., any entity or device described herein) that performs a function or is configured to perform a function (e.g., a step of a method), the entity may be configured to cause one or more elements to perform these functions (individually or collectively). One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) functions of these functions, and / or any combination thereof. In referring to an entity that performs a function, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform these functions collectively. When the entity is configured to cause more than one component to perform these functions collectively, each function does not need to be performed by each component in those components (e.g., different functions may be performed by different components) and / or each function does not need to be performed by only one component as a whole (e.g., different sub-functions of executable functions of different components).

[0112] The various exemplary logic blocks, modules, engines, circuits, and algorithmic steps described in conjunction with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate the interchangeability of hardware and software, various exemplary components, blocks, modules, engines, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. The technician may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of the present application.

[0113] The technology described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such technology may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices, mobile phones, and other devices. Any features described as modules, engines, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the technology may be implemented at least in part by a computer-readable data storage medium including a program code, which includes instructions for executing one or more of the above methods, algorithms, and / or operations when executed. A computer-readable data storage medium may form part of a computer program product, which may include packaging materials. A computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0114] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0115] Exemplary aspects of the present disclosure include:

[0116] Aspect 1. A processor-implemented method for processing image data, the method comprising: receiving a first feature set associated with an image at a convolution engine of a machine learning system, the first feature set being associated with a three-dimensional shape; applying a depth-wise separable convolution filter to the first feature set via the convolution engine to generate a first output; applying a point-wise convolution filter to the first output via the convolution engine to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; modifying the second output to the three-dimensional shape to generate a second feature set; and combining the first feature set and the second feature set to generate an output feature set.

[0117] Aspect 2. A processor-implemented method according to aspect 1, wherein the convolution engine is configured to perform transformer operations.

[0118] Aspect 3. A processor-implemented method according to Aspect 1 or Aspect 2, wherein the convolution engine is configured to perform pair-wise self-attention and global feature extraction operations.

[0119] Aspect 4. A processor-implemented method according to any preceding aspect, the method further comprising: performing image classification associated with the image based on the output feature set via the machine learning system.

[0120] Aspect 5. A processor-implemented method according to any of the preceding aspects, the method further comprising: receiving a third feature set at a second convolution engine of the machine learning system, the third feature set being generated based on the output feature set; applying an additional depth-wise separable convolution filter to the third feature set via the second convolution engine to generate a third output; applying an additional point-wise convolution filter to the third output via the second convolution engine to generate a fourth output based on the global information from the spatial dimension and the channel dimension associated with the image; modifying the fourth output to the three-dimensional shape to generate a fourth feature set; and combining the third feature set and the fourth feature set to generate an additional output feature set.

[0121] Aspect 6. The processor-implemented method of any preceding aspect, wherein the first set of features is associated with a local representation of the image and the second set of features is associated with a global representation of the image.

[0122] Aspect 7. A processor-implemented method according to any preceding aspect, wherein the depthwise separable convolutional filter comprises a spatial multilayer perceptron that extracts information from a spatial domain of the first feature set to generate the first output.

[0123] Aspect 8. A processor-implemented method according to any preceding aspect, wherein the point-wise convolution filter comprises a channel multilayer perceptron that extracts information from the channel dimension of the first feature set to generate the second output.

[0124] Aspect 9. The processor-implemented method of any preceding aspect, wherein combining the first feature set and the second feature set to generate the output feature set comprises performing an element-wise cross product between the first feature set and the second feature set.

[0125] Aspect 10. A processor-implemented method according to any preceding aspect, wherein the first feature set comprises a first tensor, wherein the second feature set comprises a second tensor, and wherein the output feature set comprises an output tensor.

[0126] Aspect 11. A processor-implemented method according to any of the preceding aspects, the method further comprising: obtaining a first feature from a first intermediate feature map based on the output feature set, the first intermediate feature map having a first resolution; obtaining a second feature from a second intermediate feature map based on the output feature set, the second intermediate feature map having a second resolution different from the first resolution; combining the first feature and the second feature via a self-attention engine to generate a fused feature; and predicting a position of an object in the image based on the fused feature.

[0127] Aspect 12. A processor-implemented method according to any of the preceding aspects, wherein combining the first feature and the second feature comprises: performing a first dot product of each of the first features in the first intermediate feature map with each of the second features in the second intermediate feature map via the self-attention engine; performing a second dot product of each of the second features in the second intermediate feature map with each of the first features in the first intermediate feature map via the self-attention engine; applying a Softmax function to an output of the first dot product and an output of the second dot product via the self-attention engine; and performing a weighted summation of the output of the Softmax function via the self-attention engine to combine the first feature and the second feature.

[0128] Aspect 13. The processor-implemented method according to any preceding aspect, further comprising: performing a first prediction based on the first feature; performing a second prediction based on the second feature; and performing a third prediction based on the fused feature.

[0129] Aspect 14. A device for processing image data, the device comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: receive a first feature set associated with an image at a convolution engine of a machine learning system, the first feature set being associated with a three-dimensional shape; apply a depth-wise separable convolution filter to the first feature set via the convolution engine to generate a first output; apply a point-wise convolution filter to the first output via the convolution engine to generate a second output based on global information from a spatial dimension and a channel dimension associated with the image; modify the second output to the three-dimensional shape to generate a second feature set; and combine the first feature set and the second feature set to generate an output feature set.

[0130] Aspect 15. The apparatus according to aspect 14, wherein the convolution engine is configured to perform a transformer operation.

[0131] Aspect 16. An apparatus according to Aspect 15, wherein the convolution engine is configured to perform pairwise self-attention and global feature extraction operations.

[0132] Aspect 17. An apparatus according to aspect 15, wherein the at least one processor coupled to at least one memory is further configured to: perform image classification associated with the image based on the output feature set via the machine learning system.

[0133] Aspect 18. An apparatus according to Aspect 15, wherein the at least one processor coupled to at least one memory is further configured to: receive a third feature set at a second convolution engine of the machine learning system, the third feature set being generated based on the output feature set; apply an additional depth-wise separable convolution filter to the third feature set via the second convolution engine to generate a third output; apply an additional point-wise convolution filter to the third output via the second convolution engine to generate a fourth output based on the global information from the spatial dimension and the channel dimension associated with the image; modify the fourth output to the three-dimensional shape to generate a fourth feature set; and combine the third feature set and the fourth feature set to generate an additional output feature set.

[0134] Aspect 19. An apparatus according to any one of aspects 14 to 18, wherein the first feature set is associated with a local representation of the image and the second feature set is associated with a global representation of the image.

[0135] Aspect 20. An apparatus according to any one of Aspects 14 to 19, wherein the depthwise separable convolutional filter comprises a spatial multilayer perceptron that extracts information from a spatial domain of the first feature set to generate the first output.

[0136] Aspect 21. An apparatus according to any one of Aspects 14 to 20, wherein the point-wise convolution filter comprises a channel multilayer perceptron that extracts information from the channel dimension of the first feature set to generate the second output.

[0137] Aspect 22. An apparatus according to any one of aspects 14 to 21, wherein combining the first feature set and the second feature set to generate the output feature set comprises: performing an element-wise cross product between the first feature set and the second feature set.

[0138] Aspect 23. An apparatus according to any one of Aspects 14 to 22, wherein the first feature set comprises a first tensor, wherein the second feature set comprises a second tensor, and wherein the output feature set comprises an output tensor.

[0139] Aspect 24. An apparatus according to any one of Aspects 14 to 23, wherein the at least one processor coupled to at least one memory is further configured to: obtain a first feature from a first intermediate feature map based on the output feature set, the first intermediate feature map having a first resolution; obtain a second feature from a second intermediate feature map based on the output feature set, the second intermediate feature map having a second resolution different from the first resolution; combine the first feature and the second feature via a self-attention engine to generate a fused feature; and predict the position of an object in the image based on the fused feature.

[0140] Aspect 25. An apparatus according to any one of Aspects 14 to 24, wherein the at least one processor coupled to at least one memory is further configured to combine the first feature and the second feature by: performing a first dot product of each of the first features in the first intermediate feature map with each of the second features in the second intermediate feature map via the self-attention engine; performing a second dot product of each of the second features in the second intermediate feature map with each of the first features in the first intermediate feature map via the self-attention engine; applying a Softmax function or a similar function to an output of the first dot product and an output of the second dot product via the self-attention engine; and performing a weighted summation of the outputs of the Softmax function or the similar function via the self-attention engine to combine the first feature and the second feature.

[0141] Aspect 26. An apparatus according to any one of Aspects 14 to 25, wherein the at least one processor coupled to at least one memory is further configured to: perform a first prediction based on the first feature; perform a second prediction based on the second feature; and perform a third prediction based on the fused feature.

[0142] Aspect 27. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform the operations according to any one of aspects 1 to 26.

[0143] Aspect 28. An apparatus for processing image data, the apparatus comprising: one or more components for performing the operations according to any one of aspects 1 to 26.

Claims

1. A processor-implemented method for processing image data, the method comprising: receiving, at a convolution engine of a machine learning system, a first feature set associated with an image, the first feature set associated with a three-dimensional shape; applying, via the convolution engine, a depthwise separable convolutional filter to the first set of features to generate a first output; applying, via the convolution engine, a point-wise convolution filter to the first output to generate a second output based on global information from spatial and channel dimensions associated with the image; modifying the second output into the three-dimensional shape to generate a second set of features; as well as The first feature set and the second feature set are combined to generate an output feature set.

2. The processor-implemented method of claim 1 , wherein the convolution engine is configured to perform transformer operations.

3. The processor-implemented method of claim 1 , wherein the convolution engine is configured to perform pairwise self-attention and global feature extraction operations.

4. The processor-implemented method of claim 1 , further comprising: Image classification associated with the image is performed via the machine learning system based on the output feature set.

5. The processor-implemented method of claim 1 , further comprising: receiving, at a second convolution engine of the machine learning system, a third feature set, the third feature set generated based on the output feature set; applying, via the second convolution engine, an additional depthwise separable convolutional filter to the third feature set to generate a third output; applying, via the second convolution engine, an additional point-wise convolution filter to the third output to generate a fourth output based on the global information from the spatial dimension and the channel dimension associated with the image; modifying the fourth output into the three-dimensional shape to generate a fourth feature set; as well as The third feature set and the fourth feature set are combined to generate an additional output feature set. 6 . The processor-implemented method of claim 1 , wherein the first set of features is associated with a local representation of the image and the second set of features is associated with a global representation of the image.

7. The processor-implemented method of claim 1, wherein the depthwise separable convolutional filter comprises a spatial multilayer perceptron that extracts information from a spatial domain of the first feature set to generate the first output.

8. The processor-implemented method of claim 1, wherein the point-wise convolutional filter comprises a channel-wise multilayer perceptron that extracts information from a channel dimension of the first feature set to generate the second output.

9. The processor-implemented method of claim 1 , wherein combining the first feature set and the second feature set to generate the output feature set comprises: An element-wise cross product between the first feature set and the second feature set is performed.

10. The processor-implemented method of claim 1, wherein the first feature set comprises a first tensor, wherein the second feature set comprises a second tensor, and wherein the output feature set comprises an output tensor.

11. The processor-implemented method of claim 1 , further comprising: Obtaining a first feature from a first intermediate feature map based on the output feature set, the first intermediate feature map having a first resolution; obtaining a second feature from a second intermediate feature map based on the output feature set, the second intermediate feature map having a second resolution different from the first resolution; combining the first feature and the second feature via a self-attention engine to generate a fused feature; as well as The location of an object in the image is predicted based on the fused features.

12. The processor-implemented method of claim 11 , wherein combining the first feature and the second feature comprises: performing, via the self-attention engine, a first dot product of each of the first features in the first intermediate feature map and each of the second features in the second intermediate feature map; performing, via the self-attention engine, a second dot product of each of the second features in the second intermediate feature map with each of the first features in the first intermediate feature map; Applying a Softmax function to an output of the first dot product and an output of the second dot product via the self-attention engine; as well as A weighted summation of the output of the Softmax function is performed via the self-attention engine to combine the first feature and the second feature.

13. The processor-implemented method of claim 11 , further comprising: performing a first prediction based on the first feature; performing a second prediction based on the second feature; as well as A third prediction is performed based on the fused features.

14. A device for processing image data, the device comprising: at least one memory; and at least one processor coupled to at least one memory and configured to: Receiving, at a convolution engine of a machine learning system, a first feature set associated with an image and having a three-dimensional shape; applying, via the convolution engine, a depthwise separable convolutional filter to the first feature set to generate a first output; applying, via the convolution engine, a point-wise convolution filter to the first output to generate a second output based on global information from spatial and channel dimensions associated with the image; modifying the second output into the three-dimensional shape to generate a second set of features; as well as The first feature set and the second feature set are combined to generate an output feature set.

15. The apparatus of claim 14, wherein the convolution engine is configured to perform a transformer operation.

16. The apparatus of claim 14, wherein the convolution engine is configured to perform pairwise self-attention and global feature extraction operations.

17. The apparatus of claim 14, wherein the at least one processor is further configured to: Image classification associated with the image is performed via the machine learning system based on the output feature set.

18. The apparatus of claim 14, wherein the at least one processor is further configured to: receiving, at a second convolution engine of the machine learning system, a third feature set, the third feature set generated based on the output feature set; applying, via the second convolution engine, an additional depthwise separable convolutional filter to the third feature set to generate a third output; applying, via the second convolution engine, an additional point-wise convolution filter to the third output to generate a fourth output based on the global information from the spatial dimension and the channel dimension associated with the image; modifying the fourth output into the three-dimensional shape to generate a fourth feature set; as well as The third feature set and the fourth feature set are combined to generate an additional output feature set.

19. The apparatus of claim 14, wherein the first set of features is associated with a local representation of the image and the second set of features is associated with a global representation of the image.

20. The apparatus of claim 14, wherein the depthwise separable convolutional filter comprises a spatial multilayer perceptron that extracts information from a spatial domain of the first feature set to generate the first output.

21. The apparatus of claim 14, wherein the point-wise convolutional filter comprises a channel-wise multilayer perceptron that extracts information from a channel dimension of the first feature set to generate the second output.

22. The apparatus of claim 14, wherein combining the first feature set and the second feature set to generate the output feature set comprises: An element-wise cross product between the first feature set and the second feature set is performed.

23. The apparatus of claim 14, wherein the first feature set comprises a first tensor, wherein the second feature set comprises a second tensor, and wherein the output feature set comprises an output tensor.

24. The apparatus of claim 14, wherein the at least one processor is further configured to: Obtaining a first feature from a first intermediate feature map based on the output feature set, the first intermediate feature map having a first resolution; obtaining a second feature from a second intermediate feature map based on the output feature set, the second intermediate feature map having a second resolution different from the first resolution; combining the first feature and the second feature via a self-attention engine to generate a fused feature; as well as The location of an object in the image is predicted based on the fused features.

25. The apparatus of claim 24, wherein to combine the first feature and the second feature, the at least one processor is further configured to: performing, via the self-attention engine, a first dot product of each of the first features in the first intermediate feature map and each of the second features in the second intermediate feature map; performing, via the self-attention engine, a second dot product of each of the second features in the second intermediate feature map with each of the first features in the first intermediate feature map; Applying a Softmax function to the output of the first dot product and the output of the second dot product via the self-attention engine; and A weighted summation of the output of the Softmax function is performed via the self-attention engine to combine the first feature and the second feature.

26. The apparatus of claim 24, wherein the coupled at least one processor is further configured to: performing a first prediction based on the first feature; performing a second prediction based on the second feature; and A third prediction is performed based on the fused features.

27. A non-transitory computer readable memory storing instructions that cause at least one processor coupled to the non-transitory computer readable memory to be configured to: Receiving, at a convolution engine of a machine learning system, a first feature set associated with an image and having a three-dimensional shape; applying, via the convolution engine, a depthwise separable convolutional filter to the first feature set to generate a first output; applying, via the convolution engine, a point-wise convolution filter to the first output to generate a second output based on global information from spatial and channel dimensions associated with the image; modifying the second output into the three-dimensional shape to generate a second set of features; as well as The first feature set and the second feature set are combined to generate an output feature set.

28. The non-transitory computer readable memory of claim 27, wherein the convolution engine is configured to perform transformer operations.