Multidimensional Spatial Decomposition for Transformer Neural Networks
Decomposing multidimensional inputs into two-dimensional subspaces with shared dimensions in transformer neural networks addresses computational inefficiencies, enabling efficient processing with reduced resources and maintaining accuracy in edge devices.
Patent Information
- Application Number
- BR112025019672
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2024-01-31
- Publication Date
- 2026-07-28
AI Technical Summary
Transformer neural networks are computationally expensive for processing multidimensional data, particularly in edge devices with limited resources, leading to inefficiencies in tasks requiring real-time processing like autonomous operations.
Decompose multidimensional inputs into multiple two-dimensional subspaces sharing a common dimension to reduce computational complexity, using attention matrices generated for each subspace to process data efficiently.
This approach reduces computational complexity to subquadratic scaling, allowing efficient processing of multidimensional data with fewer resources and maintaining inference accuracy, especially in edge devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
1 / 29 Multidimensional Spatial Decomposition for Transformer Neural Networks REFERENCE TO RELATED DEPOSIT REQUEST(S)
[0001] This application claims the benefit and priority of U.S. Patent Application No. 18 / 193,234, filed March 30, 2023, which is hereby incorporated by reference in its entirety. INTRODUCTION
[0002] The aspects of this disclosure relate to neural networks and, more particularly, to multidimensional content processing using neural networks.
[0003] Various machine learning architectures have been used to provide solutions to a wide variety of computational problems. There is a variety of machine learning model architectures, such as artificial neural networks (which may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks, generative adversarial networks (GANs), and the like), random forest models, and the like. Increasingly, transformer neural networks have been widely used in a variety of image and video processing tasks, or other tasks in which multidimensional data is processed in order to generate various inferences related to the multidimensional data.
[0004] However, transformer neural networks tend to be computationally expensive. For example, because vision transformers typically compute self-attention on each block, the computational and memory demands can increase quadratically with respect to the size of the input data. While the computational complexity of these tasks can be addressed through the use of high-performance processing units (e.g., graphics processing units, neural processing units, and / or other processing units that Petition 870250083079, dated 09 / 15 / 2025, pp. 112 / 154 2 / 29 support high degrees of parallelism) and large amounts of memory, edge devices (e.g., user equipment (UEs), such as mobile devices or autonomous vehicles, etc.) may not have sufficient computational resources to process multidimensional content (e.g., video content with two spatial dimensions (height and width) and one temporal dimension) at a performance level sufficient for the applications in which the processing must be performed. For example, these edge devices may not have sufficient computing resources to satisfy real-time or near-real-time timing constraints for applications such as autonomous operations (e.g., self-driving cars, movement in restricted environments, etc.).
[0005] Consequently, improved techniques are needed to efficiently process multidimensional content using neural networks. BRIEF SUMMARY
[0006] Certain aspects provide a processor-implemented method for processing multidimensional content using neural networks. The example method generally involves decomposing a multidimensional input into a plurality of two-dimensional subspaces, where the plurality of two-dimensional subspaces share a common dimension. A first attention matrix is generated based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network, and a second attention matrix is generated based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network. An output of the transformer neural network is generated based on a combination of the first attention matrix and the second attention matrix.
[0007] Other aspects provide configured processing systems Petition 870250083079, dated 09 / 15 / 2025, pp. 113 / 154 3 / 29 to perform the aforementioned methods, as well as those described in the present invention; non-transient computer-readable means comprising instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the aforementioned methods, as well as those described in the present invention; a computer program product embedded in a computer-readable storage medium comprising code to perform the aforementioned methods, as well as those described in the present invention; and a processing system comprising means for performing the aforementioned methods, as well as those further described in the present invention.
[0008] The following description and related drawings set out in detail certain illustrative attributes of one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The attached figures depict example attributes of certain aspects of this disclosure and, therefore, should not be considered limiting the scope of this disclosure.
[0010] Figure 1 illustrates an example transformer neural network architecture in which multidimensional data is processed.
[0011] Figure 2 illustrates several pipelines for processing multidimensional data in transformer neural networks.
[0012] Figure 3 illustrates example operations for processing multidimensional input through a transformer neural network based on the decomposition of multidimensional input into two-dimensional subspaces, according to aspects of the present disclosure.
[0013] Figure 4 depicts an example processing system configured to perform various aspects of the present disclosure.
[0014] To facilitate understanding, identical reference numbers have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that the elements and attributes of one aspect may be beneficially incorporated into other aspects without Petition 870250083079, dated 09 / 15 / 2025, pages 114 / 154 4 / 29 additional mention. DETAILED DESCRIPTION
[0015] Aspects of the present disclosure provide computer-readable apparatus, methods, processing systems and non-transient means for processing multidimensional inputs using transformer-based neural networks and decomposing multidimensional inputs into multiple two-dimensional subspaces. As used in the present invention, the term multidimensional generally refers to three or more dimensions (e.g., at least height, width and temporal).
[0016] Various types of neural networks can be used to process visual content (e.g., detecting objects, predicting the future movement of detected objects in visual content, segmenting visual content into different semantic groups, etc.), such as still images or streams of visual content (e.g., video content captured as a series of images at a given frame rate, such as 24 frames per second, 29.97 frames per second, 60 frames per second, etc.). However, these neural networks generally process visual content on a frame-by-frame basis, which can be a computationally expensive process that increases in complexity as the frame size of each frame in the visual content increases.
[0017] Transformer neural networks (also called transformers), and in particular vision transformers, have become increasingly common in a wide variety of machine learning tasks. Transformer-based architectures are generally configured to generate output based on a sequence of data (e.g., a sequence of frames in a video, a sequence of patches of a frame or image, and so on). In general, machine learning models can use any number of transformer blocks (each providing self-attention), as well as any other components (e.g., one or more neural network layers).
[0018] For multidimensional data, such as video data with two Petition 870250083079, dated 09 / 15 / 2025, pages 115 / 154 5 / 29 spatial dimensions and one temporal dimension, processing using a transformer neural network can be computationally expensive due to the amount of data to be processed. This complexity can increase significantly as the amount of data to be processed increases; for example, in video content, the amount of data to be processed can scale quadratically in both spatial and temporal dimensions. Due to the transformer block structure in a transformer neural network, the computational cost involved in processing inputs in any given transformer block can scale quadratically with the size of the input; for example, doubling the resolution for a video input while keeping the video input duration constant can result in a sixteen-fold increase in the computational cost of processing that video input through a transformer neural network.Furthermore, since a transformer neural network can include multiple transformer blocks, each transformer can perform computations on input data independently, incurring significant computational cost.
[0019] Aspects of the present disclosure provide techniques for reducing the computational cost of processing multidimensional input data in transformer neural networks. As discussed in more detail in the present invention, to reduce the computational complexity of processing a multidimensional input, aspects of the present disclosure decompose a multidimensional input into multiple two-dimensional subspaces that may share a common dimension; for example, in a video input with height, width, and temporal dimensions, the two-dimensional subspaces may be a height temporal subspace and a width temporal subspace. By decomposing a multidimensional input into multiple two-dimensional spaces sharing a common dimension, the computational cost involved in processing a multidimensional input in a transformer neural network can be reduced and can have subquadratic scaling. Thus, a quantity Petition 870250083079, dated 09 / 15 / 2025, pages 116 / 154 6 / 29 Less computing resources can be used to complete various tasks for which transformer neural networks are used, such as object detection or other computer vision tasks. In turn, this can reduce the amount of power used by computing devices to perform these tasks and / or speed up the processing of multidimensional inputs, relative to the amount of power and / or time used when a multidimensional input is not decomposed into multiple two-dimensional spaces for processing. Example transformer architecture
[0020] Figure 1 illustrates an example transformer architecture 100 in which attention data is propagated through a transformer block in a neural network (e.g., to another transformer block(s) in the network) in order to generate a neural network output.
[0021] As illustrated in Figure 1, the input data 105 is accessed by a transformer 110. As used in the present invention, access to the data may generally include receiving, retrieving, requesting, or otherwise obtaining access to the data. As discussed above, the input data 105 may correspond to the input (e.g., raw or pre-processed input data) to the first transformer block of a model, the output of a previous transformer or other model block or component, and the like. For example, the input data 105 may correspond to a multidimensional input, a tokenized version of the multidimensional input (which may optionally include positional embedding(s) and / or symbolic item(s) that may be learned), or the like.The tokenized version of the multidimensional input can also be called the attribute set for the multidimensional input generated in different portions of the multidimensional input (for example, different spatial portions, or patches, of the multidimensional input at various points in time).
[0022] Generally, the 110 transformer includes a 120 self-attention block (identified as SA) and a 140 direct power supply block (identified Petition 870250083079, dated 09 / 15 / 2025, pp. 117 / 154 7 / 29 as FF). In the self-attention block 120, the input data 105 can be linearly projected (i.e., multiplied using learned parameters) into three matrices: a query matrix Q 122 (also called, in some respects, query representation or simply queries), a key matrix K 124 (also called, in some respects, key representation or simply keys), and a value matrix V 126 (also called, in some respects, value representation or simply values). For example, during training, one or more query weights, key weights, and value weights are learned based on training data, and the queries Q 122, the keys K 124, and the values V 126 can be generated by multiplying the input data by the learned weights.
[0023] In some respects, an attention matrix A 130 (also called an attention map or simply attention in some respects) is then generated based on the queries and keys. For example, the self-attention block can, in operation 128, compute the dot product of the query matrix and the transposed key matrix (e.g., Q · KT). In some respects, the self-attention block can apply one or more operations (e.g., a row-by-row softmax operation) to the dot product to produce this attention matrix. That is, the attention matrix A 130 can be defined as A = a(Q · KT), where σ is the softmax function (or some other regularization function usable in a transformer neural network).
[0024] The resulting attributes f 134 generated by the self-attention block can then be computed, in operation 132, as the dot product of the attention matrix A 130 and the value matrix V 126. These attributes can then be provided as input to the forward feed block 140 (e.g., a neural network or subnetwork) to generate an output 150 from the transformer 110. The output 150 can be used as an input to a subsequent transformer or another block in a neural network, or it can be the final result of processing an input through the neural network. The forward feed block 140, in some respects, can be a perceptron. Petition 870250083079, dated 09 / 15 / 2025, pp. 118 / 154 8 / 29 multilayer perceptron (MLP) including a plurality of layers separated by an activation function, such as a linear unit Gaussian error activation function.
[0025] Although not depicted in Figure 1, in some respects, transformer 110 may include one or more skip or residual connections (with or without layer normalization). For example, output 150 may be generated by summing the output of operation 132 with input data 105, skipping the direct feed block 140. As another example, output 150 of transformer 110 may be generated by summing the output of the final layer of the direct feed block 140 with the previous version of output 150. Subspatial decomposition of an example of multidimensional inputs for processing in transformer neural networks.
[0026] As discussed, transformer neural networks can be used to process multidimensional inputs and generate inferences based on the processing of these multidimensional inputs. For example, these transformer neural networks can be used in performing various operations on video data, such as video enhancement (e.g., noise reduction, upscaling through super-resolution techniques that increase or otherwise enhance the resolution of an input), object detection, three-dimensional vision, medical imaging, or the like. In object detection tasks, for example, the outputs generated by transformer neural networks can be used to semantically segment an input into different segments associated with different levels of importance to the overall meaning of the scene and select different portions of the scene for monitoring (e.g., corresponding to different objects).The outputs generated by transformer neural networks can also be used, for example, to predict the movement of objects in a scene, which can then be used to apply various control inputs to an autonomous or semi-autonomous vehicle to ensure that the vehicle does not collide with these objects (or at least reduce the probability of the vehicle colliding with these objects). In the example of three-dimensional vision, transformer neural networks can be used to recreate... Petition 870250083079, dated 09 / 15 / 2025, pp. 119 / 154 9 / 29 environments in three-dimensional space based on truncated signed distance function (TSDF) data or similar. In the medical imaging example, transformer neural networks can be used for three-dimensional data segmentation to identify various structures in captured medical images, such as blood vessels, tumors, and the like.
[0027] In many of these use cases, the information captured over time can be useful data that can be used in generating various inferences from multidimensional inputs. For example, in computer vision tasks, motion information captured over time can be highly useful data for processing video content. Due to the nature of transformer neural networks, temporal dependencies (or relationships) between different portions of a multidimensional input (e.g., different video frames with different timestamps) can be used, for example, to identify moving objects, stationary objects, spatial relationships between different objects in a frame, and so on. However, as discussed above, processing multidimensional inputs using transformer neural networks can be a computationally expensive task.In a transformer neural network, each symbolic item of an input, corresponding, for example, to different distinct portions of the input, such as patches in an image (e.g., a pixel or a contiguous block of pixels in an image), can be compared with other symbolic items generated for the input. Due to these comparisons, the computational complexity involved in processing multidimensional inputs using a transformer neural network can quadratically scale with input resolution and with time (or time resolution, such as a frame rate at which video frames are captured). Thus, while some devices may be capable of processing multidimensional inputs using transformer neural networks, it may not be practical to deploy transformer neural networks. Petition 870250083079, dated 09 / 15 / 2025, pp. 120 / 154 10 / 29 to process these multidimensional inputs in other devices (for example, user equipment (UEs) in a wireless communications network, Internet of Things (IoT) devices, or other power-constrained devices that may have limited computing capabilities or computing capabilities restricted by the amount of power that can be drawn from an energy storage device.
[0028] To improve the efficiency of transformer neural networks and reduce the computational complexity involved in processing inputs using transformer neural networks, several techniques can be used. In some examples, attention can be computed on a subset of symbolic items (instead of the entirety of the set of symbolic items generated for an input) to achieve subquadratic computation complexity scaling for inputs in the neural network. This can be accomplished, for example, based on sparsity constraints, block processing, axial processing, linear decomposition of multidimensional inputs into a one-dimensional input, other decompositions, or the like. In one example, space and time can be independently decomposed in order to reduce the computational complexity of processing a video input in a transformer neural network.In this example, the cost of decomposition can be relatively small for low-resolution inputs; however, the computational complexity of processing an input can still be high for high-resolution inputs. In spatial and temporal decomposition separately, however, temporal information may not be accessible during the processing of spatial content, even though such temporal information includes significant amounts of contextual data useful for processing video data and performing various actions in relation to a video data processing output through a transformer neural network.
[0029] To further reduce the computational complexity of Petition 870250083079, dated 09 / 15 / 2025, pages 121 / 154 11 / 29 Processing video data or other multidimensional data through a transformer neural network, aspects of the present disclosure decompose multidimensional data (e.g., with three or more dimensions) into multiple two-dimensional subspaces sharing a common dimension. For example, video data with two spatial dimensions (height and width) and one temporal dimension can be decomposed into two two-dimensional subspaces: a first subspace with width and temporal dimensions and a second subspace with height and temporal dimensions, where time is the common dimension.By decomposing multidimensional data into multiple two-dimensional subspaces sharing a common dimension, the computational complexity of processing such data can be uniformly distributed, and the two-dimensional subspaces sharing a common dimension can access significant amounts of contextual data useful in self-awareness or other operations in the transformer neural network. In doing so, aspects of the present disclosure provide computationally efficient processing of multidimensional data, with efficiencies relative to other decomposition techniques (e.g., as discussed in more detail below) increasing as the input resolution increases.Furthermore, aspects of the present disclosure may provide similar inference accuracy (e.g., peak signal-to-noise ratio (PSNR)) using fewer computing resources than spatial-temporal decomposition, may provide greater inference accuracy using computing resources similar to other decomposition techniques, and may be computationally cheaper as the number of samples (e.g., frames in video content) used in a task increases relative to other decomposition techniques.
[0030] Figure 2 illustrates pipelines for processing multidimensional data in transformer neural networks. Although the pipelines in Figure 2 illustrate transformer pipelines in the context of processing video data with dimensions of height, width, and Petition 870250083079, dated 09 / 15 / 2025, pages 122 / 154 12 / 29 temporal, it should be recognized that the transformer pipelines illustrated in Figure 2 can be applied to n-dimensional data processing (with the addition or removal of attention heads in the pipeline). Furthermore, although the pipelines illustrated in Figure 2 illustrate several sequential operations for processing multidimensional data in transformer neural networks, it should be recognized that several operations illustrated in these pipelines can be performed in parallel or substantially in parallel.
[0031] Pipeline 210 illustrates a pipeline for processing multidimensional data without decomposing a multidimensional input into different subspaces. As illustrated, the multidimensional input can be processed through an attention head 212 of a transformer neural network (e.g., self-attention block 120 illustrated in Figure 1). The output of the attention head 212 (in combination with the input) of the transformer neural network can be processed through a feedforward network 214. The output of the feedforward network 214 and the sum of the output of the attention head 212 with the input can be combined to generate an output of the transformer neural network.Since multidimensional data is not decomposed into smaller portions, there can be H * W * T symbolic items generated for the multidimensional input, where H and W correspond to a number of patches in the height domain and a number of patches in the width domain, respectively, into which the input is divided, and each symbolic item can be compared with all other symbolic items generated for the multidimensional input. Thus, pipeline 210 can complete the processing of the multidimensional input in time 0(H2W2T2).
[0032] Pipeline 220 illustrates a pipeline for processing multidimensional data with spatial-temporal decomposition of a multidimensional input including spatial and temporal components. In this example, the temporal component of a multidimensional input can be processed through a first attention head 222 of a transformer neural network and the spatial component(s) of the multidimensional input. Petition 870250083079, dated 09 / 15 / 2025, pp. 123 / 154 13 / 29 can be processed through a second attention head 224 of the transformer neural network. As illustrated, the input to the second attention head 224 of the transformer neural network can be the sum of the output of the first attention head 222 with the input. The sum of the output of the first attention head, with the input and output of the second attention head 224 can subsequently be processed through a feedforward network 226. The output of the feedforward network 226 and the sum of the output of the first attention head, with the input and output of the second attention head 224 can be combined to generate an output of the transformer neural network. The computational complexity of processing the temporal component of the multidimensional input can be O(T2), and the computational complexity of processing the spatial component of the multidimensional input can be O(H2W2').Since the computational complexities involved in processing the temporal and spatial components of the multidimensional input are independent of each other (and therefore additive, not multiplicative), pipeline 220 can complete the processing of the multidimensional input in time 0(HWT2+ H2W2T), which represents a significant decrease in the computational complexity of processing a multidimensional input compared to the complexity of pipeline 210 discussed above.
[0033] However, as discussed above, further improvements in both computational complexity and accuracy of processing multidimensional inputs in a transformer neural network can be achieved by decomposing a multidimensional input into multiple two-dimensional subspaces sharing a common dimension. Pipeline 230 illustrates a pipeline for processing multidimensional data based on the decomposition of a multidimensional input into multiple two-dimensional subspaces sharing a common dimension, according to aspects of the present disclosure. In this example, a multidimensional input may have a height, a width, and a time component. Petition 870250083079, dated 09 / 15 / 2025, pages 124 / 154 14 / 29 Thus, a decomposition of this multidimensional input into a plurality of two-dimensional subspaces can be achieved by decomposing the multidimensional input into a first two-dimensional subspace including the height and time dimensions and a second two-dimensional subspace including the width and time dimensions.
[0034] As illustrated, a first attention head 232 can generate a first attention matrix based on a projection of the data in the first two-dimensional subspace (e.g., symbolic items generated for individual elements in the first two-dimensional subspace). A second attention head 234 can similarly generate a second attention matrix based on a projection of the data in the second two-dimensional subspace. The outputs of the first attention head 232 and the second attention head 234 can be combined (e.g., through a summation operation) with the input and provided as an input to a direct feed network 236.
[0035] The computational complexity of processing the data in the first two-dimensional subspace through the first attention head 232 can be O(H2T2'). Similarly, the computational complexity of processing the data in the second two-dimensional subspace through the second attention head 234 can be O(W'2T'2). As with pipeline 220, since data processing in attention heads 232 and 234 involves independent operations, the total computational complexity is additive and not multiplicative. Thus, pipeline 230 can complete the processing of the multidimensional input in O(WH2T2 + HW2T2') time, which can provide further improvements in the computational complexity of processing multidimensional data (with three or more dimensions) in a transformer neural network.Furthermore, unlike what happens in pipeline 220, attention heads 232 and 234 in pipeline 230 process data with a shared dimension, which allows useful contextual information (e.g., in the time domain) to be used in generating the attention matrices for both subspaces, instead of discarding such information. Petition 870250083079, dated 09 / 15 / 2025, pages 125 / 154 15 / 29 one of the attention heads, which, in turn, can enable reduced computational complexity in various tasks involving multidimensional data (video enhancement (e.g., noise removal, super-resolution, etc.), video classification, semantic segmentation, object detection, three-dimensional reconstruction, anatomical segmentation, etc.), allowing a small number of consecutive samples (e.g., frames) to be used to perform various tasks. Although Figure 2 illustrates that the height temporal subspace is processed before the width temporal subspace, it should be recognized that data in different subspaces can be processed in any sequential order (e.g., so that attention heads 232 and 234 are swapped in order) or can be processed in parallel or substantially in parallel.
[0036] As mentioned above, pipeline 230 can be used in data processing in any multidimensional space by decomposing the data into a plurality of two-dimensional subspaces sharing a common dimension. For example, a four-dimensional input (e.g., three spatial dimensions and one temporal dimension) can be decomposed into three two-dimensional subspaces sharing the temporal dimension as the common dimension, and the resulting transformer neural network, through which the four-dimensional input is processed, can include three attention heads (e.g., one for each two-dimensional subspace).Examples of data in a multidimensional space might include, for example, an audiovisual input including spatial, frequency, and time dimensions; an input including different spatial environments in which operations are performed along a common temporal dimension (for example, an input in a virtual reality environment in which each spatial environment corresponds to the data displayed on one of a plurality of time-synchronized displays); and so on.
[0037] Although the examples mentioned above illustrate time as a common dimension shared by the plurality of two-dimensional subspaces, it must be recognized that any other Petition 870250083079, dated 09 / 15 / 2025, pages 126 / 154 16 / 29 The appropriate dimension (e.g., frequency) can also or alternatively be used as the common dimension shared by the plurality of two-dimensional subspaces.
[0038] For an input of T samples in a D-dimensional space decomposed into a plurality of two-dimensional subspaces sharing a common dimension, a transformer neural network can complete processing the input in time 0(0 x SD+1T2'), in contrast to time O(TS2D+ SDT2') using spacetime decomposition (e.g., as in pipeline 220 discussed above) or time O(S2DT2') without decomposition. As discussed, processing multidimensional data based on decomposing such data into a plurality of two-dimensional subspaces sharing a common dimension can be significantly less computationally expensive than processing multidimensional data using spacetime decomposition, since processing data in multidimensional subspaces can reduce the amount of redundant data processed in a transformer neural network.These efficiencies can be seen in relation to the computational cost incurred in data processing, as it may be possible to generate usable inferences from multidimensional data using a smaller number of samples (e.g., frames in an input video segment) compared to other decomposition techniques. Furthermore, it can be noted that significant decreases in computational cost can be achieved by processing a smaller number of samples using a transformer neural network.
[0039] The reduction in computational complexity achieved by decomposing multidimensional data into a plurality of two-dimensional subspaces sharing a common dimension compared to processing multidimensional data without decomposition (joint attention) can be represented by the expression: "(W) Petition 870250083079, dated 09 / 15 / 2025, pages 127 / 154 17 / 29 Formula 1 where S represents the size of a non-temporal dimension in multidimensional space and D represents the number of spatial dimensions in multidimensional space.
[0040] The reduction in computational complexity achieved by decomposing multidimensional data into a plurality of two-dimensional subspaces sharing a space-time decomposition is represented by the expression: A common dimension in relation to multidimensional data can be SD SxDxT Formula 2 where T represents the number of samples in the multidimensional data.
[0041] The reduction in computational cost can be scaled based on the number of spatial dimensions included in the multidimensional input. For example, for a sequence of two-dimensional images, the computational complexity involved in processing this sequence through a transformer neural network can be reduced by O(S) time compared to joint attention and can be reduced by O(-) time compared to spatial-temporal decomposition. For a sequence of three-dimensional volumetric data, the computational complexity involved in processing this data through a transformer neural network can be reduced by O(^) time compared to joint attention and can be reduced by S3T time. O(—~) in relation to spatial-temporal decomposition.
[0042] Figure 3 illustrates example 300 operations for processing multidimensional input through a transformer neural network based on the decomposition of the multidimensional input into two-dimensional subspaces, according to aspects of the present disclosure. The 300 operations can be performed, for example, by a computing system in which a transformer neural network is deployed to process multidimensional data, such as a Petition 870250083079, dated 09 / 15 / 2025, pages 128 / 154 18 / 29 user equipment (UE), a smartphone, a tablet computer, an autonomous vehicle, an edge device or other computing system (for example, such as the 400 processing system illustrated in Figure 4 and described in more detail below).
[0043] As illustrated, operations 300 begin in block 310 with the decomposition of a multidimensional input into a plurality of two-dimensional subspaces. The plurality of two-dimensional subspaces share a common dimension. Since two-dimensional subspaces generally share a common dimension, the number of two-dimensional subspaces into which an n-dimensional input can be decomposed can be n - 1. For example, a three-dimensional space in which an input resides can be decomposed into two two-dimensional subspaces.
[0044] In some respects, multidimensional input may include an input with a plurality of spatial dimensions and a temporal dimension. For example, multidimensional input may be a video clip including a plurality of frames (the temporal dimension) including two-dimensional spatial data (e.g., height and width). In this case, the first two-dimensional subspace may be a subspace based on a first spatial dimension among the plurality of spatial dimensions and on the temporal dimension, and the second two-dimensional subspace may be a subspace based on a second spatial dimension among the plurality of spatial dimensions and on the temporal dimension.
[0045] In block 320, operations 300 proceed with the generation of a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network.
[0046] In some respects, the generation of the first attention matrix based on the projection of symbolic items onto the first two-dimensional subspace involves projecting the first two-dimensional subspace onto query, key, and value data. The first attention matrix can be generated Petition 870250083079, dated 09 / 15 / 2025, pages 129 / 154 19 / 29 based on query data, key data, and various temporal components in the first two-dimensional subspace.
[0047] In block 330, operations 300 proceed with the generation of a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network.
[0048] In some respects, the generation of the second attention matrix based on the projection of symbolic items onto the second two-dimensional subspace involves projecting the second two-dimensional subspace onto query, key, and value data. The second attention matrix can be generated based on query data, key data, and various temporal components in the second two-dimensional subspace.
[0049] In block 340, operations 300 proceed with the generation of an output from the transformer neural network based on the first attention matrix and the second attention matrix (e.g., a combination of the first attention matrix and the second attention matrix). The combination can be a sum, for example.
[0050] In some respects, generating the output of the transformer neural network involves computing a first attribute based on the first attention matrix and projected values from the symbolic items in the first two-dimensional subspace. A second attribute can be computed based on the second attention matrix and projected values from the symbolic items in the second two-dimensional subspace. The first attribute and the second attribute can be combined into a combined attribute representing the multidimensional input, and the output of the transformer neural network can be generated based on the combined attribute.
[0051] In some respects, the generation of the transformer neural network output based on the combined attribute comprises the generation of the output with the processing of the combined attribute through a direct feed component of the transformer neural network. Petition 870250083079, dated 09 / 15 / 2025, pages 130 / 154 20 / 29
[0052] In some respects, the 300 operations additionally include performing one or more actions based on the output generated by the transformer neural network. The output generated by the transformer neural network may include, for example, an identification of different portions of a multidimensional input corresponding to different objects or classes of objects and, for each object or class of object, information that identifies the importance of that object to a scene. The one or more actions may include performing actions on one or more of the objects in the scene based, at least in part, on the importance of that object to the scene.For example, varying levels of compression can be applied to different portions of the multidimensional input, with higher levels of compression (with greater compression loss) being applied to less important portions of the multidimensional input and lower levels of compression (with lower or no compression loss) being applied to more important portions of the multidimensional input. In another example, the output generated from the transformer neural network might include an identification of objects in multidimensional space and a prediction of how at least some of the identified objects will move through multidimensional space. One or more actions might include generating one or more control inputs to manage the movement of an autonomous or semi-autonomous device through multidimensional space.An autonomous or semi-autonomous device may include, for example, an autonomous vehicle, a robotic arm, or other devices that can move in a multidimensional space with limited or no human control. Obviously, it should be recognized that the outputs and actions performed based on the outputs described above are illustrative, and the outputs and actions performed based on the outputs generated by a transformer neural network may vary depending on the environment from which the multidimensional input is captured and the environment in which a device processing and / or using the multidimensional input operates. Example processing system for processing multidimensional inputs in transformer neural networks based on Petition 870250083079, dated 09 / 15 / 2025, pages 131 / 154 21 / 29 Subspatial decomposition of multidimensional inputs
[0053] Figure 4 depicts an example 400 processing system configured to perform various aspects of the present disclosure, including, for example, the techniques and methods described in relation to Figures 1 to 3. In one aspect, the 400 processing system can train, implement, or provide a machine learning model using transformer-based architectures, such as the 100 architecture of Figure 1. Although depicted as a single system for conceptual clarity, in at least some aspects, as discussed above, the operations described below in relation to the 400 processing system can be distributed across any number of devices.
[0054] The 400 processing system includes a central processing unit (CPU) 402 which, in some examples, may be a multi-core CPU. Instructions executed on CPU 402 may be loaded, for example, from a program memory associated with CPU 402, or they may be loaded from a memory partition 424.
[0055] The 400 processing system also includes additional processing components adapted for specific functions, such as a graphics processing unit (GPU) 404, a digital signal processor (DSP) 406, a neural processing unit (NPU) 408, a multimedia processing unit 410 and a wireless connectivity component 412.
[0056] An NPU, such as the NPU 408, is generally a specialized circuit configured to implement the control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), and the like. An NPU may sometimes be alternatively called a neural signal processor (NSP). Petition 870250083079, dated 09 / 15 / 2025, pages 132 / 154 22 / 29 processor), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.
[0057] NPUs, such as the NPU 408, are configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, a plurality of NPUs can be instantiated on a single chip, as a system on a chip (SoC), while in other examples, NPUs can be part of a dedicated neural network accelerator.
[0058] NPUs can be optimized for training or inference or, in some cases, configured to balance performance between the two. For NPUs capable of performing both training and inference, the two tasks can still generally be performed independently.
[0059] NPUs designed to accelerate training are generally configured to speed up the optimization of new models, which is a computationally intensive operation involving inputting data into an existing (often identified or tagged) dataset, iterating through the dataset, and then adjusting model parameters such as weights and biases to improve model performance. Generally, optimization based on a wrong prediction involves backpropagation through the model layers and gradient determination to reduce the prediction error.
[0060] NPUs designed to accelerate inference are generally configured to operate on complete models. Such NPUs can therefore be configured to take new data as input and quickly process that new data through an already trained model to generate a model output (e.g., an inference).
[0061] In an implementation, NPU 408 is part of one or more of Petition 870250083079, dated 09 / 15 / 2025, pp. 133 / 154 23 / 29 a CPU 402, a GPU 404 and / or the DSP 406.
[0062] In some examples, the wireless connectivity component 412 may include subcomponents, for example, for third-generation (3G) connectivity, fourth-generation (4G) connectivity (e.g., 4G LTE), fifth-generation connectivity (e.g., 5G or NR), WiFi connectivity, Bluetooth connectivity, and other wireless transmission standards. The wireless connectivity component 412 is additionally connected to one or more antennas 414.
[0063] The processing system 400 may also include one or more sensor processing units 416 associated with any type of sensor, one or more image signal processors (ISPs) 418 associated with any type of image sensor, and / or a navigation component 420 which may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.
[0064] The processing system 400 may also include one or more input and / or output devices 422, such as screens, touch-sensitive surfaces (including touch-sensitive displays), physical buttons, speakers, microphones and the like.
[0065] In some instances, one or more of the processors in the 400 processing system may be based on an ARM or RISC-V instruction set.
[0066] The 400 processing system also includes a 424 memory, which is representative of one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, and the like. In this example, the 424 memory includes computer executable components, which can be executed by one or more of the previously mentioned processors of the 400 processing system.
[0067] In particular, in this example, memory 424 includes a component Petition 870250083079, dated 09 / 15 / 2025, pages 134 / 154 24 / 29 multidimensional input decomposition 424A, an attention matrix generation component 424B, and an output generation component 424C. Although depicted as distinct components for conceptual clarity in Figure 4, the illustrated components (and others not depicted) may be collectively or individually implemented in various aspects.
[0068] In general, the processing system 400 and / or its components can be configured to perform the methods described in the present invention.
[0069] Notably, in other respects, aspects of the 400 processing system may be omitted, such as in the case where the 400 processing system is a server computer or similar. For example, the 410 multimedia processing unit, the 412 wireless connectivity component, the 416 sensor processing units, the 418 ISPs, and / or the 420 navigation component may be omitted in other respects. Furthermore, aspects of the 400 processing system may be distributed among multiple devices. Example clauses
[0070] Implementation details of various aspects of this disclosure are described in the following numbered clauses:
[0071] Clause 1: A processor-implemented method comprising: decomposing a multidimensional input into a plurality of two-dimensional subspaces, wherein the plurality of two-dimensional subspaces share a common dimension; generating a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network; generating a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network; and generating an output of the transformer neural network based on the first attention matrix and the second attention matrix. Petition 870250083079, dated 09 / 15 / 2025, pages 135 / 154 25 / 29
[0072] Clause 2: The method of Clause 1, in which the multidimensional input comprises an input with a plurality of spatial dimensions and a temporal dimension.
[0073] Clause 3: The method of Clause 2, wherein the plurality of spatial dimensions comprises a spatial dimension of width and a spatial dimension of height and wherein the common dimension comprises the temporal dimension, so that the computational complexity involved in generating the output of the transformer neural network is reduced in relation to the decomposition of the multidimensional input into a spatial component and a temporal component.
[0074] Clause 4: The method of Clause 2 or 3, wherein the multidimensional input comprises a video input.
[0075] Clause 5: The method of any of Clauses 2 to 4, wherein: the first two-dimensional subspace comprises a subspace based on a first spatial dimension among the plurality of spatial dimensions and on the temporal dimension and the second two-dimensional subspace comprises a subspace based on a second spatial dimension among the plurality of spatial dimensions and on the temporal dimension.
[0076] Clause 6: The method of any of Clauses 1 to 5, wherein the generation of the transformer neural network output comprises: computing a first attribute based on the first attention matrix and on projected values from the symbolic items in the first two-dimensional subspace; computing a second attribute based on the second attention matrix and on projected values from the symbolic items in the second two-dimensional subspace; combining the first attribute and the second attribute into a combined attribute representing the multidimensional input; and generating the transformer neural network output based on the combined attribute.
[0077] Clause 7: The method of Clause 6, in which the combined attribute comprises a sum of the first attribute and the second attribute.
[0078] Clause 8: The method of Clause 6 or 7, in which the generation of the transformer neural network output is based on the combined attribute. Petition 870250083079, dated 09 / 15 / 2025, pages 136 / 154 26 / 29 involves generating the output with combined attribute processing via a direct feed component of the transformer neural network.
[0079] Clause 9: The method of any of Clauses 1 to 8, wherein the generation of the first attention matrix based on the projection of symbolic items onto the first two-dimensional subspace comprises: projecting the first two-dimensional subspace onto query data, key data, and value data; and generating the first attention matrix based on the query data, key data, and various components of the common dimension onto the first two-dimensional subspace.
[0080] Clause 10: The method of any of Clauses 1 to 9, wherein the generation of the second attention matrix based on the projection of symbolic items onto the second two-dimensional subspace comprises: projecting the second two-dimensional subspace onto query data, key data, and value data; and generating the second attention matrix based on the query data, key data, and various components of the common dimension onto the second two-dimensional subspace.
[0081] Clause 11: A processing system comprising: a memory containing computer executable instructions; and one or more processors configured to execute the computer executable instructions and to make the processing system perform a method in accordance with any of clauses 1 to 10.
[0082] Clause 12: A processing system comprising means for carrying out a method in accordance with any of clauses 1 to 10.
[0083] Clause 13: A non-transient, computer-readable means comprising computer-executable instructions which, when executed by one or more processors of a processing system, cause the processing system to perform a method in accordance with any of clauses 1 to 10.
[0084] Clause 14: A computer program product embedded in a computer-readable storage medium comprising Petition 870250083079, dated 09 / 15 / 2025, pp. 137 / 154 27 / 29 code to perform a method according to any of clauses 1 to 10. Additional considerations
[0085] The preceding description is provided to enable any person skilled in the art to practice the various aspects described in the present invention. The examples discussed in the present invention are not limiting to the scope, applicability, or aspects set forth in the claims. Various modifications of these aspects will be readily apparent to those skilled in the art, and the generic principles defined in the present invention may be applied to other aspects. For example, changes may be made to the function and arrangement of the elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components, as appropriate. For example, the methods described may be performed in a different order than that described, and various steps may be added, omitted, or combined. Furthermore, attributes described in relation to some examples may be combined in some other examples.For example, an apparatus may be implemented, or a method may be practiced, using any number of the aspects set forth in the present invention. Furthermore, the scope of the disclosure is intended to cover such an apparatus or method practiced with the use of another structure, functionality, or structure and functionality in addition to or different from the various aspects of the disclosure set forth in the present invention. It should be understood that any aspect of the disclosure disclosed in this invention may be incorporated by one or more elements of a claim.
[0086] As used in the present invention, the term example means that it serves as an example, instance, or illustration. Any aspect described in the present invention as exemplary should not necessarily be interpreted as preferential or advantageous in relation to other aspects.
[0087] As used in the present invention, an expression that refers to at least one of a list of items refers to any Petition 870250083079, dated 09 / 15 / 2025, pp. 138 / 154 28 / 29 combination of these items, including unique members. As an example, at least one of: a, b, or c is intended to encompass a, b, c, ab, ac, bc, and a-bc, as well as any combination with multiples of the same element (for example, aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).
[0088] As used in the present invention, the term determine encompasses a wide variety of actions. For example, determine may include calculate, compute, process, derive, investigate, search (e.g., search in a table, a database, or other data structure), verify, and the like. Furthermore, determine may include receive (e.g., receive information), access (e.g., access data in a memory), and the like. Additionally, determine may include solve, select, choose, establish, and the like.
[0089] The methods disclosed in the present invention comprise one or more steps or actions to perform the methods. The steps and / or actions of the method can be interchanged without departing from the scope of the claims. In other words, unless a particular order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. Additionally, the various operations of the methods described above can be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including, but not limited to, a circuit, an application-specific integrated circuit (ASIC), or a processor.In general, when operations are illustrated in the figures, these operations may have equivalent components in a more functional way, corresponding to similar numbering.
[0090] The following claims are not intended to be limited to the aspects shown in the present invention, but should have the full scope consistent with the language of the claims. In a claim, reference to an element in the singular is not intended to mean one and only Petition 870250083079, dated 09 / 15 / 2025, pp. 139 / 154 29 / 29 one, unless specifically stated otherwise, but instead means one or more. Unless specifically stated otherwise, the term any refers to one or more. No element of a claim shall be interpreted in accordance with the provisions of Title 35 of the USC Code § 112(f), unless the element is expressly mentioned using the expression means to or, in the case of a method claim, the element is mentioned using the expression step to. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or may become known to those skilled in the art are expressly incorporated into the present invention by reference and are intended to be encompassed by the claims.Furthermore, nothing disclosed in this invention is intended to be exclusive to the public, regardless of whether such disclosure is explicitly mentioned in the claims. Petition 870250083079, dated 09 / 15 / 2025, pp. 140 / 154
Claims
1 / 9 CLAIMS 1. A processor-implemented method characterized by comprising: decomposing a multidimensional input into a plurality of two-dimensional subspaces, wherein the plurality of two-dimensional subspaces share a common dimension; generating a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network; generating a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network; and generating an output of the transformer neural network based on the first attention matrix and the second attention matrix.
2. Method, according to claim 1, characterized in that the multidimensional input comprises an input with a plurality of spatial dimensions and a temporal dimension.
3. A method according to claim 2, characterized in that the plurality of spatial dimensions comprises a spatial dimension of width and a spatial dimension of height, and the common dimension comprises the temporal dimension, such that the computational complexity involved in generating the output of the transformer neural network is reduced compared to decomposing the multidimensional input into a spatial component and a temporal component.
4. Method according to claim 2, characterized in that the multidimensional input comprises a video input.
5. Method, according to claim 2, characterized by: the first two-dimensional subspace comprising a subspace based on a first spatial dimension among the plurality of spatial dimensions and on the temporal dimension and Petition 870250083079, dated 09 / 15 / 2025, page 141 / 154 2 / 9 the second two-dimensional subspace comprising a subspace based on a second spatial dimension among the plurality of spatial dimensions and on the temporal dimension.
6. Method, according to claim 1, characterized in that the generation of the transformer neural network output comprises: computing a first attribute based on the first attention matrix and on projected values from the symbolic items in the first two-dimensional subspace; computing a second attribute based on the second attention matrix and on projected values from the symbolic items in the second two-dimensional subspace; combining the first attribute and the second attribute into a combined attribute representing the multidimensional input; and generating the transformer neural network output based on the combined attribute.
7. Method, according to claim 6, characterized in that the combined attribute comprises a sum of the first attribute and the second attribute.
8. Method, according to claim 6, characterized in that the generation of the transformer neural network output based on the combined attribute comprises the processing of the combined attribute through a direct feed component of the transformer neural network.
9. Method, according to claim 1, characterized by the generation of the first attention matrix based on the projection of symbolic items in the first two-dimensional subspace comprising: projecting the first two-dimensional subspace onto query data, key data and value data; and generating the first attention matrix based on the query data, the key data and several components of the common dimension in the first two-dimensional subspace.
10. Method, according to claim 1, characterized by the generation of the second attention matrix based on the projection of symbolic items in the second two-dimensional subspace comprising: projecting the second two-dimensional subspace onto query data, key data and value data; and generating the second attention matrix based on the query data, the key data and several components of the common dimension in the second two-dimensional subspace.
11. A system characterized by comprising: a memory that has executable instructions stored therein; and a processor configured to execute the executable instructions to make the system: decompose a multidimensional input into a plurality of two-dimensional subspaces, where the plurality of two-dimensional subspaces share a common dimension; generate a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network; generate a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network; and generate an output of the transformer neural network based on the first attention matrix and the second attention matrix.
12. System, according to claim 11, characterized in that the multidimensional input comprises an input with a plurality of spatial dimensions and a temporal dimension.
13. System, according to claim 12, characterized by the plurality of spatial dimensions comprising a spatial dimension of width and a spatial dimension of height and by the common dimension comprising the temporal dimension, so that the computational complexity involved in generating the output of the transformer neural network is reduced in relation to the decomposition of the multidimensional input into a spatial component and a temporal component.
14. System according to claim 12, characterized in that the multidimensional input comprises a video input.
15. System, according to claim 12, characterized by: the first two-dimensional subspace comprising a subspace based on a first spatial dimension among the plurality of spatial dimensions and on the temporal dimension, and the second two-dimensional subspace comprising a subspace based on a second spatial dimension among the plurality of spatial dimensions and on the temporal dimension.
16. System according to claim 11, characterized in that, in order to generate the output of the transformer neural network, the processor is configured to make the system: compute a first attribute based on the first attention matrix and on projected values from the symbolic items in the first two-dimensional subspace; compute a second attribute based on the second attention matrix and on projected values from the symbolic items in the second two-dimensional subspace; combine the first attribute and the second attribute into a combined attribute representing the multidimensional input; and generate the output of the transformer neural network based on the combined attribute.
17. System according to claim 16, characterized in that, in order to generate the output of the transformer neural network based on the combined attribute, the processor is configured to make the system process the combined attribute through a direct power supply component of the transformer neural network.
18. System, according to claim 11, characterized in that, in order to generate the first attention matrix based on the projection of symbolic items in the first two-dimensional subspace, the processor is configured to make the system: project the first two-dimensional subspace onto query data, key data, and value data; and generate the first attention matrix based on the query data, the key data, and several components of the common dimension in the first two-dimensional subspace.
19. System, according to claim 11, characterized in that, in order to generate the second attention matrix based on the projection of symbolic items in the second two-dimensional subspace, the processor is configured to make the system: project the second two-dimensional subspace onto query data, key data and value data; and generate the second attention matrix based on the query data, the key data and several components of the common dimension in the second two-dimensional subspace.
20. A system characterized by comprising: means to decompose a multidimensional input into a plurality of two-dimensional subspaces, wherein the plurality of two-dimensional subspaces share a common dimension; means to generate a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network; means to generate a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network; and means to generate an output of the transformer neural network based on the first attention matrix and the second attention matrix.
21. System, according to claim 20, characterized by the multidimensional input comprising an input with a plurality of spatial dimensions and a temporal dimension.
22. System, according to claim 21, characterized in that the plurality of spatial dimensions comprises a spatial dimension of width and a spatial dimension of height and in that the common dimension comprises the temporal dimension, such that the computational complexity involved in generating the output of the transformer neural network is reduced compared to decomposing the multidimensional input into a spatial component and a temporal component.
23. System, according to claim 21, characterized by: the first two-dimensional subspace comprising a subspace based on a first spatial dimension among a plurality of spatial dimensions and on the temporal dimension, and the second two-dimensional subspace comprising a subspace based on a second spatial dimension among a plurality of spatial dimensions and on the temporal dimension.
24. System, according to claim 20, characterized in that the means for generating the output of the transformer neural network comprises: means for computing a first attribute based on the first attention matrix and on projected values from the symbolic items in the first two-dimensional subspace; means for computing a second attribute based on the second attention matrix and on projected values from the symbolic items in the second two-dimensional subspace; means for combining the first attribute and the second attribute into a combined attribute representing the multidimensional input; and means for generating the output of the transformer neural network based on the combined attribute.
25. System, according to claim 24, characterized by the means for generating the output of the transformer neural network based on the combined attribute comprising the means for processing the combined attribute Petition 870250083079, dated 09 / 15 / 2025, pp. 146 / 154 7 / 9 through a direct power supply component of the transformer neural network.
26. System, according to claim 20, characterized by the means for generating the first attention matrix based on the projection of symbolic items in the first two-dimensional subspace comprising: means for projecting the first two-dimensional subspace onto query data, key data and value data; and means for generating the first attention matrix based on the query data, the key data and various components of the common dimension in the first two-dimensional subspace.
27. System, according to claim 20, characterized in that the means for generating the second attention matrix based on the projection of symbolic items in the second two-dimensional subspace comprises: means for projecting the second two-dimensional subspace onto query data, key data and value data; and means for generating the second attention matrix based on the query data, the key data and various components of the common dimension in the second two-dimensional subspace.
28. A computer-readable medium characterized by having executable instructions stored within it that, when executed by a processor, performs an operation comprising: decomposing a multidimensional input into a plurality of two-dimensional subspaces, wherein the plurality of two-dimensional subspaces share a common dimension; generating a first attention matrix based on a projection of symbolic items into a first two-dimensional subspace from among the plurality of two-dimensional subspaces through an attention block of a transformer neural network; generating a second attention matrix based on a projection of symbolic items into a second two-dimensional subspace from among the plurality of two-dimensional subspaces through the attention block of the transformer neural network; and generating an output of the transformer neural network based on the first attention matrix and the second attention matrix.
29. Computer-readable medium according to claim 28, characterized in that the multidimensional input comprises an input with a plurality of spatial dimensions and a temporal dimension.
30. A computer-readable medium according to claim 29, characterized in that the plurality of spatial dimensions comprises a spatial dimension of width and a spatial dimension of height and in that the common dimension comprises the temporal dimension, such that the computational complexity involved in generating the output of the transformer neural network is reduced compared to decomposing the multidimensional input into a spatial component and a temporal component.
31. Computer-readable medium according to claim 29, characterized in that: the first two-dimensional subspace comprises a subspace based on a first spatial dimension among a plurality of spatial dimensions and on the temporal dimension, and the second two-dimensional subspace comprises a subspace based on a second spatial dimension among a plurality of spatial dimensions and on the temporal dimension.
32. Computer-readable means according to claim 28, characterized in that the generation of the transformer neural network output comprises: computing a first attribute based on the first attention matrix and on projected values from the symbolic items in the first two-dimensional subspace; computing a second attribute based on the second attention matrix and on projected values from the symbolic items in the second two-dimensional subspace; combining the first attribute and the second attribute into a combined attribute representing the multidimensional input; and generating the transformer neural network output based on the combined attribute.
33. Computer-readable medium according to claim 32, characterized in that the generation of the transformer neural network output based on the combined attribute comprises the processing of the combined attribute through a direct feed component of the transformer neural network.
34. Computer-readable means according to claim 28, characterized in that the generation of the first attention matrix based on the projection of symbolic items in the first two-dimensional subspace comprises: projecting the first two-dimensional subspace onto query data, key data and value data; and generating the first attention matrix based on the query data, the key data and several common dimension components in the first two-dimensional subspace.
35. Computer-readable medium according to claim 28, characterized in that the generation of the second attention matrix based on the projection of symbolic items onto the second two-dimensional subspace comprises: projecting the second two-dimensional subspace onto query data, key data, and value data; and generating the second attention matrix based on the query data, key data, and various components of the common dimension in the second two-dimensional subspace. Petition 870250083079, dated 09 / 15 / 2025, pp. 149 / 154