Table folding and acceleration
Tabular convolution addresses inefficiencies in sparse data processing by directly operating on sparse data in a tabular format, enhancing efficiency and reducing complexity in machine learning operations.
Patent Information
- Application Number
- JP2023509673
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-19
- Filing Date
- 2021-08-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-08-20
AI Technical Summary
Conventional machine learning processing is inefficient for sparse data due to unnecessary operations on zero-valued inputs, leading to inefficiencies in power consumption, processing complexity, and storage requirements, particularly in edge devices.
Implementing tabular convolution, which converts sparse data into a tabular format for direct mathematical operations without converting to dense representations, using indexing rules and skip connections to maintain sparsity and reduce complexity.
Tabular convolution achieves efficient and fast processing of sparse data, reducing computational complexity and power consumption while maintaining accuracy, benefiting edge devices and mobile applications.
Smart Images

Figure 0007710507000045 
Figure 0007710507000046 
Figure 0007710507000047
Abstract
Description
Claim of Priority
[0001] Cross - Reference to Related Applications
[0001] This application claims the benefit and priority of U.S. Patent Application No. 17 / 407,046, filed on August 19, 2021, which claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 068,891, filed on August 21, 2020, the entire contents of which are incorporated herein by reference.
Technical Field
[0002]
[0002] Aspects of the present disclosure relate to improved machine - learning processing.
Background Art
[0003] Introduction
[0003] Machine learning is generally a process of creating a trained model (e.g., an artificial neural network, a tree, or other structure) that represents a generalization fit to a set of a priori known training data. Applying the trained model to new data creates an inference, which can be used to gain insights into the new data. In some cases, applying the model to new data is described as "running an inference" on the new data.
[0004]
[0004] As the use of machine learning has increased rapidly to enable various machine - learning (or artificial - intelligence) tasks, there has been a need for more efficient processing of machine - learning model data. For example, "edge - processing" devices, such as mobile devices, always - on devices, Internet of Things (IoT) devices, etc., must balance processing power with power and packaging constraints.
[0005]
[0005] One aspect of machine learning processing that has been inefficient in the past is the processing of sparse data in a machine learning model. Machine learning generally requires a large number of mathematical operations on input data, particularly including multiplication and addition operations. Conventionally, these operations are performed regardless of whether they produce a meaningful output. For example, when the input data is extremely sparse (e.g., has many zero values), conventional machine learning processing still performs multiplication and addition on the zero-valued inputs. This can result in inefficiencies because the operations on zero input values may not change or may not produce a meaningful output in some cases. For example, adding zero to any number results in the same number.
[0006]
[0006] Accordingly, there is a need for systems and methods for improving the efficiency of machine learning processing of sparse input data.
SUMMARY OF THE INVENTION
[0007]
[0007] Some aspects provide a method for performing tabular convolution that includes performing a tabularization operation on input data to generate a tabularized representation of the input data and performing a convolution operation using the tabularized representation of the input data to generate a convolution output.
[0008]
[0008] Other aspects provide a processing system configured to perform the methods described above as well as the methods described herein, a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the methods described above as well as the methods described herein, a computer program product implemented on a computer-readable storage medium comprising code for performing the methods described above as well as the methods described herein, and a processing system comprising means for performing the methods described above as well as the methods described herein.
[0009]
[0009] The following description and the related drawings detail some exemplary features of one or more embodiments.
[0010]
[0010] The accompanying drawings illustrate some aspects of one or more embodiments and should not be regarded as limiting the scope of the present disclosure.
Brief Description of the Drawings
[0011]
Figure 1A
[0011] A diagram showing an example of a weight kernel that can be used in tabular convolution.
Figure 1B
Figure 2A
[0012] A diagram showing an exemplary manner of tabulating sparse input data.
Figure 2B
Figure 3
[0013] A diagram showing an exemplary method for performing tabular convolution.
Figure 4A
[0014] A diagram showing exemplary results of tabular convolution.
Figure 4B
Figure 4C
Figure 5
[0015] A diagram showing an exemplary method for performing tabular convolution using a skip connection.
Figure 6
[0016] A diagram showing an example of hashing tabulated input data.
Figure 7
[0017] A diagram showing an exemplary method for performing tabular convolution.
Figure 8
[0018] A diagram showing an exemplary processing system configured to perform tabular convolution.
Modes for Carrying Out the Invention
[0012]
[0019] For ease of understanding, where possible, the same reference numbers are used to designate the same elements that are common to the drawings. It is contemplated that the elements and features of one embodiment may be beneficially incorporated into other embodiments without further recitation.
[0013]
[0020] Aspects of the present disclosure provide an apparatus, a method, a processing system, and a computer-readable medium for efficiently performing machine learning processing of sparse input data and, in particular, for performing tabular convolution on sparse data to improve convolution processing efficiency.
[0014]
[0021] Sparse data generally refers to data such as vectors, matrices, or tensor data that have a relatively high proportion of entries with a value of 0 (or close to 0). For example, a simple exemplary vector with an entry [0, 0, 0, 1] can be considered 75% sparse (and 25% dense). When performing a mathematical operation on this exemplary vector, such as element-wise addition, it is clear that the sparse entries do not change the output value for 75% of the vector elements. Similarly, when performing a multiplication operation, it is clear that all consecutive multiplication operations based on 0-valued elements also result in a 0 value. Considering the large number of mathematical operations in machine learning processing, these sparse entries represent an opportunity for efficiency improvement, which can also beneficially result in faster machine learning processing, lower power consumption processing, reduced complexity processing, reduced data storage (e.g., memory) requirements, lower cost hardware requirements, and other benefits.
[0015]
[0022] Sparse representations of data are increasing in relevance in many technical fields, in contrast to conventional dense representations. For example, modern sensor systems used in various types of end-user products generally can produce sparse data. One example is the three-dimensional (3D) point clouds generated by light detection and ranging (LiDAR) sensors used in automotive applications (e.g., autonomous driving systems) and in mobile device applications (e.g., for biometric user authentication). Such 3D point clouds are generally extremely sparse.
[0016]
[0023] Conventional machine learning techniques, such as deep learning, have not utilized sparsity for performance improvement. For example, conventional convolutional processing for sparse image data may rely on multi-layer perceptron (MLP) operations (e.g., PointNet, PointNet++, and PointConv), which sacrifice efficiency due to the fully-connected nature of those operations.
[0017]
[0024] Other approaches may involve conversion of sparse input data to other representations, such as voxelization for 3D convolutional neural network processing, projection / rendering for two-dimensional (2D) convolutional neural network processing, and feature extraction for fully-connected operations. In particular, none of these transformation-based approaches directly perform deep learning on sparse data representations, and thus each of these approaches incurs additional complexity and inefficiency based on the required transformations.
[0018]
[0025] In some cases, such as when processing event images from event cameras or sensors, deep learning techniques are completely avoided due to the lack of an efficient way to perform convolution on sparse input data.
[0019]
[0026] The embodiments described herein overcome these prior art problems by implementing a sparse-aware convolution, which may alternatively be referred to as tabulated convolution. Tabulated convolution advantageously enables direct, fast, and efficient convolution processing on sparse data without the need for costly and complex transformation of the input data into a dense representation. Generally, the tabulated convolution methods described herein may be applied to 2D convolution as well as higher-dimensional convolutions. Additional embodiments described herein further improve tabulated convolution by various acceleration steps. Tabulated convolution
[0027] The embodiments of tabulated convolution described herein may generally include an operation arithmetic performed on input data, followed by a special convolution that utilizes the tabulated input data. The operation arithmetic generally takes a sparse data sequence and converts it into a table data format having one column of active points and a plurality of columns indicating the relationship of neighboring points to those active points. The special convolution is equivalent to a conventional convolution but takes the tabulated input data to perform arithmetic operations with significantly reduced complexity. Thus, tabulated convolution enables performing operations directly on a sparse data representation (e.g., without an intervening conversion to a dense data representation) while performing mathematical operations in a dense format (e.g., without processing sparse data elements). As noted above, sparse data elements in a data representation may generally be zero-valued elements. In some cases, a sparse element threshold may be used to determine whether an element should be treated as a sparse element.
[0020]
[0028] As an example, a sparse input data representation (e.g., an image) has a shape
[0021]
Number
[0022] of
[0023]
Number
[0024] can be represented as, where C is the number of channels in the input data, A represents the number of "active" (non-zero) elements in the input data, and C×A is the real space
[0025]
Number
[0026] represents the dimensionality of the
[0027]
[0029] Therefore,
[0028]
Number
[0029] represents a set of active points (or elements) in an arbitrary "raw" format that depends, for example, on the sensor and the device interface.
[0030]
Number
[0031] is then
[0032]
Number
[0033] such that the shape
[0034]
Number
[0035] can be converted to a clear dictionary representation (X active ) , where
[0036] [Number]
[0037] indicates the coordinates in channel c and in dimension j for element a,
[0038] [Number]
[0039] indicates the element value for element a. Since the raw input data has an arbitrary format, there is not always a D dimension. Here, "standardization" of the data format can be performed by introducing a D factor to explicitly represent the D-component coordinates of each active point. Thus, the X active format is a "dictionary" format of active elements where each element is represented by a tuple of a D-component coordinate and the quantity at this coordinate (location).
[0040]
[0030] Next, assume a set of depthwise separable weights having K entries in each channel c (including the central entry for channel c)
[0041] [Number]
[0042] where K represents the kernel size per channel. In the examples described herein, for simplicity, K ∈ {5, 9}, but other embodiments may use different K values.
[0043] [Number]
[0044]
[0045]
[0031] Next, the creation operation can be defined as
[0046]
Number
[0047] and can be defined as
[0032] This is to convert the expression X active into the representation X T ∈
[0048]
Number
[0049] and for each channel c, X T is
[0050]
Number
[0051] , where c ∈ {0, 1,..., C - 1}, and for each c,
[0052]
Number
[0053] is
[0033] Here,
[0054]
Number
[0055] is
[0056]
[0034] In the above formula, a indicates the active point, and the creation value for the active point in the k = 0 column is (before creation) x activeNote that it is equal to the active point in the format.
[0057]
[0035] Based on the above operation calculation definition, the indexing rule k = Λ can be defined, where Λ is a set of indices for the neighborhood of the active point. For example, when K = 5, since k = 0 is the index for the active point itself, Λ = {1, 2, 3, 4}.
[0058]
[0036] For example, FIGS. 1A and 1B show two exemplary kernels 100A and 100B with weights (W c ) for channel c, where K = 9 for kernel 100A and K = 5 for kernel 100B. In this example, kernel 100A in FIG. 1A is a "square kernel" with dimension 3×3 and thus has entries for K = 9, and kernel 100B in FIG. 1B is a non-square, cross-shaped kernel with entries for K = 5.
[0059]
[0037] In these examples, each entry of the kernel weight matrices 100A and 100B (e.g.,
[0060]
Number
[0061] ) is referenced continuously, and the intermediate weight is always k = 0.
[0062]
[0038] Therefore, considering FIG. 1A with D = 2 (dimensions) and K = 9 (kernel entries), the indexing rule can be derived for each k = Λ ∈ {1,..., K - 1}, and thus,
[0063]
Number
[0064] is.
[0065]
[0039] As another example, considering FIG. 1B where D = 2 (dimensions) and K = 5 (kernel entries), the indexing rules can be derived for each k = Λ ∈ {1,..., K−1}, and thus,
[0066]
Number
[0067] it is.
[0068]
[0040] In these examples, when the index values are in the form (x, y), a negative x-index value represents a movement to the left with respect to the central k = 0 index, a negative y-index value represents a movement downward with respect to the center, a positive x-index value represents a movement to the right with respect to the center, and a positive y-index value represents a movement upward with respect to the center.
[0069]
[0041] Assuming the sparse input data representation definition and indexing rules described above, FIGS. 2A and 2B show examples of how weight kernels are used to form tabular output based on sparse data input.
[0070]
[0042] Specifically, FIG. 2A shows the active elements {x0,..., x9} plotted in D = 2 dimensions (d 0 and d 1 ). Thus, in this example, A = 10 active elements.
[0071]
[0043] Furthermore, in FIG. 2A, a weight kernel overlay 202 is shown to provide context to the indexing scheme in FIG. 2B. For example, looking from the active element x0, the active element x1 is indexed as (0, -1) since it is 1 unit down (-1) from x0 in the d 1 dimension and in the same unit in the d 0 dimension.
[0072]
[0044] Tabulation output for a single channel c of sparse data based on a weight kernel with K = 5 (as in the case of FIG. 1B)
[0073]
Number
[0074] is shown in FIG. 2B.
[0075]
[0045] In FIG. 2B, the leftmost column in the table indexed as (0,0) (e.g., the center point) represents the active point vector X of the active elements {x0,...,x9} shown in FIG. 2A active which is. The remaining columns indexed as (1,0), (-1,0), (0,1), and (0,-1) represent the tabulation relationships between the active elements based on the indexing rules described above.
[0076]
[0046] In particular, the tabulation operation T that generates the output shown in FIG. 2B K involves no multiplication and requires only bitwise logic to rearrange the sparse data (which is a significantly smaller amount than the dense data).
[0077]
[0047] Previous weight kernel definition
[0078]
Number
[0079] Assuming, for each c in C, a weight column vector (tensor) in a neural network that is trained and can be applied in testing to derive the tabulation convolution output according to the tabulation operation defined as follows
[0080]
Number
[0081] exists.
[0082]
Number
[0083] ,
[0048] Here,
[0084]
Number
[0085] is a vector for channel c in C. Or, for all channels, the operation can be
[0086]
Number
[0087] defined by.
[0088]
[0049] Figure 3 shows an exemplary tabular convolution operation 300 that starts with the tabulation of the input data X T in step 302 (described above) to generate the tabular output data X active . Then, the tabular output data X T is convolved with the kernel weight W in step 304 to generate the tabular convolution output data Y = X T W.
[0089]
[0050] The example shown in Figure 3 assumes a batch convolution of all one or more channels c in C, but the process is repeatedly iterated separately for the data of each channel, i.e., for all channels c in C
[0090]
Number
[0091] It should be noted that such can generate. In such a case, the partial (channel-specific) convolution output can be accumulated during each channel-specific iteration to generate the final tabular convolution output.
[0092]
[0051] Tabular convolution achieves a mathematically equivalent convolution result compared to conventional dense convolution, but the preliminary representation enables efficient matrix multiplication operations. Therefore, the complexity required to perform tabular convolution is O(CA) when CAK, or K is fixed, where C is the number of channels in the input data, A represents the number of active (non-zero) elements in the input data, and K is the number of kernel entries per layer. If N output channels are involved in the convolutional layer, the complexity is thus CAKN or O(CAN).
[0093]
[0052] Furthermore, according to the above definition, the sparsity of the sparse representation is maintained after tabular convolution of one or more layers. In other words, the set of active points remains unaffected, which means that there is no expansion or blurring of the convolution output as it strictly follows the destination set of points like the active points.
[0094]
[0053] For example, FIGS. 4A to 4C show examples of conventional convolution versus tabular convolution of sparse input data.
[0095]
[0054] In particular, FIG. 4A shows sparse input data 400 where pixels (squares) represent active elements with non-zero values.
[0096]
[0055] FIG. 4B shows the result 410 of conventional convolution of the sparse input data 400. In particular, the convolution created an expansion of the original active elements such that there are many more active elements with non-zero values, as indicated by the different shaded pixels around the original input data pixels.
[0097]
[0056] On the other hand, FIG. 4C shows the result 420 of the surface convolution operation. In particular, the value of the active element has changed as indicated by the change in the pattern at each pixel compared to FIG. 4A, but the number and location of the pixels having non-zero values (active elements) remain the same. Surface convolution using skip connection
[0057] In some embodiments, the surface convolution can be implemented using skip connections, as shown in FIG. 5. In such embodiments, X T is
[0098]
Number
[0099] can be redefined as follows.
[0100]
Number
[0101]
[0058] In particular, compared to k from 0 to K-1 in the surface convolution definition without using skip connections as described above, the above redefinition includes k from 1 to K-1. Therefore, this redefinition is before applying the kernel weights
[0102]
Number
[0103] to form, column 0 of X T (the dense vector component having the active element (X active ) as shown in FIG. 2B) is deleted from the tabulation result. Advantageously, the deletion of the active element is for the remaining tabulation representation
[0104]
Number
[0105] Enhance the hydrophobicity, which further improves the efficiency gain of the tabulation convolution.
[0106]
[0059] To account for the deletion of the active element (X active ) before convolution, a skip connection 502 is added. Further, the weights corresponding to the central entries of the kernels of all channels (e.g., k = 0 in FIGS. 1A and 1B) are set to 1 for all active points, and the non-central entries of the kernels of all channels remain trainable and are relative to the (1) central entry, thereby forming W’. Thus, the tabulation convolution obtained using the skip connection is
[0107]
Number
[0108] which can be represented as.
[0109]
[0060] Thus, in example 500 of FIG. 5, X active is provided as input to the tabulation step 504 and also provided to the element-wise summation (or addition) operator 508 on the skip connection 502.
[0110]
[0061] The tabulation data representation
[0111]
Number
[0112] then is
[0113]
Number
[0114] To generate, in step 506, it is convolved with the kernel weight W’. Finally, the convolution output
[0115]
Number
[0116] is added to the skip connection 502 (carrying the value X active ) to generate the tabular convolution output Y.
[0117]
[0062] Addition of Skip Connection and Redefinition of Tabular Representation
[0118]
Number
[0119] provides additional benefits over tabular convolutions without skip connections (such as, for example, those described with respect to FIG. 3), including further reduction in computational complexity (especially for non-square weight kernels such as kernel 100B in FIG. 1B and kernel 202 in FIG. 2A) and further reduction in weight memory and intermediate activation memory (especially for sparse weight kernels). Each of these processing efficiency improvements can save time and power in the processing system, which can be particularly beneficial for mobile and edge processing devices. Furthermore, using all-sparse data in the intermediate memory allows for further optimization, as described below. Hashed Tabular Convolution
[0063] It is possible to further compress the sparse data representation by creating a one-dimensional hash for each dimension of the weight kernel.
[0120]
[0064] As shown in FIG. 6,
[0121]
Number
[0122] Without populating each full column with zero entries, a one-dimensional hashed pointer list (e.g., 602 and 604) can be generated for each dimension (e.g., d 0 and d 1 ) in the data. Note here that a "hashed pointer list" describes a generalized data structure and is used to represent one-dimensional or multi-dimensional structures.
[0123]
[0065] Using the examples from FIGS. 1B, 2A, and 2B, when K = 5 and D = 2, the hashed pointer list {L o , L1} can be defined by each list that points to the values of the dimension
[0124]
Number
[0125] when D = 2. In this example, both hashed pointer lists {L o , L1} here are first constructed with row-wise pointers for L o and column-wise pointers for L1. After column-wise folding and multiplication by the corresponding kernel values for each column, row-wise folding is performed on L o to remove all empty entries. In the general case, a hashed pointer list can be generated for {L o , L1,..., L D-1}, which means that hashed table folding can be applied to higher-dimensional data compared to 2D data as in the example of FIG. 6.
[0126]
[0066] To process the table data representation, then, L1 is for the kernel index k
[0127]
Number
[0128] is multiplied by the corresponding scaler entry. Thus, for example,
[0129]
Number
[0130] is.
[0131]
[0067] Since {L o , L1} is a pointer list, the dereferenced value of another pointer list L o is automatically updated with the corresponding kernel weight. For example,
[0132]
Number
[0133] is. Finally, all the pointed entries of the hashed pointer list in L o are summed up, for example,
[0134]
Number
[0135] in accordance with,
[0068] where i is the index of each element of L 0,k .
[0136]
[0069] Thus, using the hashed pointer list further reduces the computational complexity, which beneficially reduces the processing time and the memory requirements. Exemplary Test Results and Efficiency Improvements
[0070] In the tests, the surface convolution implemented with a known convolutional architecture such as ResNet beneficially reduced both the computational complexity and the model size by at least 20% without degrading the accuracy. In the tests, all kernel weights were learned by stochastic gradient descent with respect to the kernel weight element of k = 0 with a value of 1.
[0137]
[0071] Furthermore, the tests showed that non-square kernels (such as kernel 100B in FIG. 1B) yielded further improvements in complexity and memory utilization compared to square kernels of the same extent (e.g., 3 pixels wide and high such as kernel 100A in FIG. 1A). Thus, for example, a surface convolution using a non-square kernel (as in the case of FIG. 1B) was able to outperform a surface convolution using a square kernel (as in the case of FIG. 1A) having the same outer extent (e.g., 3 units wide in the widest row and 3 units high in the tallest column).
[0138]
[0072] Table 1 below includes further computational complexity comparisons based on the methods described herein.
[0139]
Table 1
[0140]
[0073] In Table 1, C refers to the number of channels, S = HW, where H and W are the height and width of the input data (e.g., tensor), respectively. Further, in Table 1, A = S / 100 is used as an example where a density of 1 / 10 is assumed in each of the two dimensions, such as for sparse data from various types of sensors including event cameras. Further, R is defined as the average (rate) number of neighbors covered by the kernel with respect to the central "active" element. Further, when hashed surface convolution with skip connections is used, the kernel size K does not directly affect the number of multiplications. Instead, R becomes a determining factor for the number of multiplications (or activation memory size).
[0141]
[0074] As shown in Table 1, since the number of multiplications and the model size are S = HW >> A, advantageously, they are reduced by tabular convolution compared to conventional depthwise separable convolutions, where H is the height of the input data (e.g., in pixel units) and W is the width of the input data (e.g., in pixel units). Exemplary Method for Performing Tabular Convolution
[0075] FIG. 7 shows an exemplary method 700 for performing tabular convolution.
[0142]
[0076] Method 700 begins at step 702, which involves performing a tabulation operation on the input data to generate a tabular representation of the input data (as in the case of step 302 in FIG. 3 and step 504 in FIG. 5).
[0143]
[0077] Method 700 then proceeds to step 704, which involves performing a convolution operation using the tabular representation of the input data to generate a convolution output (as in the case of step 304 in FIG. 3 and step 506 in FIG. 5).
[0144]
[0078] In some embodiments, method 700 further includes determining that the input data for the convolutional layer of the machine learning model has a sparsity greater than a threshold sparsity value.
[0145]
[0079] In some embodiments of method 700, performing the convolution operation comprises performing a matrix multiplication between the weight tensor and the tabular representation of the input data to generate a convolution output, as shown in FIG. 3 (X T W) and FIG. 5(
[0146]
Number
[0147] )
[0148]
[0080] In some embodiments of method 700, performing the tabulation operation comprises using indexing rules to populate a tabular representation of the input data, where the indexing rules define a plurality of index values based on relationships between active elements of the input data and a plurality of elements of the input data adjacent to the active elements.
[0149]
[0081] In some embodiments of method 700, the sparsity of the convolutional output for a convolutional layer is the same as the sparsity of the input data to the convolutional layer, as shown in the examples of FIGS. 4A and 4C.
[0150]
[0082] In some embodiments, method 700 further comprises removing active point vector components from the tabular representation of the input data and adding the active point vector components to the convolutional output to generate a convolutional layer output, prior to performing the convolutional operation. For example, as described above with respect to FIG. 5, X active can be removed from the tabular representation to form
[0151]
Number
[0152] and X active can be added back by skip connection 502.
[0153]
[0083] In some embodiments, method 700 further includes, prior to performing the convolution operation, removing active point vector components from the tabular representation of the input data and generating a plurality of one-dimensional hashed pointer lists based on the tabular representation as described above with respect to FIG. 6. In such a case, the convolution operation includes multiplying each input value associated with a pointer in a first hashed pointer list of the plurality of one-dimensional hashed pointer lists by an associated scalar weight value based on the kernel index of that value, and summing the weighted input values associated with each pointer in a second hashed pointer list of the plurality of one-dimensional hashed pointer lists to generate a convolution output.
[0154]
[0084] In some embodiments, method 700 further includes determining a loss value associated with the convolution output and updating a plurality of weights associated with the convolutional layer of the machine learning model based on the loss value.
[0155]
[0085] In some embodiments of method 700, the convolution operation comprises a depthwise separable convolution operation.
[0156]
[0086] In some embodiments of method 700, the input data comprises sparse image sensor data. In some embodiments, the input data comprises a point cloud, such as sparse light detection and ranging (LiDAR) sensor data. Exemplary processing system for performing hardware-based voice activity detection
[0087] FIG. 8 shows an exemplary processing system 800 configured to perform tabular convolutions, such as those described herein with respect to FIGS. 2-6.
[0157]
[0088] In some examples, the processing system 800 can include a central processing unit (CPU) 802, which can be a multi-core CPU. Instructions executed in the CPU 802 can be loaded, for example, from a program memory associated with the CPU 802 or from a memory partition 824.
[0158]
[0089] The processing system 800 also includes additional processing components adapted for specific functions, such as a graphics processing unit (GPU) 804, a digital signal processor (DSP) 806, a neural processing unit (NPU) 808, a multimedia processing unit 810, a multimedia processing unit 810, and a wireless connectivity component 812.
[0159]
[0090] An NPU, such as 808, is generally a special circuit configured to implement all the necessary control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. The NPU is sometimes alternatively called a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.
[0160]
[0091] An NPU, such as 808, can be configured to accelerate the performance of common machine learning tasks, such as image classification, machine translation, object detection, and various other prediction models. In some examples, multiple NPUs can be instantiated on a single chip, such as a system-on-chip (SoC), while in other examples, multiple NPUs can be part of a dedicated neural network accelerator.
[0161]
[0092] The NPU can be optimized for training or inference, or in some cases, configured to balance performance between the two. In the case of an NPU capable of performing both training and inference, the two tasks can generally still be performed independently.
[0162]
[0093] An NPU designed to accelerate training generally involves inputting an existing (often labeled or tagged) dataset to improve model performance, iterating over that dataset, and then adjusting model parameters such as weights and biases, which is a very compute-intensive operation that involves training a new model to accelerate. Generally, optimizing based on incorrect predictions involves backpropagating through the layers of the model to determine gradients for reducing prediction error.
[0163]
[0094] An NPU designed to accelerate inference is generally configured to operate on a completed model. Thus, such an NPU can be configured to input new data and quickly process it through an already trained model to generate a model output (e.g., an inference).
[0164]
[0095] In one implementation, the NPU808 is part of one or more of the CPU802, GPU804, and / or DSP806.
[0165]
[0096] In some examples, the wireless connectivity component 812 can include sub-components for, for example, third-generation (3G) connectivity, fourth-generation (4G) connectivity (e.g., 4G LTE (registered trademark)), fifth-generation connectivity (e.g., 5G or NR), Wi-Fi (registered trademark) connectivity, Bluetooth (registered trademark) connectivity, and other wireless data transmission standards. The wireless connectivity processing component 812 is further connected to one or more antennas 814.
[0166]
[0097] The processing system 800 may also include one or more sensor processing units 816 associated with any mode of sensors, one or more image signal processors (ISPs) 818 associated with any mode of image sensors, and / or a navigation processor 820, and the navigation processor 820 may include satellite-based positioning system components (e.g., GPS or GLONASS) as well as inertial positioning system components.
[0167]
[0098] The processing system 800 may also include one or more input and / or output devices 822 such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, etc.
[0168]
[0099] In some examples, one or more of the processors of the processing system 800 may be based on the ARM or RISC-V instruction set.
[0169]
[0100] The processing system 800 also includes a memory 824 representing one or more static memories and / or dynamic memories such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 824 includes computer-executable components that may be executed by one or more of the above-described processors of the processing system 800.
[0170]
[0101] In particular, in this example, the memory 824 includes a determination component 824A, a tabulation component 824B, a convolution component 824C, a vector-matrix multiplication component 824D, a hashing component 824E, an indexing component 824F, a training component 824G, and model parameters 824G. The shown components, as well as other non-shown components, may be configured to implement various aspects of the methods described herein.
[0171] Generally, processing system 800 and / or its components may be configured to implement the methods described herein. In particular, if the processing system is primarily configured to train a machine learning model using tabular convolution, the processing system may omit some aspects such as multimedia 810, wireless connectivity 812, antenna 814, sensors 816, ISP 818, and navigation 820. Exemplary clauses
[0103] Implementation examples are described in the following numbered clauses.
[0172]
[0104] Clause 1: A method comprising performing a tabulation operation on input data to generate a tabular representation of the input data, and performing a convolution operation using the tabular representation of the input data to generate a convolution output.
[0173]
[0105] Clause 2: The method according to clause 1, further comprising performing a tabulation operation based on a determination that the input data has a sparsity greater than a threshold sparsity value.
[0174]
[0106] Clause 3: The method according to any one of clauses 1 to 2, wherein performing the convolution operation comprises performing a matrix multiplication between a weight tensor and the tabular representation of the input data to generate a convolution output.
[0175]
[0107] Clause 4: The method according to any one of clauses 1 to 3, wherein performing the tabulation operation comprises populating the tabular representation of the input data according to indexing rules, where the indexing rules define a plurality of index values based on a relationship between active elements of the input data and a plurality of elements of the input data adjacent to the active elements.
[0176]
[0108] Clause 5: The method according to any one of clauses 1 to 4, wherein the sparsity of the convolution output is the same as the sparsity of the input data.
[0177] Clause 6: Before performing the convolution operation, further comprising deleting the active point vector components from the tabular representation of the input data, and adding the active point vector components to the convolution output to generate the convolution layer output, the method according to any one of Clauses 1 to 5.
[0178]
[0110] Clause 7: Before performing the convolution operation, further comprising deleting the active point vector components from the tabular representation of the input data and generating a plurality of one-dimensional hashed pointer lists based on the tabular representation, wherein the convolution operation multiplies each input value associated with a pointer in the first hashed pointer list among the plurality of one-dimensional hashed pointer lists by a scalar weight value based on the kernel index of the input value, and adding the weighted input values associated with each pointer in the second hashed pointer list among the plurality of one-dimensional hashed pointer lists to generate the convolution output, the method according to Clause 4.
[0179]
[0111] Clause 8: Further comprising determining a loss value associated with the convolution output and updating a plurality of weights associated with the convolution layer of the machine learning model based on the loss value, the method according to any one of Clauses 1 to 7.
[0180]
[0112] Clause 9: The convolution operation comprises a depthwise separable convolution operation, the method according to any one of Clauses 1 to 8.
[0181]
[0113] Clause 10: The input data comprises sparse image sensor data, the method according to any one of Clauses 1 to 9.
[0182]
[0114] Clause 11: The input data comprises sparse light detection and ranging (LiDAR) sensor data, the method according to any one of Clauses 1 to 10.
[0183] Clause 12: A processing system comprising a memory having computer-executable instructions and one or more processors, wherein the one or more processors are configured to execute the computer-executable instructions to cause the processing system to perform the method according to any one of Clauses 1 to 11.
[0184] Clause 13: A processing system comprising means for performing the method according to any one of Clauses 1 to 11.
[0185] Clause 14: A non-transitory computer-readable medium having computer-executable instructions, wherein when the computer-executable instructions are executed by one or more processors of a processing system, the processing system is caused to perform the method according to any one of Clauses 1 to 11.
[0186] Clause 15: A computer program product embodied on a computer-readable storage medium having code for performing the method according to any one of Clauses 1 to 11. Additional Considerations
[0119] The foregoing description has been provided to enable a person of ordinary skill in the art to make and use the various embodiments described herein. The examples described herein are not intended to limit the scope, applicability, or embodiments described in the claims. Various modifications to these embodiments will be readily apparent to those of ordinary skill in the art, and the general principles defined herein may be applied to other embodiments. For example, changes may be made in the function and arrangement of the elements described without departing from the scope of the present disclosure. The various examples may, as appropriate, omit, substitute, or add various procedures or components. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, the features described with respect to some examples may be combined in some other examples. For example, the apparatus may be implemented or the method may be performed using any number of the aspects described herein. Further, the scope of the present disclosure is intended to cover such apparatus or methods implemented using other structures, functions, or structures and functions in addition to, or other than, the various aspects of the present disclosure described herein. It should be understood that any aspect of the present disclosure disclosed herein may be implemented by one or more elements of the claims.
[0187]
[0120] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other aspects.
[0188] As used herein, the phrase "at least one of" in a list of items refers to any combination of those items including a single member. By way of example, "at least one of a, b, or c" includes a, b, c, a - b, a - c, b - c, and a - b - c, as well as any combination having multiple of the same element (e.g., a - a, a - a - a, a - a - b, a - a - c, a - b - b, a - c - c, b - b, b - b - b, b - b - c, c - c, and c - c - c, or any other order of a, b, and c).
[0189] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database or another data structure), ascertaining, etc. Further, "determining" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Further, "determining" can include resolving, selecting, choosing, establishing, etc.
[0190]
[0123] The methods disclosed herein comprise one or more steps or actions for achieving the methods. The steps and / or actions of the present method may be exchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be changed without departing from the scope of the claims. Further, the various operations of the methods described above may be implemented by any suitable means capable of performing the corresponding functions. Those means may include various (one or more) hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, if there are operations shown in the figures, those operations may have corresponding means-plus-function components with similar numbers.
[0191]
[0124] The following claims are not limited to the embodiments shown herein, but should be given the full scope not inconsistent with the language of the claims. In the claims, reference to an element in the singular is not intended to mean "sole and exclusive" unless so stated, but rather "one or more." Unless otherwise specified, the term "some" refers to one or more. No claim element should be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the element is expressly recited using the phrase "step for." All structural and functional equivalents of the elements of the various aspects described throughout this disclosure, known or later to be known to those of ordinary skill in the art, are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is dedicated to the public, whether or not such disclosure is expressly recited in the claims. The invention described in the claims of the present application at the time of filing is appended below. [C1] Performing a tabulation operation on the input data to generate a tabular expression of the input data, Performing a convolution operation using the tabular expression of the input data to generate a convolution output, A method comprising: [C2] The method according to C1, further comprising performing the tabulation operation based on a determination that the input data has sparsity greater than a threshold sparsity value. [C3] The method according to C1, wherein performing the convolution operation comprises performing a matrix multiplication between a weight tensor and the tabular expression of the input data to generate the convolution output. [C4] Performing the tabulation operation comprises populating the tabular expression of the input data according to an indexing rule, wherein the indexing rule defines a plurality of index values based on a relationship between active elements of the input data and a plurality of elements of the input data adjacent to the active elements. The method according to C1. [C5] The method according to C1, wherein the sparsity of the convolution output is the same as the sparsity of the input data. [C6] Before performing the convolution operation, deleting active point vector components from the tabular expression of the input data, Adding the active point vector components to the convolution output to generate a convolution layer output. The method according to C1, further comprising: [C7] Before performing the convolution operation, deleting active point vector components from the tabular expression of the input data, Further comprising generating a plurality of one-dimensional hashed pointer lists based on the tabular expression, The convolution operation comprises multiplying input values associated with each pointer in a first hashed pointer list among the plurality of one-dimensional hashed pointer lists by a scalar weight value based on a kernel index of the input value, Adding the weighted input values associated with each pointer in a second hashed pointer list among the plurality of one-dimensional hashed pointer lists to generate the convolution output. The method according to C4, comprising: [C8] Determining a loss value associated with the convolution output Updating a plurality of weights associated with the convolutional layer of the machine learning model based on the loss value; The method according to C1, further comprising. [C9] The method according to C1, wherein the convolutional operation comprises a depthwise separable convolutional operation. [C10] The method according to C1, wherein the input data comprises sparse image sensor data. [C11] The method according to C1, wherein the input data comprises sparse light detection and ranging (LiDAR) sensor data. [C12] A memory comprising computer-executable instructions; One or more processors; A processing system comprising, wherein the one or more processors execute the computer-executable instructions, and cause the processing system to Perform a tabulation operation on the input data to generate a tabular representation of the input data; Perform a convolutional operation using the tabular representation of the input data to generate a convolutional output; A processing system configured to perform. [C13] The processing system according to C12, wherein the one or more processors are further configured to perform the tabulation operation based on a determination that the input data has a sparsity greater than a threshold sparsity value. [C14] The processing system according to C12, wherein, to perform the convolutional operation, the one or more processors are further configured to perform a matrix multiplication between a weight tensor and the tabular representation of the input data to generate the convolutional output. [C15] To perform a tabulation operation, the one or more processors Are further configured to populate the tabular representation of the input data according to an indexing rule, The indexing rule defines a plurality of index values based on a relationship between active elements of the input data and a plurality of elements of the input data adjacent to the active elements. The processing system according to C12. [C16] The processing system according to C12, wherein the sparsity of the convolutional output is the same as the sparsity of the input data. [C17] The one or more processors Before performing the convolutional operation, delete active point vector components from the tabular representation of the input data; To generate a convolutional layer output, add the active point vector components to the convolutional output; The processing system according to C12, further configured to perform. [C18] The one or more processors Before performing the convolution operation, deleting active point vector components from the tabular representation of the input data; Further configured to generate a plurality of one-dimensional hashed pointer lists based on the tabular representation; To perform the convolution operation, the one or more processors Multiplying input values associated with each pointer in a first hashed pointer list among the plurality of one-dimensional hashed pointer lists by a scalar weight value based on the kernel index of the input value; To generate the convolution output, summing the weighted input values associated with each pointer in a second hashed pointer list among the plurality of one-dimensional hashed pointer lists; The processing system according to C15, further configured to perform. [C19] The one or more processors Determining a loss value associated with the convolution output; Updating a plurality of weights associated with the convolutional layer of the machine learning model based on the loss value; The processing system according to C12, further configured to perform. [C20] The convolution operation comprises a depthwise separable convolution operation, the processing system according to C12. [C21] The input data comprises sparse image sensor data, the processing system according to C12. [C22] The input data comprises sparse light detection and ranging (LiDAR) sensor data, the processing system according to C12. [C23] A non-transitory computer-readable medium comprising computer-executable instructions, which, when executed by one or more processors of a processing system, cause the processing system to Perform a tabulation operation on the input data to generate a tabular representation of the input data; Perform a convolution operation using the tabular representation of the input data to generate a convolution output; A non-transitory computer-readable medium for implementing a method comprising. [C24] The method further comprises performing the tabulation operation based on determining that the input data has a sparsity greater than a threshold sparsity value, the non-transitory computer-readable medium according to C23. [C25] Performing the convolution operation comprises performing a matrix multiplication between a weight tensor and the tabular representation of the input data to generate the convolution output, the non-transitory computer-readable medium according to C23. [C26] Performing the arithmetic operation includes populating the tabular representation of the input data according to the indexing rules, wherein the indexing rules define a plurality of index values based on a relationship between an active element of the input data and a plurality of elements of the input data adjacent to the active element. The non-transitory computer-readable medium according to C23. [C27] The non-transitory computer-readable medium according to C23, wherein the sparsity of the convolution output is the same as the sparsity of the input data. [C28] The method further includes deleting active point vector components from the tabular representation of the input data prior to performing the convolution operation; adding the active point vector components to the convolution output to generate a convolution layer output. The non-transitory computer-readable medium according to C23, further comprising the above. [C29] The method further includes deleting active point vector components from the tabular representation of the input data prior to performing the convolution operation; generating a plurality of one-dimensional hashed pointer lists based on the tabular representation. The convolution operation includes multiplying input values associated with each pointer in a first hashed pointer list among the plurality of one-dimensional hashed pointer lists by a scalar weight value based on a kernel index of the input values; adding the weighted input values associated with each pointer in a second hashed pointer list among the plurality of one-dimensional hashed pointer lists to generate the convolution output. The non-transitory computer-readable medium according to C26, further comprising the above. [C30] The method further includes determining a loss value associated with the convolution output; updating a plurality of weights associated with the convolution layer of the machine learning model based on the loss value. The non-transitory computer-readable medium according to C23, further comprising the above.
Claims
**Claim 1** A method for improving the efficiency of convolution processing, the method being implemented by a processing system, the method comprising: performing a tabulation operation on the input data to generate a tabular representation of the input data, wherein performing the tabulation operation comprises: populating the tabular representation of the input data according to indexing rules; The indexing rules define a plurality of index values based on the relationship between the active elements of the input data and a plurality of elements of the input data adjacent to the active elements; removing active point vector components from the tabular representation of the input data prior to performing the convolution operation; performing a convolution operation using the tabular representation of the input data to generate a convolution output; adding the active point vector components to the convolution output to generate a convolution layer output; A method comprising: **Claim 2** The method according to claim 1, further comprising performing the tabulation operation based on a determination that the input data has a sparsity greater than a threshold sparsity value. **Claim 3** The method according to claim 1, wherein performing the convolution operation comprises performing a matrix multiplication between a weight tensor and the tabular representation of the input data to generate the convolution output. **Claim 4** The method according to claim 1, wherein the sparsity of the convolution output is the same as the sparsity of the input data. **Claim 5** Further comprising generating a plurality of one-dimensional hashed pointer lists based on the tabular representation, The convolution operation comprises: multiplying an input value associated with each pointer in a first hashed pointer list of the plurality of one-dimensional hashed pointer lists by a scalar weight value based on the kernel index of the input value; summing the weighted input values associated with each pointer in a second hashed pointer list of the plurality of one-dimensional hashed pointer lists to generate the convolution output; The method according to claim 1, comprising: **Claim 6** determining a loss value associated with the convolution output; updating a plurality of weights associated with the convolution layer of the machine learning model based on the loss value; The method according to claim 1, further comprising: **Claim 7** The convolution operation is the method according to claim 1, comprising a depthwise separable convolution operation.
8. The input data comprises sparse image sensor data or sparse light detection and ranging (LiDAR) sensor data, and the method according to claim 1.
9. A memory comprising computer-executable instructions, One or more processors, A processing system for improving convolution processing efficiency, comprising: the one or more processors execute the computer-executable instructions, and the processing system Performing a tabulation operation on the input data to generate a tabulated representation of the input data, wherein, to perform the tabulation operation, the one or more processors Further configured to populate the tabulated representation of the input data according to an indexing rule. The indexing rule defines a plurality of index values based on the relationship between the active elements of the input data and a plurality of elements of the input data adjacent to the active elements. Before performing the convolution operation, deleting the active point vector components from the tabulated representation of the input data; Performing a convolution operation using the tabulated representation of the input data to generate a convolution output; Adding the active point vector components to the convolution output to generate a convolution layer output; A processing system configured to perform the above.
10. The one or more processors are further configured to perform the tabulation operation based on a determination that the input data has a sparsity greater than a threshold sparsity value, and the processing system according to claim 9.
11. To perform the convolution operation, the one or more processors are further configured to perform a matrix multiplication between the weight tensor and the tabulated representation of the input data to generate the convolution output, and the processing system according to claim 9.
12. The sparsity of the convolution output is the same as the sparsity of the input data, and the processing system according to claim 9.
13. The one or more processors Are further configured to generate a plurality of one-dimensional hashed pointer lists based on the tabulated representation. To perform the convolution operation, the one or more processors multiply an input value associated with each pointer in a first hashed pointer list among the plurality of one-dimensional hashed pointer lists by a scalar weight value based on a kernel index of the input value; To generate the convolution output, sum weighted input values associated with each pointer in a second hashed pointer list among the plurality of one-dimensional hashed pointer lists; The processing system according to claim 9, further configured to perform.
14. The one or more processors determine a loss value associated with the convolution output; update a plurality of weights associated with the convolutional layer of the machine learning model based on the loss value; The processing system according to claim 9, further configured to perform.
15. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Three-dimensional convolution operation device, visual odometry system, and three-dimensional convolution program
JP2019211879A
Real time prediction of object behavior
JP2020083308A
Neural network devices and methods of operating the same
US20180253635A1