Isometric polyhedron spherical gauge Convolutional Neural Network
By generating and utilizing locally defined metrics on spherical manifolds to compute convolutions and transform results based on metric invariance, the method addresses the complexity of conventional convolution operators on spherical surfaces, achieving efficient and scalable spherical convolutional neural networks.
Patent Information
- Application Number
- CN202080065492.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-23
- Filing Date
- 2020-09-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-09-24
AI Technical Summary
Convolution operators on conventional mesh cannot be directly extended to spherical manifolds, resulting in complex and inconvenient convolution calculations on spherical manifolds.
By generating locally defined gauges and calculating convolutions on spherical manifolds based on degeneration of gauges, an efficient convolution network is designed using degeneration convolutions such as gauges.
It realizes efficient convolution operation on spherical manifold, which is suitable for processing spherical signals, and improves calculation efficiency and accuracy.
Smart Images

Figure CN114830131B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 17 / 030,361, filed on September 23, 2020, entitled "ICOSPHERICAL GAUGE CONVOLUTIONAL NEURAL NETWORK", which in turn claims the benefit of U.S. Provisional Patent Application No. 62 / 905,233, filed on September 24, 2019, entitled "ICOSPHERICAL GAUGE CONVOLUTIONAL NEURAL NETWORK", the disclosures of which are hereby incorporated by reference in their entireties.
[0003] Field of Disclosure
[0004] Aspects of the present disclosure generally relate to artificial neural networks. More specifically, the present disclosure relates to an icospherical gauge convolutional neural network.
[0005] Background
[0006] The computational simplicity and efficiency of convolutional operators on a regular grid (e.g., an image plane) do not extend to other grids / manifolds. For example, convolutional operators on a regular grid do not extend to a spherical manifold, which is the natural embedding space for omnidirectional panoramic signals obtained via an appropriate imaging setup. Additionally, due to the ambiguity and non - uniqueness of local reference frames, conventional convolution computations on a spherical manifold are not straightforward. Accordingly, coefficient kernels cannot be shifted on a spherical manifold simply by shifting.
[0007] Summary
[0008] In one aspect of the present disclosure, a method is provided. The method includes generating locally - defined gauges at multiple positions on a spherical manifold. The method also includes computing a convolution at each of the multiple positions on the spherical manifold with reference to the locally - defined gauges. Further, the method includes transforming the result of the convolution at each position based on gauge equivariance to obtain a corresponding manifold transformation.
[0009] In another aspect of the present disclosure, an apparatus is provided. The apparatus includes a memory and one or more processors coupled to the memory. The processor(s) is configured to generate locally - defined gauges at multiple positions on a spherical manifold. The processor(s) is also configured to compute a convolution at each of the multiple positions on the spherical manifold with reference to the locally - defined gauges. Additionally, the processor(s) is configured to transform the result of the convolution at each position based on gauge equivariance to obtain a corresponding manifold transformation.
[0010] In another aspect of the present disclosure, an apparatus is provided. The apparatus includes means for generating locally defined gauges at multiple positions on a spherical manifold. The apparatus further includes means for calculating a convolution at each of the multiple positions on the spherical manifold with reference to the locally defined gauges. Additionally, the apparatus includes means for transforming the result of the convolution at each position based on gauge equivariance to obtain a corresponding manifold transformation.
[0011] In a further aspect of the present disclosure, a non-transitory computer-readable medium is provided. The computer-readable medium has program code encoded thereon. The program code is executed by a processor and includes code for generating locally defined gauges at multiple positions on a spherical manifold. The program code further includes code for calculating a convolution at each of the multiple positions on the spherical manifold with reference to the locally defined gauges. Additionally, the program code includes code for transforming the result of the convolution at each position based on gauge equivariance to obtain a corresponding manifold transformation.
[0012] Additional features and advantages of the present disclosure will be described below. Those skilled in the art should appreciate that the present disclosure can be readily used as a basis for modifying or designing other structures for implementing the same purpose as the present disclosure. Those skilled in the art should also recognize that such equivalent constructs do not depart from the teachings of the present disclosure as set forth in the appended claims. The novel features, which are considered to be characteristics of the present disclosure, will be better understood in terms of their organization and method of operation, together with further objects and advantages, when considered in conjunction with the following description taken in connection with the accompanying drawings. It is to be clearly understood, however, that each of the accompanying drawings is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure. Brief Description of the Drawings
[0014] The features, nature, and advantages of the present disclosure will become more apparent when the detailed description set forth below is understood in conjunction with the accompanying drawings, in which like reference numerals throughout the drawings consistently identify corresponding elements.
[0015] Figure 1 Illustrates an example implementation of designing a neural network using a system-on-chip (SOC) (including a general-purpose processor) in accordance with certain aspects of the present disclosure.
[0016] Figure 2A 、 2B and 2C are diagrams illustrating neural networks in accordance with aspects of the present disclosure.
[0017] Figure 2D is a diagram illustrating an exemplary deep convolutional network (DCN) in accordance with aspects of the present disclosure.
[0018] Figure 3is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.
[0019] Figure 4A A conventional dodecahedron approximating a sphere with a flat surface is illustrated in accordance with aspects of the present disclosure.
[0020] Figure 4B Isosurface polyhedral spheres in accordance with aspects of the present disclosure are illustrated.
[0021] Figure 5A is a diagram illustrating an exponential mapping from a sphere or spherical manifold to a tangent plane for a gauge invariant transformation in accordance with aspects of the present disclosure.
[0022] Figure 5B A tangent plane showing a point of interest along with interpolated points on the tangent plane is illustrated in accordance with aspects of the present disclosure.
[0023] Figure 6 Methods for generating a convolutional neural network operating on a spherical manifold in accordance with aspects of the present disclosure are explained.
[0024] Detailed Description
[0025] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the described concepts may be practiced. This detailed description includes specific details in order to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0026] Based on this teaching, it will be appreciated by those skilled in the art that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether it is implemented independently or in combination with any other aspect of the present disclosure. For example, any number of aspects described can be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods that are practiced using other structures, functionality, or structures and functionality that are supplemented or different from the various aspects of the present disclosure described. It should be understood that any aspect of the present disclosure disclosed can be implemented by one or more elements of the claims.
[0027] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.
[0028] Although specific aspects are described, numerous variations and permutations of these aspects fall within the scope of the present disclosure. While some benefits and advantages of the preferred aspects are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses, or objectives. Rather, the aspects of the present disclosure are intended to be broadly applicable to different technologies, system configurations, networks, and protocols, some of which are illustrated by way of example in the figures and the following description of the preferred aspects. The detailed description and the figures merely illustrate the present disclosure and do not limit the present disclosure, the scope of which is defined by the appended claims and their equivalent technical solutions.
[0029] An artificial neural network may include a group of interconnected artificial neurons (e.g., neuron models). An artificial neural network may be a computing device or represent a method executed by a computing device. A convolutional neural network is a feedforward artificial neural network. A convolutional neural network may include a set of neurons, each of which has a receptive field and together tile an input space. Convolutional neural networks (CNNs), such as deep convolutional neural networks (DCNs), have numerous applications. In particular, these neural network architectures are used in various technologies, such as image recognition, pattern recognition, speech recognition, autonomous driving, and other classification tasks.
[0030] A spherical CNN is a convolutional neural network that can process signals on a sphere, such as global climate and weather patterns or omnidirectional images. The computational simplicity and efficiency of convolutional operators on a regular grid (e.g., an image plane) do not extend to other grids / manifolds. For example, convolutional operators on a regular grid do not extend to a spherical manifold, which is the natural embedding space for omnidirectional panoramic signals obtained via appropriate imaging setups. Additionally, due to the ambiguity and non-uniqueness of local reference frames, conventional convolution computations on a spherical manifold are not straightforward. Accordingly, shifting coefficient kernels on a spherical manifold is complex and cumbersome.
[0031] In many scientific and engineering disciplines, spherical signals arise naturally. In earth and climate science, globally distributed sensor arrays collect measurements such as temperature, pressure, wind direction, and many other variables. Cosmologists are interested in identifying physical model parameters from real and simulated cosmic microwave background measurements sampled from spherical sky maps. In robotics, omnidirectional and fisheye cameras are widely used, especially in applications such as simultaneous localization and mapping (SLAM) and visual odometry. An efficient CNN that operates directly on spherical signals can be beneficial.
[0032] Aspects of the present disclosure aim to design an efficient convolutional network implementation operable on spherical signals using gauge-equivariant convolutions. The principle of equivariance under symmetry transformations provides a theoretically grounded approach to neural network architecture design. Equivariant networks have shown excellent performance and data efficiency in visual and medical imaging problems exhibiting symmetry. This principle can be extended from global symmetries to local gauge transformations, enabling the development of a very general class of convolutional neural networks on manifolds that rely only on intrinsic geometry and include many common methods from equivariant and geometric deep learning.
[0033] The equivariance principle is used to implement a gauge-equivariant convolutional neural network (CNN) for signals defined on the surface of an isogonal polyhedron, which provides a reasonable approximation of the sphere. Gauge-equivariant convolutions can be implemented using a single two-dimensional convolution (conv2d) call, making the implementation highly scalable and a practical replacement for spherical CNNs. The theory of gauge-equivariant networks is applied to manifolds (e.g., isogonal polyhedra). The manifold includes global symmetries (e.g., discrete rotations) that exhibit the differences between local symmetries and the interaction of global symmetries. The shape of the manifold enables the implementation of gauge-equivariant convolutions in a way that is both numerically convenient (without specifying interpolation) and computationally efficient (the heavy lifting is done by a single two-dimensional convolution (conv2d) call). However, conventional implementations on isogonal polyhedra are limited to fixed kernels and are only equivariant under up to sixty rotational symmetries of the regular icosahedron.
[0034] Aspects of the present disclosure relate to a method for generating a convolutional neural network operable on a spherical manifold. The proposed method includes generating locally defined gauges at multiple positions on the spherical manifold. The method further includes defining a convolution at each of the multiple positions on the spherical manifold with reference to an arbitrarily chosen locally defined gauge. Additionally, the method includes transforming the result of the convolution defined at each position based on gauge-equivariance to obtain a manifold convolution.
[0035] In one aspect, the spherical manifold or sphere is parameterized as an isogonal polyhedron mesh. The manifold convolution can be distributed to local neighborhoods of the spherical manifold based on locally defined gauges. Each kernel associated with each of the multiple positions is a locally varying kernel derived from the same function. Each defined convolution at each position is computed by a locally connected layer. Thus, a reference frame or arbitrary gauge is selected at each position on the sphere, and the convolution is computed at each position and then combined to form a final result.
[0036] In one aspect, a gauge transformation and its corresponding representation are applied to a two-dimensional convolution to obtain a generalized definition of the convolution operation. At a particular position on an arbitrary manifold, features are transformed to the reference frame or locally defined gauge at that particular position.
[0037] Aspects of the present disclosure relate to processing signals of the spherical type. Examples of spherical signals can come from imaging devices such as fisheye, panoramic, or omnidirectional cameras. Accordingly, the proposed implementations have many practical applications including, but not limited to, image recognition, image segmentation, and detection on recording devices for the above applications.
[0038] Aspects of the present disclosure analyze inputs in a spherical domain or manifold (e.g., signals such as global temperature or climate data). The proposed implementations impact climate science, e.g., for weather forecasting using machine learning. For example, the proposed implementations can be used to analyze trends in global temperature over time in response to global warming. Similarly, it can be used to track changes on Earth via imagery obtained from satellites.
[0039] Another example application includes autonomous vehicles. For example, autonomous driving software can use the proposed implementations by processing 360-degree images collected from the surrounding environment for purposes such as collision avoidance, pedestrian detection, localization, etc. In one aspect, cameras are used to classify and identify objects, and the input to the cameras can be projected onto a spherical manifold.
[0040] The proposed system can also be used for cosmological processing. Cosmological data can include data of observations of the universe processed according to a spherical domain. For example, the proposed implementations are suitable for processing and analyzing cosmological data where the efficiency of the processing implementation is particularly important due to the large amount of data collected from the observable universe. Tasks can include detecting black holes, or other faint signals from distant galaxies or stars. Other applications for the proposed implementations include shape analysis, molecular modeling, and three-dimensional (3D) shape recognition. For example, according to aspects of the present invention, shape models can be indirectly projected onto a sphere and analyzed.
[0041] Although aspects of the present disclosure are described with reference to a spherical manifold, the same implementations can be extended to an arbitrary manifold with little or no modification, thereby making it a general process for geometric deep learning.
[0042] Figure 1An example implementation of a system-on-chip (SOC) 100 is described, which may include a central processing unit (CPU) 102 or a multi-core CPU configured for efficient processing of convolutional neural networks. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., neural network with weights), latencies, frequency slot information, and task information may be stored in a memory block associated with the neural processing unit (NPU) 108, in a memory block associated with the CPU 102, in a memory block associated with the graphics processing unit (GPU) 104, in a memory block associated with the digital signal processor (DSP) 106, in memory block 118, or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from memory block 118.
[0043] The SOC 100 may also include additional processing blocks customized for specific functions, such as the GPU 104, the DSP 106, the connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that can detect and recognize poses, for example. In one implementation, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SOC 100 may also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which may include a global positioning system).
[0044] The SOC 100 may be based on the ARM instruction set. In one aspect of the present disclosure, the instructions loaded into the general-purpose processor 102 may include: code for generating locally defined gauges at multiple locations on a spherical manifold, code for calculating a convolution at each of the multiple locations on the spherical manifold with reference to the locally defined gauges, and code for transforming the result of the convolution at each location based on gauge equivariance to obtain a corresponding manifold transformation.
[0045] Deep learning architectures can perform object recognition tasks by learning to represent the input at successively higher levels of abstraction in each layer, thereby constructing useful feature representations of the input data. In this way, deep learning addresses the main bottlenecks of traditional machine learning. Before the advent of deep learning, machine learning approaches for object recognition problems might rely heavily on human-engineered features, perhaps combined with shallow classifiers. Shallow classifiers can be two-class linear classifiers, for example, where the weighted sum of the feature vector components is compared to a threshold to predict which class the input belongs to. Human-engineered features can be templates or kernels customized by engineers with domain expertise for a specific problem domain. In contrast, deep learning architectures can learn to represent features similar to those that human engineers might design, but it learns through training. Additionally, deep networks can learn to represent and identify new types of features that humans might not have considered yet.
[0046] Deep learning architectures can learn a hierarchy of features. For example, if visual data is presented to the first layer, the first layer can learn to identify relatively simple features (such as edges) in the input stream. In another example, if auditory data is presented to the first layer, the first layer can learn to identify spectral power in specific frequencies. A second layer that takes the output of the first layer as input can learn to identify combinations of features, such as identifying simple shapes for visual data or sound combinations for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to identify common visual objects or spoken phrases.
[0047] Deep learning architectures may perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motor vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0048] Neural networks can be designed with various connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, where each neuron in a given layer communicates to neurons in higher layers. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one chunk of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be beneficial when the recognition of high-level concepts can assist in discerning specific low-level features of the input.
[0049] The connections between the layers of a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, a neuron in the first layer can convey its output to every neuron in the second layer, such that every neuron in the second layer will receive inputs from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, a neuron in the first layer can be connected to a limited number of neurons in the second layer. More generally, the locally connected layer of the locally connected neural network 204 can be configured such that each neuron in one layer will have the same or similar connectivity pattern, although the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern may result in spatially distinct receptive fields in higher layers, since the higher layer neurons in a given region can receive inputs that are tuned by training to the properties of a restricted portion of the total input to the network.
[0050] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. The convolutional neural network 206 can be configured such that the connection strengths associated with the inputs for each neuron in the second layer are shared (e.g., 208). Convolutional neural networks can be well-suited to problems where the spatial location of the input is meaningful.
[0051] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to recognize visual features from an image 226 input from an image capture device 230 (such as an in-vehicle camera) is illustrated. The DCN 200 of the current example can be trained to identify traffic signs and the digits provided on the traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signals.
[0052] The DCN 200 can be trained using supervised learning. During training, an image (such as the image 226 of a speed limit sign) can be presented to the DCN 200, and then a "forward pass" can be computed to produce an output 222. The DCN 200 can include a feature extraction section and a classification section. Upon receiving the image 226, the convolutional layer 232 can apply a convolutional kernel (not shown) to the image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of the convolutional layer 232 can be a 5x5 kernel that generates 28x28 feature maps. In this example, since four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to the image 226 at the convolutional layer 232. The convolutional kernel can also be referred to as a filter or a convolutional filter.
[0053] The first set of feature maps 218 can be subsampled by a max pooling layer (not shown) to generate a second set of feature maps 220. The max pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (such as 14x14) is smaller than the size of the first set of feature maps 218 (such as 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate subsequent sets of one or more feature maps (not shown).
[0054] In Figure 2D the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Additionally, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 can include a number corresponding to a possible feature of the image 226 (such as "sign", "60", and "100"). A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0055] In this example, the probabilities of "sign" and "60" in the output 222 are higher than the probabilities of other features (such as "30", "40", "50", "70", "80", "90", and "100") of the output 222. Before training, the output 222 produced by the DCN 200 is likely to be incorrect. Thus, the error between the output 222 and the target output can be computed. The target output is the ground truth of the image 226 (e.g., "sign" and "60"). The weights of the DCN 200 can then be adjusted so that the output 222 of the DCN 200 aligns more closely with the target output.
[0056] To adjust the weights, the learning algorithm can compute a gradient vector for the weights. This gradient can indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top layer, this gradient can directly correspond to the value of the weights connecting the activated neurons in the penultimate layer to the neurons in the output layer. In the lower layers, the gradient can depend on the values of the weights and the error gradients computed for the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be called "backpropagation" because it involves a "backward pass" in the neural network.
[0057] In practice, the error gradient of the weights may be computed on a small number of examples so that the computed gradient approximates the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the error rate achievable by the entire system has stopped decreasing or until the error rate has reached a target level. After learning, new images can be presented to the DCN and a forward pass in the network can produce an output 222, which can be considered an inference or prediction of the DCN.
[0058] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training data set. A DBN can be obtained by stacking multiple layers of restricted Boltzmann machines (RBMs). An RBM is a type of artificial neural network that can learn a probability distribution over an input set. Since an RBM can learn a probability distribution without information about which class each input should be classified into, RBMs are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained in an unsupervised manner and used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the inputs from the previous layer and the target classes) and used as a classifier.
[0059] A deep convolutional network (DCN) is a network of convolutional networks that is configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. A DCN can be trained using supervised learning, where both the inputs and the output targets are known for many paradigms and are used to modify the weights of the network using gradient descent.
[0060] A DCN can be a feedforward network. Additionally, as described above, the connections from the neurons in the first layer of the DCN to a group of neurons in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of the DCN can be exploited for fast processing. The computational burden of a DCN can be much smaller than that of a neural network of a similar size that includes recurrent or feedback connections, for example.
[0061] The processing of each layer of a convolutional network can be considered as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then the convolutional network trained on this input can be considered three-dimensional, having two spatial dimensions along the axes of the image and a third dimension capturing color information. The output of a convolutional connection can be considered to form a feature map in subsequent layers, where each element in the feature map (e.g., 220) receives inputs from a certain range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with a non-linearity such as rectification, max(0, x). Values from adjacent neurons can be further pooled (which corresponds to downsampling) and can provide additional local invariance as well as dimensionality reduction. Normalization can also be applied through lateral inhibition between neurons in the feature map, which corresponds to whitening.
[0062] The performance of deep learning architectures can improve as more labeled data points become available or as computing power increases. Modern deep neural networks are routinely trained with thousands of times more computing resources than were available to a typical researcher just fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Rectified linear units can reduce a training problem known as vanishing gradients. New training techniques can reduce over-fitting and thus enable larger models to achieve better generalization. Encapsulation techniques can abstract the data within a given receptive field and further enhance overall performance.
[0063] Figure 3 is a block diagram illustrating a deep convolutional network 350. The deep convolutional network 350 can include multiple different types of layers based on connectivity and weight sharing. As Figure 3 shown, the deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.
[0064] The convolutional layer 356 can include one or more convolutional filters, which can be applied to the input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not limited thereto, but instead any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preferences. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max pooling layer 360 can provide spatially downsampled aggregation to achieve local invariance as well as dimensionality reduction.
[0065] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the deep convolutional network 350 can access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and the navigation module 120 dedicated to sensors and navigation, respectively.
[0066] The deep convolutional network 350 may further include one or more fully connected layers 362 (FC1 and FC2). The deep convolutional network 350 may further include a logistic regression (LR) layer 364. Weights (not shown) to be updated exist between each layer 356, 358, 360, 362, 364 of the deep convolutional network 350. The output of each layer (e.g., 356, 358, 360, 362, 364) can be used as the input to a subsequent layer (e.g., 356, 358, 360, 362, 364) in the deep convolutional network 350 to learn a hierarchical feature representation from the input data 352 (e.g., an image, audio, video, sensor data, and / or other input data) supplied at the first convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 for the input data 352. The classification score 366 can be a set of probabilities, where each probability is the probability that the input data includes features from a feature set.
[0067] Aspects of the present disclosure aim to design an efficient convolutional network process that can operate on non-planar (e.g., spherical) signals using gauge-equivariant convolutions. The locally defined gauge on the spherical manifold S 2 assigns to each point p on the manifold a linear map from the standard plane to the tangent plane T p S 2 at the point p of the sphere. The locally defined gauge allows the computation of the manifold convolution to be distributed over the local neighborhood. For example, at each position or point p in the spherical manifold S 2 , the convolution can be defined with reference to an arbitrarily chosen gauge. Finally, gauge-equivariance ensures that the results of the local computations can be meaningfully transformed into each other. That is, the corresponding manifold transformation can be determined based on gauge-equivariance.
[0068] The feature space in a gauge convolutional neural network (CNN) can be modeled as a field f on a manifold M. For example, the input data can be a vector field of wind directions on the Earth, or a scalar field of intensity values on a plane (e.g., a grayscale image), or a diffusion tensor field on. Such quantities (e.g., scalars, vectors, tensors, etc.) can be referred to as geometric features, and these geometric features can be applied to the geometric feature field.
[0069] In computer science, a vector or tensor can be regarded as a list or array of numbers, but from a physical or mathematical perspective, these are geometric quantities that exist independently of the choice of coordinates or basis. However, to represent geometric features numerically, a frame for the tangent plane space T can be chosen at each position p ∈ M. A smooth choice of frame is a gauge. Mathematically, a gauge on a d-dimensional manifold M can be defined as a set of linear maps, which are smoothly parameterized by points p on the manifold: p (see Equation 1). On a manifold with an orientable metric tensor (such as a sphere), the choice of gauge can be restricted to a set of oriented orthonormal gauges. In this case, any two gauges w, w′ are related at point p by an element r in the d-dimensional rotation group SO(d) such that (see Equation 1). On a manifold with an orientable metric tensor (such as a sphere), the choice of gauge can be restricted to a set of oriented orthonormal gauges. In this case, any two gauges w, w′ are related at point p by an element r in the d-dimensional rotation group SO(d) such that p related, such that
[0070] The application of a gauge transformation can affect the coefficients of geometric features. This is because the choice of gauge is arbitrary. First, consider the coefficients of the standard basis vectors (c1, e2) for at a point or position p. The tangent vector V at position p ∈ M with f(p) = v, expressed as a pair of numbers v = (v1, v2) relative to the orthonormal frame (w p M in the tangent space T p (e1), w p (e2)). If the frame is rotated at position p by an element r in the plane rotation group SO(2) using Equation then the coefficient vector transforms as Regarding the plane rotation r ∈ SO(2) as a two-by-two matrix acting on these two coefficients of v in its matrix representation. A vector is an abstract geometric quantity that is invariant under gauge transformation, such that: V = (w p r)r -1 v = w p v. In some aspects, a gauge transformation can be defined as a smoothly varying choice of rotation r p ∈ SO(2). However, the present disclosure is not limited to this, and a gauge transformation can also be defined otherwise.
[0071] Beyond scalars (which are invariant under gauge transformation) and vectors (which transform like ), more general kinds of geometric features can be considered. For example, a (2,0) tensor is a linear combination of the tensor product p of vectors V, W ∈ T (see Equation 1). Given a frame, such a tensor can be represented as a d×d matrix. Under a change of frame, the matrix f(p) can transform like The matrix f(p) can be flattened into a d 2 dimensional coordinate vector f(p), and the transformation can be expressed as where is the Kronecker product.
[0072] Tensor product is an example of a group representation. The group representation can be a mapping that takes each element r of G (where G is SO(2)) and maps it to an invertible matrix ρ(r) that acts on a C-dimensional eigenvector. If the invertible matrix ρ(r) satisfies ρ(rr′) = ρ(r)ρ(r′) (which can be checked for the tensor / Kronecker product), then the matrix is considered a representation.
[0073] Thus, the image under a gauge transformation of a geometric feature field that transforms as such can be generalized for any group representation ρ of SO(2). Such fields are called ρ-fields or ρ-type fields. In a gauge-equivariant CNN, a representation ρ that determines the type of features learned by that layer can be chosen for each feature space of the network. The network can be constructed such that a gauge transformation applied to the input can result in a corresponding gauge transformation in each feature space. In one example, ρ can be chosen to be block diagonal, containing, for example, several scalar fields (1×1 blocks p i (r) = 1), several vector fields, etc. The number of copies of each type of feature can be called its multiplicity.
[0074] For each layer of the network, both the input and output can be interpreted as geometric feature fields. Applying a gauge transformation, the input coefficients can change and a gauge transformation is also performed on the output. That is, in aspects of the present disclosure, gauge equivariance can be determined.
[0075] According to aspects of the present disclosure, parallel transport can be applied to the eigenvectors before summing them. Given a curve from q to p, a vector W ∈ T q M can be transported to T p←q M by applying the rotation r p ∈ SO(2) to its coefficient vector w. Since r p←q w can be interpreted as a vector in T p M, the addition of a vector v ∈ T p M, v + r p←q w is well-defined. For other types of geometric features, parallel transport can act via ρ, e.g., the addition v + ρ(r p←q )w.
[0076] A local neighborhood around p ∈ M can be parameterized by the tangent plane via the exponential map. That is, by defining q as v = exp p w pv (which can be referred to as the "normal coordinate") can be used to index the points q near the tangent vector pair by the exponential mapping. Subsequently, the convolution can be defined by the following operation: for each point q near v , by calculating transport the eigenvector f(q v ) to p, use the learned kernel to transform the resulting eigenvector at p, and integrate the result over the support of K in . The convolution operation of the kernel K on the eigenvector f is denoted by :
[0077]
[0078] This operation can be regarded as gauge equivariance if and only if in some aspects K(v) satisfies the following condition:
[0079] K(r -1 v) = ρ out (r -1 )K(v)ρ in (r) (2)
[0080] In addition to gauge equivariance, spherical CNN can also achieve equivariance to any rotation of the sphere by elements of the three-dimensional (3D) rotation group SO(3). That is, if a 3D rotation is applied to the input of the network, the output also rotates.
[0081] In one example, consider a local patch on the sphere (e.g., the support of the kernel) and the signal defined there. If the sphere rotates, the patch is moved to another place, and its orientation may change. Moving the patch may not be a problem: apply the same kernel K at the new position, so it can be expected that the convolution result at the new position is equal to the convolution result of the original signal at the old position. However, since the orientation of the kernel is determined by the gauge (which is arbitrary but fixed), and since the orientation of the patch can be arbitrarily changed by rotating around its center, after applying the rotation, the kernel and the patch may match in different relative orientations. Fortunately, since the kernel satisfies Equation 2, the result is equivalent to the gauge transformation by ρ out acting, and thus SO(3) equivariance is also achieved. Correspondingly, in the continuous theory, gauge-equivariant convolution is also SO(3) equivariant.
[0082] The signal can be represented as a list of values f associated with a finite number of points i = f(p i ). It can be assumed that the kernel K(v) has local support, such that for some radius R, if ||v|| > R, then K(v) = 0. Equivalently, q ∈ S only when the geodesic distance between p and q is less than R.2 can contribute to the convolution result at p ∈ S 2 Accordingly, the neighbor set N(p) of p can be defined as the set of points q within radius R starting from p.
[0083] One way to discretize the gauge convolution (see Equation 1) is to replace the integral over (denoted by T p M) with a sum over the neighbors of p. Each neighbor can be associated with a tangent vector via the logarithmic map: v pq = log p q. This results in the following approximation:
[0084]
[0085] The gauge convolution sums over messages of the form K(v pq )ρ in (r p←q )f(p). Thus, the feature vectors f(q) of the neighbors q are transformed in a way that can depend on, for example: i) the intrinsic geometry of the manifold via r p←q and v pq , and ii) a non-isotropic (but gauge-equivariant) learnable kernel K(v pq ).
[0086] The discrete gauge convolution can be computed in several procedural steps, some of which can be made during precomputation and others during the forward pass: i) compute the logarithmic map v pq = log p q:, ii) compute the parallel transporter r p←q , iii) construction / parameterization of the kernel, and iv) linear contraction of the kernel and the signal.
[0087] Computing the logarithmic map and the parallel transporter on a general manifold or mesh can be complex. Additionally, since the actual geometry of M = S 2 is known (not just a discrete approximation), the accuracy of the computed logarithmic map and transporter is not affected by the mesh type or resolution as it would be in the case where the mesh only approximates a sphere.
[0088] Note that since r p←q is a planar rotation, it can be determined by the position where it sends a single (non-zero) vector. The first basis vector can be expressed in 3D Euclidean coordinates. The first basis vector is rotated by the angle ∠(p, q) = arccos <p, q> between p and q about an axis orthogonal to the pq plane. The resulting vector lies in the tangent plane at p. Subsequently, r p←qThis is the angle between the vector and T p S 2 and the first basis vector in .
[0089] In some aspects, for each point p in the grid V and each the transfer angle can be pre-computed. This results in an array of angles of size num_v×num_neigh, where num_v = and is the maximum neighborhood size. Nodes with a non-maximal number of neighbors can be padded with zeros.
[0090] For each and compute v pq = log p q, which is the vector in T p S 2 pointing in the direction of q and having a length equal to the geodesic distance between p and q. One way to compute the logarithmic map is to project the 3D Euclidean difference vector q - p onto the tangent plane at p. This results in a vector with the correct direction. Then, the length of can be scaled to match the geodesic distance d(p, q) (which can be referred to as the arc length):
[0091]
[0092] The resulting v pq = log p q can be expressed in polar coordinates. This provides two arrays log_map_r (the length / radius coordinate of v) and log_map_angle (the angular part of v, relative to the gauge at p). These two arrays ((log_map_r and log_map_angle) are shaped as num_v×num_neigh as before. Since the geometry and the grid are fixed, these arrays are only computed once before training.
[0093] The kernel K(v) can be defined as a continuous matrix-valued function in satisfying the kernel constraints (see Equation 2). In a classical CNN, operating on a uniform pixel grid in a small (e.g., 3×3) set of neighboring pixels can be defined as (i) such that the kernel can be evaluated at a small number (e.g., 9) of points v out ×C in ×3×3 learnable coefficients.
[0094] On a sphere, there may not be a perfectly uniform grid, so depending on the convolution Points to be evaluated The neighborhood structure N(p) can be different. Therefore, the points at which K is evaluated can also be different. For this reason, K can be parameterized as a linear combination of analytically determined continuous basis kernels. The linear coefficients of K can be learned.
[0095] Assume that ρ_in and ρ_out are block diagonal, with blocks being irreducible representations (irreps). Then any SO(2) representation can be put into this form by a change of basis. In this case, the kernel can also have a block structure, with each block corresponding to a specific input / output irreducible representation, having irreducible representations labeled by integer frequencies n ≥ 0. The full kernel can be constructed block by block, where both the input and output representations are single irreducible representations.
[0096] The analytical solution of Equation 2 can be split into independent radial and angular parts. For the kernel mapping from ρ n to ρ m , the solution for the angular part K(θ) is shown in Table 1, while the radial part is unconstrained. In Table 1, c ± = cos(m ± n)θ, s ± = sin(n ± n)θ. Correspondingly, if a set of radial functions {R a (r)} is chosen, and {K b (θ)} is the complete set of angular solutions, then for weights w, the parameterized kernel is: This solution can be expressed as K i , such that the parameterized kernel is ∑ i w i K i . The number of basis kernels is called num_basis.
[0097]
[0098] Table 1
[0099] Since the geometry and the grid are fixed, the basis kernels evaluated at all points can be precomputed. That is, for each and each basis kernel contracted with the input representation K i (v pq )ρ in (r p←q ) can be evaluated. The result of this precomputation is an array of shape num_basis × num_v × num_neigh × c_out × c_in, where c_in and c_out are the dimensions of p in and ρ out and also the number of channels of the input and output signals.
[0100] After the basis kernel is computed at each v pq the discretized gauge convolution (see Equation 3) can be computed as a linear contraction. In doing so, the signal f(p) of shape num_v×c_in can be extended to a signal of shape num_v×num_neigh×c_in Such that is the signal value at the q-th neighbor of p.
[0101] Subsequently, the signal with the basis kernel K i (v pq )ρ in (r p←q ) and weights w i can be contracted to obtain the convolution result ψ*f of shape num_v×c_out. Since the basis kernel K acts on only one in / out irreducible representation pair, it may mostly be zero. i In some respects, each layer of the network can be gauge equivariant, including non-linearities. Irreducible representation features do not commute with pointwise non-linearities. However, the basis can be transformed to a basis in which the pointwise non-linearity is approximately gauge equivariant. Thereafter, the basis can be transformed back to the irreducible representation.
[0102] For simplicity, assume that the representation is
[0103] U copies of One such copy can be regarded as the discrete Fourier modes of a circular signal band-limited to M. The inverse discrete Fourier transform (DFT) matrix can map these modes to N spatial samples. Under a gauge transformation by a multiple of 2π / N, the sampling can be circularly shifted. The resulting representation can thus be called the regular representation, and hence the procedure can be called regular non-linearity. Non-linearities acting pointwise on these samples (such as the rectified linear unit (ReLU)) commute with such gauge transformations.
[0104] One way to compute is to interpolate the sampled values at N(p) to obtain a continuous function on and then use orthogonal integration to obtain a more accurate integral value. Orthogonal integration is a general numerical technique for approximating an integral with a finite sum. For a region A and a function g, the integral ∫ A g(x)dx can be approximated by ∑ x∈l ω x g(x), where is a finite set of orthogonal points, each with a weight ω x . The goal is to select I and ω x, such that the approximation is accurate (or even exact) for functions g that satisfy certain regularity assumptions (e.g., band-limited). For example, region A can be a disc with a radius such as the support radius R of the kernel.
[0105] The signal at c ∈ I Inferred from the signal at N(p) by interpolation:
[0106]
[0107] where k(c, q) = exp(-||c - Iogp(q)|| 2 / σ 2 ) is a Gaussian kernel scaled by σ, which measures the distance between c and q in the tangent space, and is a normalization constant.
[0108] Can be calculated by orthogonal integration The integral over:
[0109]
[0110] The convolution of Equation 6 can be summed over the homogenized neighborhood and is thus more equivalent to the rotation of the sphere. For example, if a large number of orthogonality points are used, the equivariance can be improved, which may increase the computational cost. However, since the combination of linear operations is linear, it can be simplified:
[0111]
[0112]
[0113]
[0114] To obtain a new kernel This new kernel Can be pre-computed once, so that the convolution during runtime only involves the sum over the neighbors, just like the convolution in Equation 3. Thus, interpolation does not affect the computational cost.
[0115] Figure 4A Is a diagram illustrating the example icosahedron 400. Refer to Figure 4A, the icosahedron 400 is a rough approximation of a sphere. The icosahedron 400 is a convex polyhedron. The icosahedron 400 is a spherical-like Platonic solid. The icosahedron 400 has twenty flat faces 402a-t, thirty edges 404a-dd, and twelve vertices 406a-l. Points on the icosahedron 400 or its grid are at different distances from the origin (0, 0, 0) in three-dimensional Euclidean space. Due to local flatness, gauge-equivariant convolutions can be reduced to conventional two-dimensional convolutions (conv2d) with feature transport performed via simple indexing. Accordingly, many mathematical definitions associated with the icosahedron 400 (e.g., exponential mapping) are trivialized.
[0116] Figure 4B illustrates an example isogonal polyhedron mesh S according to aspects of the present disclosure 2 450. Referring to Figure 4B , the isogonal polyhedron mesh 450 is a specific sampling of this Platonic solid of the icosahedron 400. Points or positions on the triangular faces (e.g., 454a-n) can be selected and connected such that they cover the object like a net. The mesh can sample or discretize a continuous signal on the sphere.
[0117] Figure 5A illustrates a diagram 500 of an exponential mapping from a sphere or spherical manifold 502 to a tangent plane 504 for gauge-invariant transformation according to aspects of the present disclosure. Referring to Figure 5A , a point p on the spherical manifold 502 is projected onto the tangent plane 504. The linear mapping or gauge w p is defined as Using the gauge w p , the exponential mapping takes a tangent vector V ∈ T p M, and starting from the point p, travels along the geodesic 506 at a speed of one unit of time ||V|| to reach the point q on the spherical manifold 502 v = exp p at V ∈ M.
[0118] Figure 5B illustrates a tangent plane 550 showing a point of interest (p) 552 and interpolation points 554 (e.g., 554a-f) on the tangent plane 550. For each point of interest 552, the set of points is in the same position relative to its corresponding point of interest. Signals from the neighbors p (on the isogonal polyhedron sphere) are interpolated to the interpolation points 554a - 554f, thereby appropriately steering the gauge. A convolution (e.g., the convolution operation proposed in Equation 1) can be performed on these interpolation points 554a - 554f. In one aspect, the interpolation is performed as a pre-computation step such that the number of interpolations does not affect the training time. Thus, during runtime, the summation is only for the neighbors, rather than all interpolation points.
[0119] However, due to different interpolation weights and different gauge turns, the convolution operation does not treat the neighbors of one vertex (e.g., the first focus point p) the same as the neighbors of another vertex (e.g., the focus point p1 (not shown)). The difference in the treatment of different neighbors is in contrast to both planar CNNs and isopolyhedral CNNs (e.g., the convolution operation proposed in Equation 1), which can use a single kernel and apply a conventional single two-dimensional convolution (conv2d).
[0120] The neighborhood expansion implementation improves the convolution implementation. The neighborhood expansion implementation starts with the signal vector f(p) at each vertex p. For each focus point p, up to M neighbors q are assigned. The signal vector f(p) is indexed to form a tensor f(p, q) for q = 0...M. For example, the tensor f(p, 3) is the signal at the third neighbor of vertex p. Subsequently, the convolution operation proposed in Equation 1 can be applied to the tensor f(p, q). The execution order of some operations of the neighborhood expansion implementation can be adjusted to reduce the additional memory requirements compared to convolution on a regular grid. For example, the efficiency of the neighborhood expansion implementation can be improved by focusing on the non-zero blocks of the resulting matrix of the tensor. That is, the non-zero blocks can be applied to the relevant input and output vectors.
[0121] Aspects of the present disclosure are more robust (approximately equivariant) for any group action of SO(3) (e.g., the 3D rotation group), faster, and more scalable than other implementations (e.g., implementations based on Fourier domain operations, which are computationally restrictive).
[0122] Figure 6 is a diagram illustrating a method 600 for generating a convolutional neural network operating on a spherical manifold according to aspects of the present disclosure. As Figure 6 shown, at block 602, locally defined gauges are generated at multiple positions on the spherical manifold. The locally defined gauges correspond to the tangent planes and their corresponding focus positions. For example, referring to Figure 5A , the linear map or gauge w p is defined as Using the gauge w p , the exponential map takes a tangent vector V ∈ T p M, and starting from point p, travels along the geodesic 506 to point q on the spherical manifold 502 at a speed of one unit of time ||V|| v = exp p at V ∈ M.
[0123] At block 604, a convolution is computed at each of the multiple positions on the spherical manifold with reference to the locally defined gauges. For example, as described with reference to Equation 1, the convolution can be computed by to each neighboring point q v, the eigenvector f(q v ) is defined by being transported to p. In some aspects, the locally defined gauge can be arbitrarily selected. Additionally, the locally defined gauge can be defined differently at different positions on the manifold.
[0124] At block 606, the result of the convolution is transformed at each position based on gauge equivariance to obtain the corresponding manifold transformation. As described with reference to Equation 1, the convolution is defined such that the resulting feature at point p can use the learned kernel to perform a transformation and integrate the result over the support of K in .
[0125] The various operations of the methods described above can be performed by any suitable device capable of performing the corresponding functions. These devices can include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, in cases where there are operations illustrated in the figures, those operations can have corresponding paired device plus functional components with similar numbers.
[0126] As used, the term "determine" encompasses a variety of actions. For example, "determine" can include calculating, computing, processing, deriving, researching, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and the like. Additionally, "determine" can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and similar actions. Furthermore, "determine" can include parsing, selecting, choosing, establishing, and similar actions.
[0127] As used, the phrase "at least one of" recited in a list of items refers to any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c.
[0128] The various illustrative logical blocks, modules, and circuits described in connection with the present disclosure can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array signal (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the described functions. The general purpose processor can be a microprocessor, but in an alternative, the processor can be any commercially available processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0129] The steps of the methods or algorithms described in connection with the present disclosure may be implemented directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules may reside in any form of storage medium known in the art. Some examples of storage media that may be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, hard disks, removable disks, CD-ROMs, and the like. The software modules may include a single instruction, or many instructions, and may be distributed over several different code segments, distributed among different programs, and across multiple storage media. The storage medium may be coupled to the processor such that the processor can read from, and write to, the storage medium. In an alternative, the storage medium may be integrated into the processor.
[0130] The disclosed methods include one or more steps or acts for achieving the described methods. These method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the specific order and / or use of the steps and / or acts may be altered without departing from the scope of the claims.
[0131] The described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration may include a processing system in a device. The processing system may be implemented with a bus architecture. Depending on the particular application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect, among other things, a network adapter to the processing system via the bus. The network adapter may be used to implement signal processing functions. For certain aspects, a user interface (such as a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits such as a timing source, peripherals, voltage regulators, power management circuits, and similar circuits that are well known in the art and will not be described further herein.
[0132] The processor may be responsible for managing the bus and general processing, including executing software stored on a machine-readable medium. The processor may be implemented with one or more general-purpose and / or special-purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry capable of executing software. Software should be construed broadly to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. By way of example, the machine-readable medium may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, or any other suitable storage medium, or any combination thereof. The machine-readable medium may be embodied in a computer program product. The computer program product may include packaging materials.
[0133] In a hardware implementation, the machine-readable medium may be a part of the processing system separate from the processor. However, as will be readily appreciated by those skilled in the art, the machine-readable medium or any part thereof may be external to the processing system. By way of example, the machine-readable medium may include transmission lines, carrier waves modulated with data, and / or computer products separate from the device, all of which may be accessed by the processor via a bus interface. Alternatively or additionally, the machine-readable medium or any part thereof may be integrated into the processor, such as may be the case with a cache and / or a general register file. Although the various components discussed may be described as having a particular location, such as local components, they may also be configured in various ways, such as some components being configured as part of a distributed computing system.
[0134] The processing system may be configured as a general-purpose processing system having one or more microprocessors providing processor functionality and an external memory providing at least a portion of the machine-readable medium, all linked together via an external bus architecture with other support circuitry. Alternatively, the processing system may include one or more neuromorphic processors for implementing the described neuron models and nervous system models. As another alternative, the processing system may be implemented with an application-specific integrated circuit (ASIC) with a processor, bus interface, user interface, support circuitry, and at least a portion of the machine-readable medium integrated on a single chip, or with one or more field-programmable gate arrays (FPGA), programmable logic devices (PLD), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuits capable of performing the various functions described throughout this disclosure. Depending on the particular application and the overall design constraints imposed on the system, those skilled in the art will recognize how best to implement the functionality described with respect to the processing system.
[0135] A machine-readable medium may include several software modules. These software modules include instructions that cause a processing system to perform various functions when executed by a processor. These software modules may include a transmission module and a reception module. Each software module may reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, the software module may be loaded from a hard drive into RAM. During the execution of the software module, the processor may load some instructions into the cache to improve access speed. One or more cache lines may then be loaded into the general register file for the processor to execute. When referring to the functionality of the software module below, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module. In addition, it should be appreciated that aspects of the present disclosure result in improvements to the functionality of a processor, computer, machine, or other system that implements such aspects.
[0136] If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared (IR), radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, and microwave is included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk often magnetically reproduces data, while disc optically reproduces data with a laser. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transitory computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0137] Accordingly, some aspects may include a computer program product for performing a given operation. For example, such a computer program product may include a computer-readable medium having (and / or encoded with) instructions that, when executed by one or more processors, perform the described operations. For some aspects, the computer program product may include packaging material.
[0138] In addition, it should be appreciated that modules and / or other suitable means for performing the described methods and techniques may be downloaded and / or otherwise obtained by a user terminal and / or a base station where applicable. For example, such devices can be coupled to a server to facilitate the transfer of means for performing the described methods. Alternatively, the various methods can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk, etc.) such that once the storage device is coupled to or provided to the user terminal and / or the base station, the device can obtain the various methods. In addition, any other suitable technique for providing the described methods and techniques to a device can be utilized.
[0139] It will be understood that the claims are not limited to the exact configurations and components described above. Various changes, substitutions, and modifications can be made in the layout, operation, and details of the methods and apparatuses described above without departing from the scope of the claims.
Claims
1. A method, comprising: receiving, by an artificial neural network (ANN), an input, the input including a spherical signal on a spherical manifold, the spherical signal being generated by an imaging device; generating, at a plurality of positions on the spherical manifold, a plurality of locally defined gauges; calculating, by the ANN, at each of the plurality of positions on the spherical manifold, a gauge-equivariant convolution for the input with reference to a locally defined gauge among the plurality of locally defined gauges, the plurality of locally defined gauges assigning a linear mapping from a standard plane to a position on a tangent plane of the spherical manifold for each of the plurality of positions, and the gauge-equivariant convolution being based on a learned kernel constraint for determining a transformation equivalent to a rotation of the spherical manifold; and performing, by the ANN, image recognition by applying the gauge-equivariant convolution to the input to identify and classify an object on the spherical manifold.
2. The method of claim 1, wherein for each focus position, a set of interpolation positions is included in each tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
3. The method of claim 1, further comprising interpolating signals from adjacent positions of a focus position on the spherical manifold to adjacent interpolation points on a tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
4. The method of claim 3, further comprising defining the gauge-equivariant convolution on the adjacent interpolation points.
5. The method of claim 3, further comprising: indexing a signal vector of the focus position to generate a tensor associated with adjacent positions; and performing a convolution operation on the tensor by applying a non-zero block of a resulting matrix of the tensor.
6. The method of claim 1, further comprising parameterizing the spherical manifold as one of a predefined array of shapes.
7. The method of claim 1, further comprising distributing a corresponding manifold transformation to a local neighborhood of the spherical manifold based on the plurality of locally defined gauges.
8. The method of claim 1, wherein each kernel associated with each of the plurality of positions is a locally varying kernel derived from the same function.
9. The method of claim 1, wherein each gauge-equivariant convolution calculated at each position is calculated using a locally connected layer.
10. An apparatus, comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: receive, by an artificial neural network (ANN), an input, the input including a spherical signal on a spherical manifold, the spherical signal being generated by an imaging device; generate, at a plurality of positions on the spherical manifold, a plurality of locally defined gauges; The gauge-equivariant convolution for the input is calculated by the ANN at each of the plurality of positions on the spherical manifold with reference to a locally defined gauge among the plurality of locally defined gauges, the plurality of locally defined gauges assigning a linear mapping from a standard plane to a position on the tangent plane of the spherical manifold, and the gauge-equivariant convolution is based on a learned kernel constraint for determining a transformation equivalent to a rotation of the spherical manifold; and The ANN performs image recognition by applying the gauge-equivariant convolution to the input to identify and classify an object on the spherical manifold.
11. The apparatus according to claim 10, wherein for each focus position, a set of interpolation positions is included in each tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
12. The apparatus according to claim 10, wherein the at least one processor is further configured to interpolate signals from adjacent positions of a focus position on the spherical manifold to adjacent interpolation points on the tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
13. The apparatus according to claim 12, wherein the at least one processor is further configured to define the gauge-equivariant convolution on the adjacent interpolation points.
14. The apparatus according to claim 12, wherein the at least one processor is further configured to: index a signal vector of the focus position to generate a tensor associated with adjacent positions; and perform a convolution operation on the tensor by applying non-zero blocks of the resulting matrix of the tensor.
15. The apparatus according to claim 10, wherein the at least one processor is further configured to parameterize the spherical manifold as one of a predefined array of shapes.
16. The apparatus according to claim 10, wherein the at least one processor is further configured to distribute a corresponding manifold transformation to a local neighborhood of the spherical manifold based on the plurality of locally defined gauges.
17. The apparatus according to claim 10, wherein each kernel associated with each of the plurality of positions is a locally varying kernel derived from the same function.
18. The apparatus according to claim 10, wherein each gauge-equivariant convolution defined at each position is calculated with a locally connected layer.
19. An apparatus comprising: means for receiving, by an artificial neural network (ANN), an input including a spherical signal on a spherical manifold, the spherical signal being generated by an imaging device; means for generating a plurality of locally defined gauges at a plurality of positions on the spherical manifold; Apparatus for calculating gauge-equivariant convolutions for the input with respect to a locally defined gauge among the plurality of locally defined gauges at each of the plurality of positions on the spherical manifold by the ANN, the plurality of locally defined gauges assigning a linear mapping from a standard plane to a position on the tangent plane of the spherical manifold to each of the plurality of positions, and the gauge-equivariant convolution being based on learned kernel constraints for determining a transformation equivalent to a rotation of the spherical manifold; and Apparatus for performing image recognition by the ANN to identify and classify objects on the spherical manifold by applying the gauge-equivariant convolution to the input.
20. The apparatus of claim 19, wherein for each focus position, a set of interpolation positions is included in each tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
21. The apparatus of claim 19, further comprising apparatus for interpolating signals from adjacent positions of a focus position on the spherical manifold to adjacent interpolation points on the tangent plane corresponding to the plurality of locally defined gauges and their corresponding focus positions.
22. The apparatus of claim 21, further comprising apparatus for calculating the gauge-equivariant convolution on the adjacent interpolation points.
23. The apparatus of claim 21, further comprising: Apparatus for indexing a signal vector of the focus position to generate a tensor associated with adjacent positions; and Apparatus for performing a convolution operation on the tensor by applying non-zero blocks of the resulting matrix of the tensor.
24. The apparatus of claim 19, further comprising apparatus for parameterizing the spherical manifold into one of a predefined array of shapes.
25. The apparatus of claim 19, further comprising apparatus for distributing a corresponding manifold transformation to a local neighborhood of the spherical manifold based on the plurality of locally defined gauges.
26. A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and comprising: Program code for receiving an input by an artificial neural network (ANN), the input including a spherical signal on a spherical manifold, the spherical signal being generated by an imaging device; Program code for generating a plurality of locally defined gauges at a plurality of positions on the spherical manifold; Program code for calculating gauge-equivariant convolutions for the input with respect to a locally defined gauge among the plurality of locally defined gauges at each of the plurality of positions on the spherical manifold by the ANN, the plurality of locally defined gauges assigning a linear mapping from a standard plane to a position on the tangent plane of the spherical manifold to each of the plurality of positions, and the gauge-equivariant convolution being based on learned kernel constraints for determining a transformation equivalent to a rotation of the spherical manifold; and Program code for performing image recognition by the ANN to identify and classify objects on the spherical manifold by applying the gauge-equivariant convolution to the input.
27. The non-transitory computer-readable medium according to claim 26, wherein for each location of interest, a set of interpolation locations is included in each tangent plane corresponding to the plurality of locally defined gauges and their corresponding locations of interest.
28. The non-transitory computer-readable medium according to claim 26, wherein the program code further includes program code for interpolating signals from adjacent locations of interest on the spherical manifold to adjacent interpolation points on the tangent plane corresponding to the plurality of locally defined gauges and their corresponding locations of interest.