Identifying multi-scale features using neural networks
Spatially adaptive separable convolutional layers optimize filter usage and parallel processing to address the computational complexity of convolutional neural networks, enhancing feature identification efficiency and reducing resource demands.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-03-13
AI Technical Summary
The complexity and resource requirements of convolutional neural networks increase significantly with the identification of multiple features in input data, leading to increased memory, time, and computational demands due to the need for multiple filters across multiple layers.
The implementation of spatially adaptive separable convolutional layers that apply multiple filters of varying sizes to different layers, reducing the computational load by parallel processing and optimizing filter usage through backpropagation.
This approach significantly reduces the computational and memory requirements while maintaining feature identification efficiency, allowing for faster and more effective processing of complex features in input data.
Smart Images

Figure 0007829647000011 
Figure 0007829647000012 
Figure 0007829647000013
Abstract
Description
[Technical Field]
[0001] At least one embodiment relates to an improved convolutional layer within a convolutional neural network used to facilitate and implement artificial intelligence. For example, at least one embodiment relates to a processor and computing system used to identify multiple features in parallel within a convolutional layer by various novel techniques described herein. [Overview of the project]
[0002] The complexity of a convolutional neural network can increase based on the fact that features are identified within one or more input data items, as a result of performing convolutional layer steps to generate feature maps within the network. The amount of memory, time, and computational resources required to identify a particular feature in the data increases with each feature to be identified. Resource and computational complexity generally increase due to each filter within the convolutional layer required to identify each individual feature. These resource and computational complexity increase further when the input data contains multiple individual information layers. In particular, for each information layer in the input data item set, the convolutional layers in the convolutional neural network must apply individual filters of a specific size to each layer in the input data item set for each feature to be identified. The operation of applying filters to each layer in the input data item set generates individual outputs for each filter applied, such as within separable convolutional layers, so point-by-point operations must be performed to aggregate the various filter information for each input data item into an output feature map. As more complex features are identified in the data set, the filter count and associated resource requirements increase significantly. [Brief explanation of the drawing]
[0003] [Figure 1] This figure shows an exemplary convolutional neural network according to at least one embodiment. [Figure 2A]A diagram showing vertical line features within an input data item according to at least one embodiment. [Figure 2B] A diagram showing horizontal line features within an input data item according to at least one embodiment. [Figure 2C] A diagram showing multi-scale features within an input data item according to at least one embodiment. [Figure 3] A diagram showing convolution of depth units in a convolutional layer of a convolutional neural network according to at least one embodiment. [Figure 4] A diagram showing convolution of point units in a convolutional layer of a convolutional neural network according to at least one embodiment. [Figure 5] A diagram showing a separable convolutional layer within a convolutional neural network according to at least one embodiment. [Figure 6] A diagram showing an architecture for performing convolution of depth units in a spatially adaptive separable convolutional layer of a convolutional neural network according to at least one embodiment. [Figure 7] A diagram showing a spatially adaptive separable convolutional layer within a convolutional neural network according to at least one embodiment. [Figure 8] A diagram showing a process for generating a feature map within a spatially adaptive separable convolutional layer according to at least one embodiment. [Figure 9A] A diagram showing inference and / or training logic according to at least one embodiment. [Figure 9B] A diagram showing inference and / or training logic according to at least one embodiment. [Figure 10] A diagram showing training and deployment of a neural network according to at least one embodiment. [Figure 11] A diagram showing an exemplary data center system according to at least one embodiment. [Figure 12A] A diagram showing an example of an autonomous vehicle according to at least one embodiment. [Figure 12B]This figure shows an example of the camera location and field of view of the autonomous vehicle shown in Figure 12A, according to at least one embodiment. [Figure 12C] This is a block diagram illustrating an exemplary system architecture of an autonomous vehicle in at least one embodiment, as shown in Figure 12A. [Figure 12D] This figure shows a system for communication between a cloud-based server and the autonomous vehicle shown in Figure 12A, according to at least one embodiment. [Figure 13] A block diagram of a computer system according to at least one embodiment. [Figure 14] A block diagram of a computer system according to at least one embodiment. [Figure 15] This figure shows a computer system according to at least one embodiment. [Figure 16] This figure shows a computer system according to at least one embodiment. [Figure 17A] This figure shows a computer system according to at least one embodiment. [Figure 17B] This figure shows a computer system according to at least one embodiment. [Figure 17C] This figure shows a computer system according to at least one embodiment. [Figure 17D] This figure shows a computer system according to at least one embodiment. [Figure 17E] This figure shows a shared programming model according to at least one embodiment. [Figure 17F] This figure shows a shared programming model according to at least one embodiment. [Figure 18] This figure shows an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 19A] This figure shows an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 19B]This figure shows an exemplary integrated circuit and associated graphics processor according to at least one embodiment. [Figure 20A] This figure shows additional exemplary graphics processor logic according to at least one embodiment. [Figure 20B] This figure shows additional exemplary graphics processor logic according to at least one embodiment. [Figure 21] This figure shows a computer system according to at least one embodiment. [Figure 22A] This figure shows a parallel processor according to at least one embodiment. [Figure 22B] This figure shows a partition unit according to at least one embodiment. [Figure 22C] This figure shows a processing cluster according to at least one embodiment. [Figure 22D] This figure shows a graphics multiprocessor according to at least one embodiment. [Figure 23] This figure shows a multi-graphics processing unit (GPU) system according to at least one embodiment. [Figure 24] This figure shows a graphics processor according to at least one embodiment. [Figure 25] A block diagram showing a processor microarchitecture for a processor, according to at least one embodiment. [Figure 26] This figure shows a deep learning application processor according to at least one embodiment. [Figure 27] A block diagram showing an exemplary neuromorphic processor according to at least one embodiment. [Figure 28] This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 29] This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 30]This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 31] This is a block diagram of the graphics processing engine 3110 of a graphics processor according to at least one embodiment. [Figure 32] This is a block diagram of at least a portion of a graphics processor core according to at least one embodiment. [Figure 33A] This figure shows a thread execution logic 3300 including an array of processing elements of a graphics processor core, according to at least one embodiment. [Figure 33B] This figure shows a thread execution logic 3300 including an array of processing elements of a graphics processor core, according to at least one embodiment. [Figure 34] This figure shows a parallel processing unit ("PPU") according to at least one embodiment. [Figure 35] This figure shows a general-purpose processing cluster ("GPC") according to at least one embodiment. [Figure 36] This figure shows a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment. [Figure 37] This figure shows a streaming multiprocessor according to at least one embodiment. [Modes for carrying out the invention]
[0004] Figure 1 shows a convolutional neural network comprising several convolutional layers 104, 112 that are responsible for identifying features contained in or displayed in one or more input data items 102, according to at least one embodiment. In at least one embodiment, this convolutional neural network assists computer vision applications using one or more input data items 102, such as an input image with several layers representing color channels and other information, if available.
[0005] In at least one embodiment, the convolutional neural network includes multiple layers, including those for feature learning 104, 106, 108, 110, 112, 114, 116, 118 and those for classification 120, 122. In at least one embodiment, the convolutional neural network takes in an input image including multiple layers and assigns importance to each aspect, feature, or object in each input image during the feature learning stages 104, 106, 108, 110, 112, 114, 116, 118. In at least one embodiment, the classification stage takes in information about aspects, features, or objects in one or more input images and uses the trained neural network to classify those aspects, features, or objects in the one or more input images.
[0006] In at least one embodiment, the system trains a convolutional neural network to generate feature maps 106, 110, 114, 118 containing information about aspects, features, or objects in an input image during feature learning stages 104, 106, 108, 110, 112, 114, 116, 118. In at least one embodiment, the feature learning stages 104, 106, 108, 110, 112, 114, 116, 118 include a plurality of steps in which convolutional layers 104, 112 are applied to a two-dimensional data matrix such as a feature map.
[0007] In at least one embodiment, one or more convolutional layers form feature learning stages 104, 106, 108, 110, 112, 114, 116, and 118. In at least one embodiment, one or more convolutional layers include a set of filters (also called kernels, feature extractors, and matrices), each filter applied to each layer in the image across the data in the image with respect to the width and height of the image. In at least one embodiment, convolutional layers 104 and 112 take as input multiple feature maps 106 and 114 representing data layers, such as the red, green, and blue layers of an RGB image. In at least one embodiment, the feature maps 106 and 114 are two-dimensional matrices of values associated with images, each value may represent an aspect of one or more input images 102. In at least one embodiment, the aspects of one or more input images 102 represented in the feature map may be probabilities associated with the location in the image 102 or the likelihood of the aspect existing in a layer of the image.
[0008] In at least one embodiment, the filter is a submatrix. In at least one embodiment, the filter may be used to identify features of an image and to perform blurring, sharpening, embossing, edge detection, or other image-related filtering operations. In at least one embodiment, the filter or kernel is applied to the image through a convolution operation, which may consist of performing a dot product operation across all dimensions of the input image. In at least one embodiment, the convolutional layers 104, 112 may include one or more filters.
[0009] In at least one embodiment, the convolutional layers 104, 112 may include depth-based and point-based convolutional stages. In at least one embodiment, the depth-based convolution applies one or more filters across the width and height of the input image 102 or feature maps 106, 114. In at least one embodiment, the depth-based convolution applies one or more filters or kernels across a subset of the width and height of the image 102 or feature maps 106, 114. In at least one embodiment, the application of one or more filters or kernels can be used to identify features within the image 102 or feature maps 106, 114. In at least one embodiment, the set of feature maps output from the depth-based convolution contains information about aspects, features, or objects within the input set of image 102 or feature map 114. In at least one embodiment, the point-based convolution in the convolutional layer applies a pixel-by-pixel convolution to each feature map generated by the depth-based convolution. In at least one embodiment, point-level convolution combines the information contained within each feature map generated by depth-level convolution, as further described below.
[0010] In at least one embodiment, the convolutional layers 104, 112 generate one or more feature maps 106, 114. In at least one embodiment, one or more pooling layers 108, 116 reduce the size of the feature maps 106, 114 after the convolutional layers 104, 112. In at least one embodiment, one or more pooling layers 108, 116 reduce the size of the feature maps to reduce the computational power requirements for subsequent convolutional layers 104, 112 or other operations 120 within the convolutional neural network.
[0011] In at least one embodiment, one or more pooling layers 108, 116 may be either a maximum value pooling layer or an average value pooling layer. In at least one embodiment, the maximum value pooling layer returns the maximum value from a portion of the input image or feature map processed by the kernel in the convolutional layers 104, 112. In at least one embodiment, the average value pooling layer returns the average of all values from a portion of the input image or feature map processed by the kernel in the convolutional layers 104, 112.
[0012] In at least one embodiment, a classification stage 120 is performed to generate an output 122 containing the respective probabilities for each classification, after a predetermined number of convolutional layers 104, 112 and subsequent pooling layers 108, 116 have extracted aspects, features, or objects from an input image or feature map. In at least one embodiment, a fully connected layer 120 can be used to learn a nonlinear combination of high-level features identified by feature learning stages 104, 106, 108, 110, 112, 114, 116, 118. In at least one embodiment, the fully connected layer 120 may include several steps. In at least one embodiment, the fully connected layer 120 can flatten the input set of feature maps into a single-dimensional vector. In at least one embodiment, additional steps may be performed in 120. In at least one embodiment, the flattened representation of the data is fed into a feedforward neural network, as described later, and the neural network is trained through iterative training using backpropagation. In at least one embodiment, after a predetermined number of iterations or epochs, the feedforward neural network is able to distinguish between features initially specified in the filter or kernel between the preceding convolutional layers 104, 112. In at least one embodiment, the output is classified using a softmax classification technique.
[0013] Figure 2A shows a vertical line feature @102@02 in an input data item, such as an image. In at least one embodiment, the convolutional layer may take one or more images as input. In at least one embodiment, one or more images may contain aspects, features, or objects @102@02 that will be identified by the convolutional layer. In at least one embodiment, the vertical line @102@02 may be an aspect, feature, or object that can be identified by a filter or kernel in the convolutional layer. In at least one embodiment, one or more filters in the convolutional layer may be responsible for identifying a single feature, such as the vertical line @102@02, in the input image.
[0014] Figure 2B shows a horizontal line feature @102@04 in an input data item, such as an image. In at least one embodiment, the convolutional layer may take one or more images as input, and one or more images may contain aspects, features, or objects @102@04 that will be identified or highlighted in the feature map by the convolutional layer. In at least one embodiment, the horizontal line @102@04 may be an aspect, feature, or object that can be identified, extracted, or highlighted in the feature map by a filter or kernel in the convolutional layer. In at least one embodiment, one or more filters in the convolutional layer may be responsible for identifying a single feature, such as the horizontal line @102@04, in the input image.
[0015] Figure 2C shows a multi-scale feature @102@06 in an input data item, such as a variable-size object. In at least one embodiment, a convolutional layer may take one or more images as input, and one or more images may contain one or more variable-size aspects, features, or objects @102@06 that will be identified or highlighted in a feature map by a convolutional layer containing one or more filters or kernels. In at least one embodiment, a multi-scale feature @102@06 may contain one or more aspects, features, or objects that can be identified, extracted, or highlighted in a feature map by filters or kernels in the convolutional layer. In at least one embodiment, one or more filters in the convolutional layer may be responsible for identifying a single feature in the multi-scale @102@06 in the input image. In at least one embodiment, multiple filters can be position-aligned to each location of each multi-scale feature @102@06 in a feature map or other input data item. In at least one embodiment, the filter can be position-aligned by padding or linearly scaling filters of various sizes.
[0016] Figure 3 shows a depth-based convolution in a convolutional layer of a convolutional neural network. In at least one embodiment, an input data item 302, such as an image, may include multiple layers. In at least one embodiment, the layers within the input data image 302 may include color information, such as the red, green, and blue layers in an RGB input image 302. In at least one embodiment, the depth-based convolution in a convolutional layer of a convolutional neural network can first separate each layer 304 into individual layers 306, 308, 310, which are represented as two-dimensional matrices or feature maps as described above. In at least one embodiment, each separated layer or feature map 306, 308, 310 may represent a subset of information about the input data item or image 302.
[0017] In at least one embodiment, each separated layer or feature map 306, 308, 310 is then filtered or processed through convolution. In at least one embodiment, the separated layers or feature maps 306, 308, 310 may be two-dimensional matrices having dimensions K × K. In at least one embodiment, the input data item or image 302 is C in It may have layers or feature maps 306, 308, 310. In at least one embodiment, the convolution in units of depth in the convolutional layer of the convolutional neural network is C out The number of layers or feature maps 322 may be output.
[0018] In at least one embodiment, a conventional convolutional layer or separable convolutional layer shares a single K×K filter across all layers or feature maps 306, 308, and 310 during the convolution step 312. In at least one embodiment, a spatially adaptive separable convolutional layer can apply multiple filters or kernels of varying sizes, not limited to K×K, to different layers or feature maps 306, 308, and 310 during the convolution step 312, as described later. In at least one embodiment, a spatially adaptive separable convolutional layer can apply multiple filters or kernels of size K×K to different layers or feature maps 306, 308, and 310 during the convolution step 312, as described later, with each filter corresponding to a different feature.
[0019] In at least one embodiment, each K×K filter, or each filter of smaller dimensions, is applied to each input layer or feature map 306, 308, 310 in the convolution step 312. In at least one embodiment, the convolution step 312 may include calculating the dot product between the input matrix and the filter. In at least one embodiment, the output feature maps 314, 316, 318 may be two-dimensional matrices of dimensions K×K representing the dot product between the filter or kernel and the input layer or feature map 306, 308, 310. In at least one embodiment, the output feature maps 314, 316, 318 may be two-dimensional matrices of variable dimensions representing the dot product between filters of smaller sizes than K×K and a subset of values in the two-dimensional matrices representing the input layer or feature map 306, 308, 310. In at least one embodiment, the output feature maps 314, 316, 318 can be combined into a multilayer output 322 so that point-by-point convolution can be applied, as described later.
[0020] Figure 4 shows point-level convolution in a convolutional layer of a convolutional neural network. In at least one embodiment, the input feature map 402 can be generated as an output from depth-level convolution in the convolutional layer, as described later. In at least one embodiment, the input feature map 402 may be a two-dimensional matrix containing information about the layer in the image after a filter or kernel has been applied during the depth-level convolution in the convolutional layer for each layer. In at least one embodiment, the input feature map 402 may have dimensions of K × K within a separable convolutional layer. In at least one embodiment, the input feature map 402 may have variable dimensions within a spatially adaptive separable convolutional layer.
[0021] In at least one embodiment, a point-by-point convolution operation 404 in the convolution layer applies a convolution operation to the input feature map 402 generated by the depth-by-point convolution, either pixel-by-pixel or matrix-by-matrix element-by-matrix. In at least one embodiment, the point-by-point convolution 404 combines the spatial information contained within the input feature map 402 generated by the depth-by-point convolution to generate an output feature map 406.
[0022] Figure 5 shows a separable convolutional layer within a convolutional neural network. In at least one embodiment, the separable convolutional layer is a factorized convolutional layer consisting of depth-unit convolutions 508 as described above and 1x1 point-unit convolutions 522 as described above. In at least one embodiment, the factorized convolutional layer performs convolutions simultaneously across input channels 502, 504, and 506.
[0023] In at least one embodiment, the depth-based convolution 508 applies one or more K×K filters or kernels per input feature map to the K×K feature map U p,q We obtain spatial information within, where p,q ∈ [1..Cin]510, 512, 514, 516, 518, 520. In at least one embodiment, a 1×1 point-by-point convolution 522 is U p,q The spatial information in 510, 512, 514, 516, 518, and 520 is combined to generate output feature maps 524, 526, and 528. In at least one embodiment, multiple filters may be applied to each input image layer or feature map 502, 504, and 506 during the depth-based convolution 508. Figure 5 shows the depth-based convolution 508, with a K×K filter or kernel (where M, the number of filters or kernels applied, is equal to 2).
[0024] In at least one embodiment, a separable convolutional layer, as shown in Figure 5, receives input image layers or feature maps 502, 504, 506 as input C inReceived as, each input image layer or feature map 502, 504, 506 has a width W and a height H. In at least one embodiment, the activation size can describe the number of calculations to be performed in a separable convolutional layer. In at least one embodiment, the activation size or calculation of a separable convolutional layer is as follows. W×H×K×K×C in +W×H×C in ×C out
[0025] In at least one embodiment, a convolutional neural network as described herein can attempt to learn the values of filters or kernels to be applied to the input data 502, 504, 506 within the convolutional layer using backpropagation. In at least one embodiment, each layer within a convolutional neural network, or a convolutional layer including matrices 510, 512, 514, 516, 518, 520, 524, 526, 528 that describe weights, such as feature maps, can be considered a learnable layer or a layer including learnable elements. In at least one embodiment, the element count of the filters or kernels within a layer that can be learned is the parameter considered for the filters or kernels within the layer. In at least one embodiment, for a separable convolutional layer, the parameter count, or the elements that can be learned, is as follows. K×K×C in +C in [[ID=ID=17]]×C out
[0026] In at least one embodiment, a separable convolutional layer can include two or more K×K filters or kernels to be applied to each input feature map or image layer 502, 504, 506. In at least one embodiment, the total number of filters to be applied to each input layer or feature map 502, 504, 506 is M. FIG. 5 shows a separable convolutional layer with M = 2 according to at least one embodiment. In at least one embodiment, the activation size of a separable convolutional layer including two or more K×K filters or kernels is as follows. W×H×K×K×M×C in +W×H×C in ×M×C out
[0027] In at least one embodiment, the activation size of a separable convolutional layer containing two or more K×K filters or kernels is as follows: K×K×M×C in +C in ×M×C out
[0028] Figure 6 shows an architecture for performing depth-unit convolution in a spatially adaptive separable convolutional layer of a convolutional neural network. In at least one embodiment, the spatially adaptive separable convolutional layer allows neurons in the convolutional neural network to adaptively adjust to different locations in the input data items or image 602, as described herein. In at least one embodiment, the input values 602 may be an image layer or feature map, as described above in conjunction with the exemplary convolutional neural network.
[0029] In at least one embodiment, the “split” operation generates multiple paths or matrices 604, 606 from the input image layer or feature map 602, and the different spatial filters used to create the intermediate feature maps or matrices 604, 606 may have a variable kernel size or a uniform kernel size, but may be constructed to identify different features. In at least one embodiment, the total number of filters to be applied is M. In at least one embodiment, a depth-unit convolutional layer applies multiple filters or spatial transformations 604, 606 per input image layer or feature map 602. In at least one embodiment, the multiple filters or spatial transformations 604, 606 may be of variable size or uniform size. In at least one embodiment, there are M different spatial transformations F mc :X c →U mc ∈R W×H 604, 606, where c∈[1,...,C in], is the input feature map X∈R W×H×Cin This applies to 602. In at least one embodiment, F mc 604 and 606 are the Cth input feature map X c This is the mth filter applied to [the specified location].
[0030] In at least one embodiment, the “conflict” operation generates a compact feature descriptor 624 encoded with more information about the multi-scale features in the input 602. In at least one embodiment, the “conflict” operation uses various filters F mc Soft attention is applied across different branches or filter channels 604, 606 at each spatial location identified by the system. In at least one embodiment, a softmax operation 608 is applied to the feature map specific to each filter channel at each pixel to select a weight map S m ∈R W×H×Cin 610 and 612 are obtained, where m = 1, ..., M. In at least one embodiment, the dot product is each selected weight map S m 610, 612 and Feature Map U mc The information D was selected and performed between 604 and 606. mc 618 and 620 are generated.
[0031] In at least one embodiment, element-level addition is performed on selected information D mc Filter channel F mc By aggregating across these elements, a compact feature descriptor V c ∈R W×H Generate 624, where c∈[1,···,C in In at least one embodiment, the selected weight map value S is... i j mc 610, 612, pixel value U i j mc It can be calculated from 604 and 606, where i is a specific row, j is a specific column, c is a specific input 602, and m is a specific filter channel. In at least one embodiment, softmax can be used to determine the selected weight map values 610 and 612, which can be calculated as follows:
number
[0032] In at least one embodiment, a compact feature descriptor or a final spatially adaptive feature map V c ∈R W×H 624, c∈[1,···,C in ] can be calculated as follows:
number
number
[0033] Figure 7 shows a spatially adaptive separable convolutional layer within a convolutional neural network. In at least one embodiment, a depth-based convolution 728, as described in conjunction with Figure 6, is applied to each input image layer or feature map X. c This was performed on 702, 704, and 706, where c∈[1,···,C in In at least one embodiment, the filter or kernel F mc :X c →U mc ∈R W×H is the input feature map X c ∈R W×H Applies to 702, 704, and 706, where c∈[1,···,C in ]. In at least one embodiment, F mc U mc c-th input feature map X to generate 708, 710, 712, 714, 716, 718 c This is the mth filter applied to [the specified location].
[0034] In at least one embodiment, each filter channel U mcFor 708, 710, 712, 714, 716, and 718, as described above, the softmax function 720 is applied after the convolution in units of depth. In at least one embodiment, each softmax 720 output and each filter channel U mc Soft information is collected from the dot products 722, 724 between 708, 710, 712, 714, 716, and 718. In at least one embodiment, the soft information output from the dot products 722, 724 is a compact feature descriptor or a final spatially adaptive feature map V c ∈R W×H 730, 732, 734, c∈[1,...,C in ], as described above, is aggregated for each filter channel to generate 726.
[0035] In at least one embodiment, the point-by-point convolution 736 described above is performed on each compact feature descriptor V c This is performed on 730, 732, and 734 to generate output feature maps 738, 740, and 742. In at least one embodiment, the input values for the point-level convolution 736, or the compact feature descriptors 730, 732, and 734, do not change as the number of “splits” or the number of filters applied in the depth-level convolution 728 increases. In at least one embodiment, each pixel in each “split” and “conflict” operation performed in the depth-level convolution 728 is independent of all other pixels, so each depth-level convolution 728 for each input feature map 702, 704, and 706 may be performed in parallel. In at least one embodiment, each parallel depth-level convolution 728 for each input feature map 702, 704, and 706 may be performed on a graphics processing unit (GPU) or any other parallel processing unit (PPU) as described herein.
[0036] In at least one embodiment, a spatially adaptive separable convolutional layer, as shown in Figure 7, receives an image layer or feature map 702, 704, 706 as input C inThe input image layers or feature maps 702, 704, and 706 are received as such, and each input image layer or feature map has a width W and a height H. In at least one embodiment, the activation size can describe the number of calculations to be performed in the separable convolutional layer. In at least one embodiment, the activation size or calculations for a spatially adaptive separable convolutional layer are as follows:
number
[0037] In at least one embodiment, a convolutional neural network as described herein may attempt to learn filter or kernel values to be applied to input data 702, 704, 706 in a spatially adaptive separable convolutional layer using backpropagation. In at least one embodiment, each layer in the convolutional neural network, or a convolutional layer containing matrix 702, 704, 706 describing weights such as feature maps, can be considered a learnable layer or a layer containing learnable elements. In at least one embodiment, the number of elements of the filter or kernel in the layer that can be learned is the parameter considered for the filter or kernel in the layer. In at least one embodiment, for a spatially adaptive separable convolutional layer as described herein, the parameter count, or learnable elements, is described as follows:
number
number
[0038] Figure 8 illustrates the process for generating an output feature map by a spatially adaptive separable convolutional layer in a convolutional neural network. In at least one embodiment, the spatially adaptive separable convolutional layer begins by performing a depth-unit convolution 816. In at least one embodiment, several steps in the depth-unit convolution 816 may be performed in parallel on a graphics processing unit (GPU) or any other parallel processing unit (PPU), as described herein.
[0039] In at least one embodiment, a spatial filter as described above is applied to an input feature map 804. In at least one embodiment, a softmax function is applied to the spatial information generated by applying the spatial filter to the input feature map 804 806. In at least one embodiment, selected information is generated across the filter channel 808 based on the output from the softmax function 806 and the spatial information generated by applying the spatial filter to the input feature map 804 808. In at least one embodiment, the selected information generated across the filter channel 808 is aggregated into a compact feature descriptor 810. In at least one embodiment, the compact feature descriptor aggregated from the selected information generated across the filter channel 808 810 is combined through point-by-point convolution 812 to generate an output feature map in order to complete the process for generating an output feature map by a spatially adaptive separable convolutional layer 814.
[0040] Logic of reasoning and training Figure 9A shows the inference and / or training logic 915 used to perform inference and / or training operations with respect to one or more embodiments. Further details regarding the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B.
[0041] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, code and / or data storage 901 for storing forward and / or output weights and / or input / output data and / or other parameters for constituting neurons or layers of a neural network that are trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 915 may include, or be coupled to, code and / or data storage 901 for storing graph code or other software for controlling timing and / or sequence, the code and / or data storage 901 being loaded with weight and / or other parameter information to constitute logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, the code, such as graph code, processes the weight or other parameter information based on the architecture of the neural network to which the code corresponds. Load into the ALU. In at least one embodiment, the code and / or data storage 901 stores the weight parameters and / or input / output data of each layer of the neural network being trained or used in conjunction with one or more embodiments, while forward-propagating the input / output data and / or weight parameters during training and / or inference using the embodiments of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 901 may be included together with other on-chip or off-chip data storage, including L1, L2, or L3 caches of the processor or system memory.
[0042] In at least one embodiment, any portion of the code and / or data storage 901 may be inside or outside one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data storage 901 may be cache memory, dynamic randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or code and / or data storage 901 is inside or outside a processor, or whether it consists of DRAM, SRAM, flash, or some other type of storage, may depend on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors.
[0043] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, code and / or data storage 905 for storing backpropagated and / or output weights and / or input / output data corresponding to neurons or layers of a neural network used to train and / or infer in one or more embodiments. In at least one embodiment, the code and / or data storage 905 stores the weight parameters and / or input / output data of each layer of the neural network used to train or in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, the training logic 915 may include, or be coupled to, code and / or data storage 905 for storing graph code or other software for controlling timing and / or sequence, the code and / or data storage 905 being loaded with weight and / or other parameter information to constitute logic including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs). In at least one embodiment, the code, such as graph code, loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In one embodiment, any portion of the code and / or data storage 905 may be included together with other on-chip or off-chip data storage, including L1, L2, or L3 caches of the processor, or system memory. In at least one embodiment, any portion of the code and / or data storage 905 may be inside or outside one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 905 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 905 is, for example, internal or external to the processor, or whether it consists of DRAM, SRAM, flash, or any other type of storage, may be determined depending on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors.
[0044] In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be separate storage structures. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be the same storage structure. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of the code and / or data storage 901 and the code and / or data storage 905 may be included together with other on-chip or off-chip data storage, including L1, L2, or L3 caches of the processor or system memory.
[0045] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, one or more arithmetic logic units ("ALUs") 910, including integer and / or floating-point units, for performing logical and / or arithmetic operations that are at least partially based on or shown by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 920, which are functions of input / output and / or weight parameter data stored in code and / or data storage 901 and / or code and / or data storage 905. In at least one embodiment, the activation stored in the activation storage 920 is generated according to linear algebraic calculations and / or matrix-based calculations performed by the ALU 910 in response to the execution of an instruction or other code, where the weight values stored in the code and / or data storage 905 and / or data 901 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 905, or the code and / or data storage 901, or other on-chip or off-chip storage.
[0046] In at least one embodiment, the ALU910 is contained within one or more processors or other hardware logic devices or circuits, while in another embodiment, the ALU910 may be outside of the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, the ALU910 may be contained within an execution unit of a processor, or otherwise contained within an ALU bank accessible by execution units of a processor, either within the same processor or distributed among different types of processors (e.g., a central processing unit, graphics processing unit, fixed-function unit, etc.). In at least one embodiment, the data storage 901, code and / or data storage 905, and activation storage 920 may be in the same processor or other hardware logic devices or circuits, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activated storage 920 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache, or system memory. Furthermore, the inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic circuits.
[0047] In at least one embodiment, the activated storage 920 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activated storage 920 may be entirely or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activated storage 920 is, for example, inside or outside the processor, or whether it consists of DRAM, SRAM, flash, or some other type of storage, may be determined depending on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used for neural network inference and / or training, or any combination of these factors. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9A may be used in conjunction with application-specific integrated circuits ("ASICs") such as Google's Tensorflow® processing unit, Graphcore® inference processing unit (IPU), or Intel Corp.'s Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9A may be used in conjunction with other hardware such as central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or field-programmable gate array ("FPGA").
[0048] Figure 9B shows an inference and / or training logic 915 in various forms of at least one embodiment. In at least one embodiment, the inference and / or training logic 915 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used in conjunction with, weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9B may be used in conjunction with application-specific integrated circuits ("ASICs") such as Google's Tensorflow® processing unit, Graphcore® inference processing unit (IPU), or Intel Corp's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 915 shown in Figure 9B may be used in conjunction with other hardware such as central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or field-programmable gate array ("FPGA"). In at least one embodiment, the inference and / or training logic 915 may include, without limitation, code and / or data storage 901 and code and / or data storage 905, which may be used to store other information including code (e.g., graph code), weight values and / or bias values, gradient information, momentum values and / or other parameter or hyperparameter information. In at least one embodiment shown in Figure 9B, each of the code and / or data storage 901 and code and / or data storage 905 is associated with dedicated computing resources such as computing hardware 902 and computing hardware 906, respectively. In at least one embodiment, each of the computing hardware 902 and computing hardware 906 includes one or more ALUs that execute mathematical functions, such as linear algebraic functions, only on the information stored in the code and / or data storage 901 and code and / or data storage 905, respectively, and the results are stored in activation storage 920.
[0049] In at least one embodiment, each of the code and / or data storages 901 and 905, and the corresponding computing hardware 902 and 906, corresponds to different layers of a neural network, and the activation resulting from one “storage / computation pair 901 / 902” of the code and / or data storage 901 and computing hardware 902 is provided as input to the next “storage / computation pair 905 / 906” of the code and / or data storage 905 and computing hardware 906 to reflect the conceptual organization of the neural network. In at least one embodiment, the storage / computation pairs 901 / 902 and 905 / 906 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 915 after or in parallel with the storage / computation pairs 901 / 902 and 905 / 906.
[0050] Training and implementation of neural networks Figure 10 shows the training and deployment of a deep neural network in at least one embodiment. In at least one embodiment, an untrained neural network 91006 is trained using the training dataset 1002. In at least one embodiment, the training framework 1004 is the PyTorch framework, while in other embodiments, the training framework 1004 is Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 and generates a trained neural network 1008, enabling it to be trained using the processing resources described herein. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0051] In at least one embodiment, the untrained neural network 1006 is trained using supervised learning, where the training dataset 1002 includes inputs paired with desired outputs for each input, or the training dataset 1002 includes inputs with known outputs, and the outputs of the neural network 1006 are manually scored. In at least one embodiment, the untrained neural network 1006 is trained in a supervised manner, processing inputs from the training dataset 1002 and comparing the resulting outputs with a set of expected or desired outputs. In at least one embodiment, the error is then backpropagated through the untrained neural network 1006. In at least one embodiment, the training framework 1004 adjusts the weights controlling the untrained neural network 1006. In at least one embodiment, the training framework 1004 includes a tool to monitor how well the untrained neural network 1006 is converging toward a model such as a trained neural network 1008 that is better suited to generating the correct answer in result 1014, etc., based on known input data such as new data 1012. In at least one embodiment, the training framework 1004 iteratively trains the untrained neural network 1006 while adjusting the weights to refine the output of the untrained neural network 1006 using a loss function and tuning algorithms such as stochastic gradient descent. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 until it reaches a desired accuracy. In at least one embodiment, the trained neural network 1008 can then be introduced to perform any number of machine learning operations.
[0052] In at least one embodiment, an untrained neural network 1006 is trained using unsupervised learning, where the untrained neural network 1006 attempts to train itself using unlabeled data. In at least one embodiment, the training dataset 1002 for unsupervised learning includes input data that has no associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 1006 can learn grouping within the training dataset 1002 and determine how individual inputs relate to the untrained dataset 1002. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, which is a type of trained neural network 1008 that can perform useful operations to reduce the dimensionality of the novel data 1012. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for the identification of data points in the novel dataset 1012 that deviate from the normal patterns of the novel dataset 1012.
[0053] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled and unlabeled data are mixed in the training dataset 1002. In at least one embodiment, incremental learning, such as a transmission learning technique, may be performed using the training framework 1004. In at least one embodiment, incremental learning enables the trained neural network 1008 to adapt to new data 1012 without forgetting the knowledge taught to the network during initial training.
[0054] Data center Figure 11 shows an exemplary data center 1100 in which at least one embodiment may be used. In at least one embodiment, the data center 1100 includes a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and an application layer 1140.
[0055] As shown in Figure 11, in at least one embodiment, the data center infrastructure layer 1110 may include a resource orchestrator 1112, grouped computing resources 1114, and node computing resources ("node CRs") 1116(1) to 1116(N), where "N" represents any positive integer. In at least one embodiment, the node CRs 1116(1) to 1116(N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (semiconductor drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules. In at least one embodiment, one or more of the nodes CR1116(1) to 1116(N) may be servers having one or more of the computing resources described above.
[0056] In at least one embodiment, the grouped computing resources 1114 may include separate groups of node CRs housed in one or more racks (not shown), or a number of racks housed in a data center in various graphical locations (also not shown). Separate groups of node CRs within the grouped computing resources 1114 may include grouped compute resources, network resources, memory resources, or storage resources that are configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped in one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0057] In at least one embodiment, the resource orchestrator 1112 may constitute or otherwise control one or more nodes CR1116(1) to 1116(N) and / or grouped computing resources 1114. In at least one embodiment, the resource orchestrator 1112 may include a software design infrastructure ("SDI") management entity for the data center 1100. In at least one embodiment, the resource orchestrator may include hardware, software, or any combination thereof.
[0058] In at least one embodiment shown in Figure 11, the framework layer 1120 includes a job scheduler 1132, a configuration manager 1134, a resource manager 1136, and a distribution file system 1138. In at least one embodiment, the framework layer 1120 may include a framework for supporting software 1132 of the software layer 1130 and / or one or more applications 1142 of the application layer 1140. In at least one embodiment, the software 1132 or the application 1142 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1120 may be, but is not limited to, a type of free, open-source software web application framework, such as Apache Spark® ("Spark"), which can use the distribution file system 1138 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 1132 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1100. In at least one embodiment, the configuration manager 1134 may be capable of configuring different layers, such as the software layer 1130 and the framework layer 1120, which includes Spark and a distribution file system 1138 to support large-scale data processing. In at least one embodiment, the resource manager 1136 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support the distribution file system 1138 and the job scheduler 1132. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1114 located in the data center infrastructure layer 1110.In at least one embodiment, the resource manager 1136 may work in conjunction with the resource orchestrator 1112 to manage these mapped or allocated computing resources.
[0059] In at least one embodiment, the software 1132 included in the software layer 1130 may include software used by at least a portion of the nodes CR1116(1) to 1116(N), the grouped computing resources 1114, and / or the distribution file system 1138 of the framework layer 1120. One or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.
[0060] In at least one embodiment, application 1142 included in application layer 1140 may include one or more types of applications used by at least a portion of nodes CR1116(1) to 1116(N), grouped computing resources 1114, and / or the distribution file system 1138 of framework layer 1120. One or more types of applications may include, but are not limited to, any number of genomics applications, recognition compute, and machine learning applications including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0061] In at least one embodiment, any of the configuration manager 1134, resource manager 1136, and resource orchestrator 1112 may implement any number and type of self-correcting measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting measures may enable the data center operator of data center 1100 to avoid determining potentially faulty configurations and to eliminate underutilized and / or underperforming portions of the data center.
[0062] In at least one embodiment, the data center 1100 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 1100. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1100 by using weight parameters computed by one or more techniques described herein.
[0063] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to perform training or inference on information such as image recognition, speech recognition, or other artificial intelligence services.
[0064] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 11 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0065] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system diagram 11 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operation, neural network functionality and / or architecture, or neural network use case described herein.
[0066] Autonomous vehicles Figure 12A shows an example of the autonomous vehicle 1200 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1200 (or, as referred to herein as "vehicle 1200") may be, without limitation, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle accommodating one or more occupants. In at least one embodiment, vehicle 1200 may be a semi-tractor trailer truck for cargo transport. In at least one embodiment, vehicle 1200 may be an aircraft, a robotic vehicle, or another type of vehicle.
[0067] Autonomous vehicles may also be described in terms of automation levels as defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") in their "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (for example, standard No. J3016-201806, published June 15, 2018; standard No. J3016-201609, published September 30, 2016; and previous and newer versions of this standard). In one or more embodiments, the vehicle 1200 may be capable of functioning at one or more of the autonomous driving levels from Level 1 to Level 5. For example, in at least one embodiment, the vehicle 1200 may be capable of conditional automation (level 3), high automation (level 4), and / or full automation (level 5), depending on the embodiment.
[0068] In at least one embodiment, the vehicle 1200 may include, without limitation, components such as a chassis, a vehicle body, wheels (2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, the vehicle 1200 may include, without limitation, a propulsion system 1250 such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, the propulsion system 1250 may be coupled to the drivetrain of the vehicle 1200, and the drivetrain may include, without limitation, a transmission to enable the propulsion of the vehicle 1200. In at least one embodiment, the propulsion system 1250 may be controlled in response to receiving a signal from a throttle / accelerator 1252.
[0069] In at least one embodiment, a steering system 1254, which may include a steering wheel (without limitation), is used to steer the vehicle 1200 (for example, along a desired path or route) when the propulsion system 1250 is operating (for example, when the vehicle is moving). In at least one embodiment, the steering system 1254 may receive signals from a steering actuator 1256. The steering wheel may be optional with respect to fully automated (Level 5) functionality. In at least one embodiment, a brake sensor system 1246 may be used to actuate the vehicle brakes in response to receiving signals from a brake actuator 1248 and / or brake sensor.
[0070] In at least one embodiment, the controller 1236, which may include, without limitation, one or more system-on-chip ("SoC") (not shown in Figure 12A) and / or graphics processing unit ("GPU"), provides signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1200. For example, in at least one embodiment, the controller 1236 may transmit signals to operate the vehicle brakes via the brake actuator 1248, signals to operate the steering system 1254 via the steering actuator 1256, and signals to operate the propulsion system 1250 via the throttle / accelerator 1252. The controller 1236 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in the driving vehicle 1200. In at least one embodiment, the controller 1236 may include a first controller 1236 for autonomous driving functions, a second controller 1236 for functional safety functions, a third controller 1236 for artificial intelligence functions (e.g., computer vision), a fourth controller 1236 for infotainment functions, a fifth controller 1236 for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller 1236 may address two or more of the above functionalities, two or more controllers 1236 may address a single functionality, and / or any combination thereof.
[0071] In at least one embodiment, the controller 1236 provides signals for controlling one or more components and / or systems of the vehicle 1200 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may include, for example, a global navigation satellite system ("GNSS") sensor 1258 (e.g., a global positioning system sensor), a radar sensor 1260, an ultrasonic sensor 1262, a lithium-ion radar sensor 1264, and an inertial measurement unit ("IMU"), without limiting them. The unit may receive signals from sensors 1266 (e.g., accelerometer, gyroscope, magnetic compass, magnetometer, etc.), microphone 1296, stereo camera 1268, wide-angle camera 1270 (e.g., fisheye camera), infrared camera 1272, ambient camera 1274 (e.g., 360-degree camera), long-range camera (not shown in Figure 12A), medium-range camera (not shown in Figure 12A), speed sensor 1244 (e.g., for measuring the speed of vehicle 1200), vibration sensor 1242, steering sensor 1240, brake sensor (e.g., as part of brake sensor system 1246), and / or other types of sensors.
[0072] In at least one embodiment, one or more of the controllers 1236 may receive inputs (represented, for example, by input data) from the instrument cluster 1232 of the vehicle 1200 and provide outputs (represented, for example, by output data, display data, etc.) via the human-machine interface ("HMI") display 1234, an audible annunciator, a loudspeaker, and / or other components of the vehicle 1200. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high-definition map (not shown in Figure 12A)), location data (e.g., the location of vehicle 1200 on a map), direction, location of other vehicles (e.g., occupying grid), and information about objects and the state of objects sensed by controller 1236. For example, in at least one embodiment, HMI display 1234 may display information about the presence of one or more objects (e.g., road signs, warning signs, changes in traffic lights, etc.) and / or information about driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at exit 34B 3.22 km (2 miles) ahead, etc.).
[0073] In at least one embodiment, the vehicle 1200 further includes a network interface 1224, which may use a wireless antenna 1226 and / or a modem for communication over one or more networks. For example, in at least one embodiment, the network interface 1224 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA®"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), and the like. In at least one embodiment, the wireless antenna 1226 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, and / or low power wide-area networks ("LPWAN") such as LoRaWAN and SigFox.
[0074] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 12A for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0075] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with system diagram 12A for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operation, neural network functionality and / or architecture, or neural network use case described herein.
[0076] Figure 12B shows an example of camera locations and fields of view for the autonomous vehicle 1200 of Figure 12A according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are illustrative and not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included, and / or the cameras may be positioned at different locations on the vehicle 1200.
[0077] In at least one embodiment, the camera type of the camera may include, but is not limited to, a digital camera which may be adapted for use with components and / or systems of the vehicle 1200. The camera may operate at Automotive Safety Integrity Level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera type may be capable of handling any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a roll shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-clear-clear-clear ("RCCC") color filter array, a red-clear-clear-blue ("RCCB") color filter array, a red-blue-green-clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a clear pixel camera, such as a camera having RCCC, RCCB, and / or RBGC color filter arrays, may be used to increase light sensitivity.
[0078] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance system ("ADAS") functions (for example, as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (for example, all of the cameras) may simultaneously record and provide image data (for example, video).
[0079] In at least one embodiment, one or more of the cameras may be mounted on a mounting assembly, such as a custom-designed (three-dimensionally printed) assembly, to eliminate stray light and reflections from inside the vehicle (e.g., reflections from the dashboard to the windshield) that could interfere with the camera's image data acquisition performance. Referring to a door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed so that the camera mounting plate conforms to the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. For side-view cameras, the camera may also be integrated into the four pillars at each corner of the cabin.
[0080] In at least one embodiment, a camera having a field of view that includes a portion of the environment in front of the vehicle 1200 (e.g., a front camera) may be used for the surrounding view to facilitate the identification of the path and obstacles ahead, and may be used in conjunction with one or more controllers 1236 and / or control SoCs to assist in providing information essential for generating an occupied grid and / or determining a preferred vehicle path. In at least one embodiment, the front camera may be used to perform many of the same ADAS functions as LIDAR, including, but not limited to, emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the front camera may also be used for ADAS functions and systems, including, but not limited to, other functions such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0081] In at least one embodiment, various cameras, including a monocular camera platform including, for example, a CMOS (complementary metal oxide semiconductor) color imaging device, may be used in a front configuration. In at least one embodiment, a wide-angle camera 1270 may be used to sense objects entering the view from the surroundings (e.g., pedestrians, cross-traffic, or bicycles). Although only one wide-angle camera 1270 is shown in Figure 12B, in other embodiments, the vehicle 1200 may have any number of wide-angle cameras 1270 (including zero). In at least one embodiment, any number of long-range cameras 1298 (e.g., a pair of stereo cameras with long-range views) may be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. In at least one embodiment, the long-range cameras 1298 may also be used for object detection and classification, as well as basic object tracking.
[0082] In at least one embodiment, any number of stereo cameras 1268 may also be included in a front configuration. In at least one embodiment, one or more stereo cameras 1268 may include an integrated control unit with an expandable processing unit, which may provide a programmable logic ("FPGA") and a multi-core microprocessor having an integrated Controller Area Network ("CAN") or Ethernet® interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the vehicle 1200's environment, including distance estimation for all points in the image. In at least one embodiment, one or more of the stereo cameras 1268 may include, without limitation, a compact stereo vision sensor, which may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle 1200 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1268 may be used in addition to or instead of those described herein.
[0083] In at least one embodiment, a camera having a field of view that includes a portion of the environment to the sides of the vehicle 1200 (e.g., a side-view camera) may be used for the surrounding view to provide information used for creating and updating the occupancy grid and generating side collision warnings. For example, in at least one embodiment, surrounding cameras 1274 (e.g., four surrounding cameras 1274 as shown in Figure 12B) may be positioned on the vehicle 1200. The surrounding cameras 1274 may include, without limitation, any number and combination of wide-angle cameras 1270, fisheye cameras, and / or 360-degree cameras. For example, in at least one embodiment, four fisheye cameras may be positioned in front of, behind, and to the sides of the vehicle 1200. In at least one embodiment, the vehicle 1200 may use three surrounding cameras 1274 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a front camera) as a fourth surrounding camera.
[0084] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 1200 (e.g., a rear-view camera) may be used for parking assistance, surrounding view, and rear collision warning to create and update the occupancy grid. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as front cameras as described herein (e.g., long-range camera 1298 and / or medium-range camera 1276, stereo camera 1268, infrared camera 1272, etc.).
[0085] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 12B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0086] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with system diagram 12B for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operation, neural network functionality and / or architecture, or neural network use case described herein.
[0087] Figure 12C is a block diagram illustrating an exemplary system architecture of the autonomous vehicle 1200 of Figure 12A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1200 in Figure 12C is shown as being connected via a bus 1202. In at least one embodiment, the bus 1202 may include, without limitation, a CAN data interface (or, as referred to herein, the “CAN bus”). In at least one embodiment, the CAN may be an internal network of the vehicle 1200 used to assist in the control of various features and functions of the vehicle 1200, such as brake activation, acceleration, brake control, steering, and windshield wipers. In at least one embodiment, the bus 1202 may be configured to have tens or even hundreds of nodes, each having its own unique identifier (e.g., a CAN ID). In at least one embodiment, the bus 1202 may be read to find steering angle, ground speed, engine revolutions per minute ("RPM"), button position, and / or other vehicle state indicators. In at least one embodiment, bus 1202 may be a CAN bus compliant with ASIL B.
[0088] In at least one embodiment, FlexRay and / or Ethernet® may be used in addition to or instead of CAN. In at least one embodiment, there may be any number of buses forming bus 1202, which may include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 1202 may be used to perform different functions and / or to provide redundancy. For example, a first bus 1202 may be used for collision avoidance functions and a second bus 1202 may be used for operation control. In at least one embodiment, each bus 1202 may communicate with any of the components of the vehicle 1200, and two or more buses 1202 may communicate with the same component. In at least one embodiment, each of any number of system-on-chip ("SoC") 1204, each of the controllers 1236, and / or each computer in the vehicle may have access to the same input data (e.g., inputs from sensors in the vehicle 1200) and may be connected to a common bus such as a CAN bus.
[0089] In at least one embodiment, the vehicle 1200 may include one or more controllers 1236, such as those described herein with respect to Figure 12A. The controllers 1236 may be used for a variety of functions. In at least one embodiment, the controllers 1236 may be coupled to any of the various other components and systems of the vehicle 1200, and may be used for the control of the vehicle 1200, the artificial intelligence of the vehicle 1200, and / or the infotainment of the vehicle 1200, etc.
[0090] In at least one embodiment, the vehicle 1200 may include any number of SoCs 1204. Each of the SoCs 1204 may include, without limitation, a central processing unit ("CPU") 1206, a graphics processing unit ("GPU") 1208, a processor 1210, a cache 1212, an accelerator 1214, a data store 1216, and / or other components and features not shown. In at least one embodiment, the SoCs 1204 may be used to control the vehicle 1200 in various platforms and systems. For example, in at least one embodiment, the SoCs 1204 may be incorporated into a system (e.g., the system of the vehicle 1200) having a high-definition ("HD") map 1222 that can be refreshed and / or updated from one or more servers (not shown in Figure 12C) via a network interface 1224.
[0091] In at least one embodiment, the CPU 1206 may include a CPU cluster, or CPU complex (or referred to herein as “CCPLEX”). In at least one embodiment, the CPU 1206 may include multiple cores and / or Level 2 ("L2") caches. For example, in at least one embodiment, the CPU 1206 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 1206 may include four dual-core clusters, each having its own dedicated L2 cache (e.g., 2MB of L2 cache). In at least one embodiment, the CPU 1206 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, which allows any combination of clusters of the CPU 1206 to be activated at any given time.
[0092] In at least one embodiment, one or more of the CPUs 1206 may implement a power management function, which includes, but is not limited to, one or more of the following features: individual hardware blocks may be automatically clock-gate controlled when idle to conserve dynamic power; each core clock may be gated when a core is not actively executing instructions due to the execution of an interrupt-wait ("WFI") / event-wait ("WFE") instruction; each core may be power-gated independently; each core cluster may be clock-gated independently when all cores are clock-gated or power-gated; and / or each core cluster may be power-gated independently when all cores are power-gated. In at least one embodiment, the CPU 1206 may further implement an extended algorithm for managing power states, where acceptable power states and expected wake-up times are specified, and the hardware / microcode determines the best power state in which the cores, clusters, and CCPLEX should be. In at least one embodiment, the processing core may support in software a simple sequence for entering a power state with the work offloaded to microcode.
[0093] In at least one embodiment, the GPU1208 may include an integrated GPU (or, as referred to herein, an "iGPU"). In at least one embodiment, the GPU1208 may be programmable and efficient for parallel workloads. In at least one embodiment, the GPU1208 may use an extended tensor instruction set. In one embodiment, the GPU1208 may include one or more streaming microprocessors, each of which may include a Level 1 ("L1") cache (e.g., an L1 cache with a storage capacity of at least 96KB), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache with a storage capacity of 512KB). In at least one embodiment, the GPU1208 may include at least eight streaming microprocessors. In at least one embodiment, the GPU1208 may use a compute application programming interface (API). In at least one embodiment, the GPU1208 may use one or more parallel computing platforms and / or programming modules (for example, NVIDIA's CUDA model).
[0094] In at least one embodiment, one or more of the GPU1208s may be power-optimized to best perform in automotive and embedded use cases. For example, in one embodiment, the GPU1208 may be fabricated on a Finn field-effect transistor ("FinFET"). In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores divided into multiple blocks. For example, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks, without limitation. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR cores for deep learning matrix operations, a level zero ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor may include independent parallel data paths for integers and floating-point numbers, enabling efficient execution of workloads by combining computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities, enabling finer-grained synchronization and coordination between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0095] In at least one embodiment, one or more of the GPU1208s include high-bandwidth memory ("HBM") and / or a 16GB HBM2 memory subsystem, which in some examples may provide a peak memory bandwidth of approximately 900GB / s. In at least one embodiment, in addition to or instead of HBM memory, synchronous graphics random-access memory ("SGRAM") such as graphics double data rate type five synchronous random-access memory ("GDDR5") may be used.
[0096] In at least one embodiment, the GPU1208 may include integrated memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU1208 to directly access the CPU1206's page table. In at least one embodiment, when the GPU1208 memory management unit ("MMU") encounters a miss, an address translation request may be sent to the CPU1206. In at least one embodiment, in response, the CPU1206 may look up the virtual-to-physical address mapping in its own page table and send the translation back to the GPU1208. In at least one embodiment, integrated memory technology makes it possible to provide a single, unified virtual address space for the memories of both the CPU1206 and the GPU1208, thereby simplifying the programming of the GPU1208 and the porting of applications to the GPU1208.
[0097] In at least one embodiment, the GPU 1208 may include any number of access counters that can record how often the GPU 1208 accesses the memory of other processors. In at least one embodiment, the access counters may help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently, thereby improving the efficiency of memory ranges shared between processors.
[0098] In at least one embodiment, one or more of the SoC1204 may include any number of caches 1212, including those described herein. For example, in at least one embodiment, the cache 1212 may include a Level 3 ("L3") cache that is available to both the CPU 1206 and the GPU 1208 (e.g., connected to both the CPU 1206 and the GPU 1208). In at least one embodiment, the cache 1212 may include a write-back cache that can record the state of the line, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4 MB or more, depending on the embodiment, but smaller cache sizes may be used.
[0099] In at least one embodiment, one or more of the SoC1204 may include one or more accelerators 1214 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC1204 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU1208 and offload some of the tasks of the GPU1208 (e.g., free up more cycles of the GPU1208 to perform other tasks). In at least one embodiment, accelerators 1214 can be used for target workloads that are stable enough to accept acceleration (e.g., perception, convolutional neural networks ("CNNs"), recurrent neural networks ("RNNs"), etc.). In at least one embodiment, the CNN may include region-based, i.e., regional convolutional neural networks ("RCNN"), and fast RCNNs (for example, used for object detection), or other types of CNNs.
[0100] In at least one embodiment, the accelerator 1214 (e.g., a hardware acceleration cluster) may include a deep learning accelerator ("DLA"). The DLA may include, without limitation, one or more Tensor processing units ("TPUs"), which may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be configured to perform image processing functions (e.g., CNN, RCNN, etc.) and may be an accelerator optimized for that purpose. The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inference. In at least one embodiment, the design of the DLA can improve performance per millisecond compared to a typical general-purpose GPU, and typically far exceed the performance of a CPU. In at least one embodiment, the TPU may execute several functions, including, for example, a single-instance convolution function supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions. In at least one embodiment, the DLA may rapidly and efficiently run a neural network, particularly a CNN, on processed or raw data for any of a variety of functions, including, for example, a CNN for object recognition and detection using data from camera sensors, a CNN for distance estimation using data from camera sensors, a CNN for emergency vehicle detection and identification and detection using data from microphone 1296, a CNN for facial recognition and vehicle owner identification using data from camera sensors, and / or a CNN for security and / or safety-related events.
[0101] In at least one embodiment, the DLA may perform any function of the GPU1208, and the designer may target either the DLA or the GPU1208 for any function, for example by using an inference accelerator. For example, in at least one embodiment, the designer may concentrate CNN and floating-point arithmetic processing on the DLA and offload other functions to the GPU1208 and / or other accelerators 1214.
[0102] In at least one embodiment, the accelerator 1214 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator ("PVA"), which may be referred to herein as a computer vision accelerator instead. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver-assistance systems ("ADAS") 1238, autonomous driving, augmented reality ("AR") applications, and / or virtual reality ("VR") applications. The PVA may provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, for example, any number of reduced instruction set computer ("RISC") cores, direct memory access ("DMA"), and / or any number of vector processors.
[0103] In at least one embodiment, the RISC core may interact with an image sensor (e.g., an image sensor from any of the cameras described herein) and / or an image signal processor, etc. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols, depending on the embodiment. In at least one embodiment, the RISC core may run a real-time operating system ("RTOS"). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits ("ASICs") and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0104] In at least one embodiment, the DMA may allow components of the PVA to access system memory independently of the CPU 1206. In at least one embodiment, the DMA may support any number of features used to provide optimization to the PVA, including but not limited to multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0105] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly perform programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or vector memory (e.g., "VMEM"). In at least one embodiment, the VPU may include digital signal processors such as single instruction, multiple data ("SIMD") and very long instruction word ("VLIW") digital signal processors. In at least one embodiment, a combination of SIMD and VLIW may improve throughput and speed.
[0106] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to run independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or furthermore, execute different algorithms on consecutive images or on parts of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to enhance the overall safety of the system.
[0107] In at least one embodiment, the accelerator 1214 may include (e.g., a hardware acceleration cluster), an on-chip computer vision network, and static random access memory ("SRAM"), providing high-bandwidth, low-latency SRAM for the accelerator 1214. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM consisting of, for example, eight field-configurable memory blocks, which may be accessible from both the PVA and the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone that provides the PVA and DLA with high-speed access to the memory. In at least one embodiment, the backbone may include an on-chip computer vision network interconnecting the PVA and DLA to the memory (e.g., using the APB).
[0108] In at least one embodiment, the on-chip computer vision network may include an interface that determines whether both the PVA and DLA provide ready and enable signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transfer. In at least one embodiment, the interface may conform to the International Organization for Standardization ("ISO") 26262 or the International Electrotechnical Commission ("IEC") 61508 standard, but other standards and protocols may be used.
[0109] In at least one embodiment, one or more of the SoC1204s may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the position and extent of an object (e.g., in a world model) and generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general waveform propagation simulation, comparison with LIDAR data for localization and / or other functions, and / or other uses.
[0110] In at least one embodiment, accelerator 1214 (e.g., a hardware accelerator cluster) has a variety of applications for autonomous driving. In at least one embodiment, the PVA may be a programmable vision accelerator that can be used in key processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well suited to algorithmic domains that require low-power and low-latency predictable processing. In other words, the PVA performs well even with small datasets for semi-dense or dense regular computations that require low-latency and low-power predictable runtime. In at least one embodiment, in an autonomous vehicle such as vehicle 1200, the PVA is designed to run conventional computer vision algorithms because they are effective for object detection and integer calculations.
[0111] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using PVA. In at least one embodiment, algorithms based on semi-global matching may be used in some examples, but this is not limited to them. In at least one embodiment, applications for Level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structuring from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may perform computer stereo vision functions for input from two monocular cameras.
[0112] In at least one embodiment, PVA may be used to perform high-density optical flow. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and processed time-of-flight data is provided by processing raw time-of-flight data, for example.
[0113] In at least one embodiment, DLA may be used to run any type of network for enhancing control and driving safety, including, for example, a neural network that outputs a measure of reliability for each object detection. In at least one embodiment, reliability may be expressed or interpreted as the probability of each detection compared to other detections, or as providing a relative “weight.” In at least one embodiment, reliability allows the system to make further decisions about which detections should be considered positive rather than false positives. For example, in at least one embodiment, the system may set a threshold for reliability and consider only detections exceeding the threshold as positive. In embodiments where an automatic emergency braking ("AEB") system is used, a false positive would cause the vehicle to automatically apply the emergency brakes, which is obviously undesirable. In at least one embodiment, a highly reliable detection may be considered a trigger for the AEB. In at least one embodiment, DLA may run a neural network to regress the confidence value. In at least one embodiment, the neural network may take as its input at least a subset of parameters such as the dimensions of the bounding box, ground estimation obtained (e.g., from another subsystem), output from IMU sensor 1266 correlated with the orientation of vehicle 1200, distance, and 3D location estimation of an object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 1264 or RADAR sensor 1260).
[0114] In at least one embodiment, one or more of the SoC1204 may include a data store 1216 (e.g., memory). In at least one embodiment, the data store 1216 may be on-chip memory of the SoC1204, which may store a neural network running on the GPU1208 and / or DLA. In at least one embodiment, the capacity of the data store 1216 may be large enough to store multiple instances of the neural network for redundancy and safety. In at least one embodiment, the data store 1212 may include an L2 or L3 cache.
[0115] In at least one embodiment, one or more of the SoC1204 may include any number of processors 1210 (e.g., embedded processors). The processors 1210 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and associated security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC1204 and may provide run-time power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance in transitioning the system to a low-power state, management of thermal and temperature sensors of the SoC1204, and / or management of the power state of the SoC1204. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the SoC1204 may use the ring oscillator to detect the temperature of the CPU 1206, GPU 1208, and / or accelerator 1214. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature failure routine, put the SoC 1204 into a low-power state, and / or put the vehicle 1200 into driver-safe stop mode (for example, safely stop the vehicle 1200).
[0116] In at least one embodiment, the processor 1210 may further include a set of embedded processors that can serve as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide range of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0117] In at least one embodiment, the processor 1210 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0118] In at least one embodiment, the processor 1210 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for addressing safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), and / or routing logic. In safety mode, in at least one embodiment, two or more cores operate in lockstep mode and may function as a single core with comparison logic for detecting any differences between their operations. In at least one embodiment, the processor 1210 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for addressing real-time camera management. In at least one embodiment, the processor 1210 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0119] In at least one embodiment, the processor 1210 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that performs video post-processing functions required by a video playback application to produce a final image for the playback device window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1270, the ambient camera 1274, and / or the in-cabin surveillance camera sensor. In at least one embodiment, the in-cabin surveillance camera sensor is preferably monitored by a neural network running on another instance of the SoC 1204, which is configured to identify events in the cabin and respond thereto accordingly. In at least one embodiment, the in-cabin system may perform lip-reading, without limitation, to activate cellular services, make phone calls, write emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable at other times.
[0120] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, if motion occurs in the video, the noise reduction appropriately weights the spatial information to reduce the weight of the information provided by adjacent frames. In at least one embodiment, if the image or part of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from previous images to reduce noise in the current image.
[0121] In at least one embodiment, the video image synthesizer may also be configured to perform stereo parallelization on the input stereo lens frame. In at least one embodiment, the video image synthesizer may further be used to synthesize the user interface when the operating system desktop is in use, so that the GPU1208 does not need to continuously render new surfaces. In at least one embodiment, when the GPU1208 is powered on and actively performing 3D rendering, the video image synthesizer may be used to offload the GPU1208 to improve performance and responsiveness.
[0122] In at least one embodiment, one or more of the SoC1204 SoCs may further include a camera serial interface, a high-speed interface, and / or a video input block which may be used for input functions of the camera and associated pixels of a Mobile Industry Processor Interface ("MIPI") for receiving video and camera inputs. In at least one embodiment, one or more of the SoC1204 SoCs may further include an input / output controller which may be software-controlled and may be used to receive I / O signals that are not tied to a specific role.
[0123] In at least one embodiment, one or more of the SoC1204 may further include a broad peripheral interface for enabling communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. The SoC1204 may be used to process data from cameras (connected, for example, via Gigabit Multimedia Serial Link and Ethernet®), data from sensors (e.g., LiDAR sensor 1264, RADAR sensor 1260, etc., which may be connected via Ethernet®), data from bus 1202 (e.g., vehicle speed, steering wheel position, etc.), data from GNSS sensor 1258 (connected, for example, via Ethernet® or CAN bus), and the like. In at least one embodiment, one or more of the SoC1204 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to free the CPU 1206 from routine data management tasks.
[0124] In at least one embodiment, the SoC1204 may be an end-to-end platform with a flexible architecture spanning automation levels 3 to 5, thereby providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques to achieve diversity and redundancy, and a platform for a flexible and reliable driving software stack, along with deep learning tools. In at least one embodiment, the SoC1204 may be faster, more reliable, and more energy-efficient and space-efficient than conventional systems. For example, in at least one embodiment, the accelerator 1214, when combined with the CPU 1206, GPU 1208, and data store 1216, can realize a fast and efficient platform for Level 3 to 5 autonomous vehicles.
[0125] In at least one embodiment, the computer vision algorithm may run on a CPU, which may be constructed using a high-level programming language such as the C programming language, and may execute a variety of processing algorithms across a variety of visual data. However, in at least one embodiment, the CPU often fails to meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms used in in-vehicle ADAS applications and in realistic Level 3-5 autonomous vehicles in real time.
[0126] The embodiments described herein allow multiple neural networks to run simultaneously and / or sequentially, and the results can be combined to enable Level 3–5 autonomous driving capabilities. For example, in at least one embodiment, a DLA or a CNN running on a separate GPU (e.g., GPU1220) may include text and word recognition, enabling a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained to read. In at least one embodiment, the DLA may further include a neural network that can identify, interpret, and provide a semantic understanding of the signs, which can then pass that semantic understanding to a route planning module running on a CPU complex.
[0127] In at least one embodiment, multiple neural networks may run simultaneously with respect to Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign consisting of an electric light and the text "Caution: Flashing indicates frozen conditions" may be interpreted separately or collectively by several neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), and the text "Flashing indicates frozen conditions" may be interpreted by a second introduced neural network, which, if flashing light is detected, notifies the vehicle's route planning software (preferably running on the CPU complex) that frozen conditions are present. In at least one embodiment, the flashing light may also be identified by running a third introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may run simultaneously within the DLA and / or on the GPU1208, etc.
[0128] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1200. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves the vehicle. Thus, SoC1204 provides security against theft and / or vehicle hijacking.
[0129] In at least one embodiment, the CNN for emergency vehicle detection and identification may use data from microphone 1296 to detect and identify emergency vehicle sirens. In at least one embodiment, SoC 1204 uses the CNN to classify visual data as well as environmental and urban sounds. In at least one embodiment, the CNN running on DLA is trained to identify the relative speed of approaching emergency vehicles (for example, by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area in which the vehicle is operating, which are identified by GNSS sensor 1258. In at least one embodiment, if operating in Europe, the CNN attempts to detect European sirens, and if in the United States, it attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine may be used to slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle using ultrasonic sensor 1262 until the emergency vehicle has passed.
[0130] In at least one embodiment, the vehicle 1200 may include a CPU 1218 (e.g., a separate CPU or dCPU) which may be coupled to the SoC 1204 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU 1218 may include, for example, an x86 processor. The CPU 1218 may be used to perform any of a variety of functions, including, for example, mediating potentially inconsistent results between ADAS sensors and the SoC 1204, and / or monitoring the status and health of the controller 1236 and / or the infotainment system ("Infotainment SoC") 1230 on the chip.
[0131] In at least one embodiment, the vehicle 1200 may include a GPU 1220 (e.g., a separate GPU or dGPU) which may be coupled to the SoC 1204 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). In at least one embodiment, the GPU 1220 may provide additional artificial intelligence capabilities, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on input from the vehicle 1200's sensors (e.g., sensor data).
[0132] In at least one embodiment, the vehicle 1200 may further include a network interface 1224, which may include, but is not limited to, one or more wireless antennas 1226 for different communication protocols, such as a cellular antenna or a Bluetooth antenna. In at least one embodiment, the network interface 1224 may be used to enable wireless connectivity over the Internet to a cloud (e.g., a server and / or other network devices), other vehicles, and / or computing devices (e.g., occupant client devices). In at least one embodiment, a direct link may be established between the vehicle 120 and other vehicles for communication with other vehicles, and / or an indirect link may be established (e.g., over a network and over the Internet). In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 1200 with information about nearby vehicles (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1200). In at least one embodiment, the functions described above may be part of the cooperative adaptive cruise control function of the vehicle 1200.
[0133] In at least one embodiment, the network interface 1224 may include an SoC that provides modulation and demodulation functions, enabling the controller 1236 to communicate over a wireless network. In at least one embodiment, the network interface 1224 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, frequency conversion may be performed in any technically feasible manner. For example, frequency conversion can be performed by a well-known process and / or using a superheterodyne process. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functionality for communication over LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0134] In at least one embodiment, the vehicle 1200 may further include a data store 1228, which may include, without limitation, off-chip (e.g., not on the SoC 1204) storage. In at least one embodiment, the data store 1228 may include, without limitation, one or more storage elements, including RAM, SRAM, dynamic random-access memory ("DRAM"), video random-access memory ("VRAM"), flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0135] In at least one embodiment, the vehicle 1200 may further include GNSS sensors 1258 (e.g., GPS and / or auxiliary GPS sensors) to assist in mapping, perception, occupy grid generation, and / or route planning functions. In at least one embodiment, any number of GNSS sensors 1258, including, for example, GPS, may be used, using a USB connector having an Ethernet® to serial (e.g., RS-232) bridge.
[0136] In at least one embodiment, the vehicle 1200 may further include a RADAR sensor 1260. The RADAR sensor 1260 may be used by the vehicle 1200 to perform long-range vehicle detection even in darkness and / or severe weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. The RADAR sensor 1260 may use CAN and / or bus 1202 for control (e.g., to transmit data generated by the RADAR sensor 1260) and to access object tracking data, and in some examples, it may have access to Ethernet® to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example, without limitation, the RADAR sensor 1260 may be suitable for forward, rear, and side RADAR use. In at least one embodiment, one or more of the RADAR sensors 1260 are pulse-Doppler RADAR sensors.
[0137] In at least one embodiment, the RADAR sensor 1260 may include different configurations, such as a narrow-field long-range, a wide-field short-range, and a lateral-covering short-range. In at least one embodiment, the long-range RADAR may be used for adaptive cruise control functionality. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within 250m, achieved by two or more independent scans. In at least one embodiment, the RADAR sensor 1260 may be designed to facilitate the distinction between static and moving objects and may be used by the ADAS system 1238 to provide emergency braking assistance and forward collision warning. The sensor 1260 included in the long-range RADAR system may include, without limitation, multiple (e.g., six or more) fixed RADAR antennas, as well as a monostatic multimode RADAR with high-speed CAN and FlexRay interfaces. In at least one embodiment, if there are six antennas, the four central antennas may generate a concentrated beam pattern designed to record the area around vehicle 1200 at a higher speed with minimal interference from adjacent lanes. In at least one embodiment, the other two antennas may extend the field of view, enabling rapid detection of vehicles entering or leaving the lane of vehicle 1200.
[0138] In at least one embodiment, the medium-range RADAR system may include, for example, a range of up to 160 m (forward) or 80 m (rear) and a field of view of up to 42 degrees (forward) or 150 degrees (rear). In at least one embodiment, the short-range RADAR system may include, without limitation, any number of RADAR sensors 1260 designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor blind spots behind and beside the vehicle. In at least one embodiment, the short-range RADAR system may be used in an ADAS system 1238 to perform blind spot detection and / or lane change assistance.
[0139] In at least one embodiment, the vehicle 1200 may further include an ultrasonic sensor 1262. The ultrasonic sensor 1262 may be positioned in front of, behind, and / or to the side of the vehicle 1200 and may be used for parking assistance and / or to generate and update the occupancy grid. In at least one embodiment, a variety of ultrasonic sensors 1262 may be used, and different ultrasonic sensors 1262 may be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, the ultrasonic sensor 1262 may operate at functional safety level ASIL B.
[0140] In at least one embodiment, the vehicle 1200 may include a LiDAR sensor 1264. The LiDAR sensor 1264 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LiDAR sensor 1264 may have a functional safety level of ASIL B. In at least one embodiment, the vehicle 1200 may include a plurality of LiDAR sensors 1264 (e.g., two, four, six, etc.), and these sensors may use Ethernet® (for example, to provide data to a Gigabit Ethernet® switch).
[0141] In at least one embodiment, the LIDAR sensor 1264 may be capable of providing a list of objects and their distances over a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 1264 may, for example, have an advertised range of approximately 100m, an accuracy of 2cm to 3cm, and support a 100Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors 1264 may be used. In such embodiments, the LIDAR sensor 1264 may be implemented as a small device that can be incorporated into the front, rear, side, and / or corners of a vehicle 1200. In at least one embodiment, the LIDAR sensor 1264 of such embodiments may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, over a range of 200m. In at least one embodiment, a front-mounted LIDAR sensor 1264 may be configured to provide a horizontal field of view of 45 to 135 degrees.
[0142] In at least one embodiment, LiDAR technology such as 3D flash LiDAR may also be used. 3D flash LiDAR uses a laser flash as a source to illuminate the area around the vehicle 1200 up to approximately 200m. In at least one embodiment, the flash LiDAR unit includes, but is not limited to, a receptor that records the transit time of the laser pulse and the reflected light at each pixel, corresponding to the range from the vehicle 1200 to the object. In at least one embodiment, the flash LiDAR enables the generation of highly accurate and distortion-free ambient images with each laser flash. In at least one embodiment, four flash LiDAR sensors may be introduced, one on each side of the vehicle 1200. In at least one embodiment, the 3D flash LiDAR system includes, but is not limited to, a semiconductor 3D staring array LiDAR camera (e.g., a non-scanning LiDAR device) with no moving parts other than a fan. In at least one embodiment, the flash LiDAR device may use a Class I (eye-safe) laser pulse of 5 nanoseconds per frame to capture the reflected laser light in the form of a 3D range point cloud and position-synchronized (co-registered) intensity data.
[0143] In at least one embodiment, the vehicle may further include an IMU sensor 1266. In at least one embodiment, the IMU sensor 1266 may be positioned in the center of the rear axle of the vehicle 1200. In at least one embodiment, the IMU sensor 1266 may include, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other types of sensors, without limitation. In at least one embodiment, such as a 6-axis application, the IMU sensor 1266 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as a 9-axis application, the IMU sensor 1266 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0144] In at least one embodiment, the IMU sensor 1266 may be implemented as a small, high-performance GPS-Aided Inertial Navigation System ("GPS / INS") that combines a micro-electro-mechanical system ("MEMS") inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 1266 enables the vehicle 1200 to estimate its heading without requiring input from a magnetic sensor by directly observing velocity changes and correlating them from GPS to the IMU sensor 1266. In at least one embodiment, the IMU sensor 1266 and the GNSS sensor 1258 may be combined into a single integrated unit.
[0145] In at least one embodiment, the vehicle 1200 may include a microphone 1296 installed inside and / or around the vehicle 1200. In at least one embodiment, the microphone 1296 may be used, in particular, for the detection and identification of emergency vehicles.
[0146] In at least one embodiment, the vehicle 1200 may further include any number of camera types, including a stereo camera 1268, a wide-angle camera 1270, an infrared camera 1272, a perimeter camera 1274, a long-range camera 1298, a medium-range camera 1276, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of the vehicle 1200. In at least one embodiment, the types of cameras used will vary depending on the vehicle 1200. In at least one embodiment, any combination of camera types may be used to provide the required coverage area around the vehicle 1200. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, the vehicle 1200 may include six cameras, seven cameras, ten cameras, twelve cameras, or any other number of cameras. The cameras may support Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet®, without being limited to examples. In at least one embodiment, each of the cameras is described in further detail above with respect to Figures 12A and 12B.
[0147] In at least one embodiment, the vehicle 1200 may further include a vibration sensor 1242. The vibration sensor 1242 may measure vibrations of components of the vehicle 1200, such as axles. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 1242 are used, the difference in vibration may be used to determine the amount of friction or slip on the road surface (for example, if there is a difference in vibration between a power-driven axle and a freely rotating axle).
[0148] In at least one embodiment, the vehicle 1200 may include an ADAS system 1238. The ADAS system 1238 may include a SoC in some examples, without limitation. In at least one embodiment, the ADAS system 1238 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control ("ACC") systems, cooperative adaptive cruise control ("CACC") systems, forward crash warning ("FCW") systems, automatic emergency braking ("AEB") systems, lane departure warning ("LDW") systems, lane keep assist ("LKA") systems, blind spot warning ("BSW") systems, rear cross-traffic warning ("RCTW") systems, collision warning ("CW") systems, lane centering ("LC") systems, and / or other systems, features, and / or functions.
[0149] In at least one embodiment, the ACC system may use a RADAR sensor 1260, a LIDAR sensor 1264, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance of vehicle 1200 to the vehicle immediately in front and automatically adjusts the speed of vehicle 1200 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 1200 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications such as LC and CW.
[0150] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a network interface 1224 and / or wireless antenna 1226, either via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, the link may be provided directly by a vehicle-to-vehicle ("V2V") communication link, while the link may be provided indirectly by an infrastructure-to-vehicle ("I2V") communication link. Generally, the concept of V2V communication provides information about the vehicle immediately ahead (e.g., a vehicle in the same lane immediately ahead of vehicle 1200), while the concept of I2V communication provides information about traffic further ahead. In at least one embodiment, the CACC system may include either or both I2V and V2V information sources. In at least one embodiment, having information about the vehicle ahead of vehicle 1200 can further enhance the reliability of the CACC system, potentially leading to smoother traffic flow and reduced congestion on the road.
[0151] In at least one embodiment, the FCW system is designed to alert the driver to hazardous materials so that the driver can take corrective action. In at least one embodiment, the FCW system uses a front camera and / or a radar sensor 1260, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibration, and / or quick brake pulses.
[0152] In at least one embodiment, the AEB system may detect an imminent head-on collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazardous object, the AEB system usually first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the anticipated collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-collision braking.
[0153] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the vehicle 1200 crosses a lane marker. In at least one embodiment, the LDW system does not activate if the driver indicates an intentional lane departure by activating the turn signal. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variation of the LDW system. The LKA system provides steering input or brake control to correct the vehicle 1200 if the vehicle 1200 begins to drift out of its lane.
[0154] In at least one embodiment, the BSW system detects vehicles in the vehicle's blind spot and warns the driver. In at least one embodiment, the BSW system may provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses the turn signal. In at least one embodiment, the BSW system may use a rear camera and / or radar sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration components.
[0155] In at least one embodiment, the RCTW system may provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1200 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear-facing radar sensors 1260, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibration components.
[0156] In at least one embodiment, conventional ADAS systems are prone to producing false positives, which can be annoying and distracting to the driver, but are usually not a major issue. This is because conventional ADAS systems advise the driver, allowing the driver to determine whether a safety-critical condition truly exists and whether appropriate action should be taken. In at least one embodiment, if the results are contradictory, the vehicle 1200 itself determines whether to follow the result from the primary computer (e.g., the first controller 1236) or the result from the secondary computer (e.g., the second controller 1236). For example, in at least one embodiment, the ADAS system 1238 may be a backup and / or secondary computer for providing perceptual information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may run various software for redundancy on hardware components to detect perceptual errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1238 may be provided to a monitoring MCU. In at least one embodiment, if the output from the primary computer and the output from the secondary computer are inconsistent, the monitoring MCU determines how to reconcile the inconsistency to ensure safe operation.
[0157] In at least one embodiment, the primary computer may be configured to provide the monitoring MCU with a reliability score indicating the reliability of the result selected by the primary computer. In at least one embodiment, if the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer, regardless of whether the secondary computer provides contradictory or inconsistent results. In at least one embodiment, if the reliability score does not satisfy the threshold and the primary and secondary computers provide different results (e.g., contradictory), the monitoring MCU may mediate between the computers to determine an appropriate result.
[0158] In at least one embodiment, the monitoring MCU may be configured to run a neural network trained and configured to determine, at least in part, the conditions under which a secondary computer provides a false alarm, based on the output from the primary computer and the output from the secondary computer. In at least one embodiment, the neural network of the monitoring MCU may learn when the output from the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, if the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metallic object that is not actually a hazard, such as a drain grate or manhole cover, which triggers an alarm. In at least one embodiment, if the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when there are cyclists or pedestrians and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or GPU suitable for running the neural network with associated memory. In at least one embodiment, the monitoring MCU may comprise and / or be included as a component of the SoC1204.
[0159] In at least one embodiment, the ADAS system 1238 may include a secondary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and reliability, safety, and performance may be improved by the presence of a neural network in the monitoring MCU. For example, in at least one embodiment, diverse implementations and intentional non-identities increase the overall fault tolerance of the system, particularly to errors caused by the functionality of the software (or software-hardware interface). For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer, and non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU may have greater confidence that the overall result is correct and that the software or hardware bug on the primary computer has not caused a critical error.
[0160] In at least one embodiment, the output of the ADAS system 1238 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1238 is issuing a head-on collision warning due to an object immediately preceding, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have its own trained, and therefore false-detection-reducing, neural network, as described herein.
[0161] In at least one embodiment, the vehicle 1200 may further include an infotainment SoC 1230 (for example, an in-vehicle infotainment system (IVI)). The infotainment system 1230 is illustrated and described as an SoC, but in at least one embodiment, it does not have to be an SoC and may include, without limitation, two or more separate components. In at least one embodiment, the infotainment SoC 1230 may include, without limitation, a combination of hardware and software that can be used to provide the vehicle 1200 with audio (e.g., music, personal digital assistant, navigation commands, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., a navigation system, rear parking assist, wireless data system, vehicle-related information, such as fuel level, total mileage, brake fuel level, oil level, door open / closed, air filter information, etc.). For example, the infotainment SoC 1230 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car computer, in-car entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, head-up display ("HUD"), HMI display 1234, telematics device, control panel (for example, for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 1230 may further be used to provide the vehicle user with (e.g., visual and / or auditory) information such as information from the ADAS system 1238, autonomous driving information such as vehicle operation plans and trajectories, ambient information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0162] In at least one embodiment, the infotainment SoC 1230 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1230 may communicate with other devices, systems, and / or components of the vehicle 1200 via bus 1202 (e.g., CAN bus, Ethernet®, etc.). In at least one embodiment, the infotainment SoC 1230 may be coupled to a monitoring MCU so that the GPU of the infotainment system can perform some self-driving functions when the primary controller 1236 (e.g., the primary and / or backup computer of the vehicle 1200) fails. In at least one embodiment, the infotainment SoC 1230 may put the vehicle 1200 into driver-safe stop mode as described herein.
[0163] In at least one embodiment, the vehicle 1200 may further include an instrument cluster 1232 (e.g., a digital dashboard, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 1232 may include, without limitation, a controller and / or a supercomputer (e.g., a separate controller or supercomputer). In at least one embodiment, the instrument cluster 1232 may include, without limitation, any number and combination of instrument sets such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift lever position indicator, seat belt warning light, parking brake warning light, engine fault light, auxiliary restraint system (e.g., airbag) information, light control, safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1230 and the instrument cluster 1232. In at least one embodiment, the instrument cluster 1232 may be included as part of the infotainment SoC 1230, or vice versa.
[0164] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 12C for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0165] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system in Figure 12C for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0166] Figure 12D is a diagram of a system 1276 for communication between a cloud-based server and the autonomous vehicle 1200 of Figure 12A, according to at least one embodiment. In at least one embodiment, the system 1276 may include, without limitation, a server 1278, a network 1290, and any number and type of vehicles, including the vehicle 1200. The server 1278 may include, without limitation, a plurality of GPUs 1284(A) to 1284(H) (collectively referred to herein as GPU 1284), PCIe switches 1282(A) to 1282(H) (collectively referred to herein as PCIe switch 1282), and / or CPUs 1280(A) to 1280(B) (collectively referred to herein as CPU 1280). The GPU 1284, CPU 1280, and PCIe switch 1282 may be interconnected by high-speed interconnects such as the NVLink interface 1288 and / or PCIe connection 1286 developed by NVIDIA, for example, without limitation. In at least one embodiment, the GPUs 1284 are connected to each other via NVLink and / or NVS switch SoC, and the GPUs 1284 and PCIe switch 1282 are connected via PCIe interconnects. In at least one embodiment, eight GPUs 1284, two CPUs 1280, and four PCIe switches 1282 are illustrated, but this is not limited to. In at least one embodiment, each server 1278 may include any number of GPUs 1284, CPUs 1280, and / or PCIe switches 1282 in any combination, without limitation. For example, in at least one embodiment, each server 1278 may include 8, 16, 32, and / or more GPUs 1284.
[0167] In at least one embodiment, the server 1278 may receive from the vehicle via the network 1290 image data representing images showing unexpected or altered road conditions, such as recently started road construction. In at least one embodiment, the server 1278 may transmit map information 1294, including information about traffic and road conditions, including the neural network 1292, the updated neural network 1292, and / or, without limitation, information about traffic and road conditions, to the vehicle via the network 1290. In at least one embodiment, updates to the map information 1294 may include, without limitation, updates to the HD map 1222, including information about construction sites, potholes, detours, floods, and / or other obstacles. In at least one embodiment, the neural network 1292, the updated neural network 1292, and / or the map information 1294 may be derived from new training and / or experience represented in data received from any number of vehicles in the environment, and / or may be derived at least in part from training performed in a data center (for example, using server 1278 and / or other servers).
[0168] In at least one embodiment, a machine learning model (e.g., a neural network) may be trained using server 1278, at least in part, on training data. The training data may be generated by the vehicle and / or generated by a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or otherwise preprocessed (e.g., if the relevant neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or preprocessed (e.g., if the relevant neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by the vehicle (e.g., transmitted to the vehicle via network 1290 and / or the machine learning model may be used by server 1278 to remotely monitor the vehicle).
[0169] In at least one embodiment, server 1278 may receive data from the vehicle and apply the data to a state-of-the-art real-time neural network to enable real-time intelligent reasoning. In at least one embodiment, server 1278 may include a deep learning supercomputer and / or dedicated AI computer powered by GPU 1284, such as DGX and DGX Station Machines developed by NVIDIA. However, in at least one embodiment, server 1278 may include a deep learning infrastructure using a CPU-powered data center.
[0170] In at least one embodiment, the deep learning infrastructure of server 1278 may be capable of high-speed, real-time inference and may use this capability to evaluate and verify the health of the vehicle 1200's processor, software, and / or associated hardware. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1200, such as a series of images and / or objects located in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by vehicle 1200. If the results do not match and the deep learning infrastructure concludes that the vehicle 1200's AI is malfunctioning, server 1278 may send a signal to vehicle 1200 instructing the vehicle 1200's fail-safe computer to take control, notify the occupants, and complete a safe stopping operation.
[0171] In at least one embodiment, server 1278 may include a GPU 1284 and one or more programmable inference accelerators (e.g., NVIDIA TensorRT3). In at least one embodiment, a server powered by a GPU and inference acceleration can be combined to enable real-time response. In at least one embodiment, a server powered by a CPU, FPGA, and other processors may be used for inference, for example, when performance is not critical. In at least one embodiment, a hardware structure 915 is used to perform one or more embodiments. Details relating to hardware structure (x)915 are provided herein in conjunction with Figures 9A and / or 9B.
[0172] Computer system Figure 13 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or any combination thereof 1300, formed together with a processor which may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, the computer system 1300 may include, without limitation, components such as a processor 1302 for using an execution unit which includes logic for executing algorithms for processing data in accordance with the Disclosure, such as in the embodiments described herein. In at least one embodiment, the computer system 1300 may include a processor such as the PENTIUM® processor family, Xeon®, Itanium®, XScale®, and / or StrongARM®, Intel® Core®, or Intel® Nervana® microprocessors, available from Intel Corporation in Santa Clara, California, but other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, the computer system 1300 may run a version of the WINDOWS® operating system available from Microsoft Corporation in Redmond, Washington, but other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0173] The embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-a-chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0174] In at least one embodiment, the computer system 1300 may include, without limitation, a processor 1302 which may include, without limitation, one or more execution units 1308 for training and / or inferring machine learning models using the techniques described herein. In at least one embodiment, the system 13 is a single-processor desktop or server system, but in another embodiment, the system 13 may be a multi-processor system. In at least one embodiment, the processor 1302 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 1302 may be coupled to a processor bus 1310 which may transmit data signals between the processor 1302 and other components in the computer system 1300.
[0175] In at least one embodiment, the processor 1302 may include, without limitation, a level 1 ("L1") internal cache memory ("cache") 1304. In at least one embodiment, the processor 1302 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to the processor 1302. Other embodiments may include a combination of both internal and external caches, depending on the particular embodiment and necessity. In at least one embodiment, the register file 1306 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, state registers, and instruction pointer registers.
[0176] In at least one embodiment, the processor 1302 also includes an execution unit 1308 which includes, without limitation, logic for performing integer and floating-point arithmetic. The processor 1302 may also include a microcode ("u-code") read-only memory ("ROM") for storing microcode for certain macro instructions. In at least one embodiment, the execution unit 1308 may include logic for handling a packed instruction set 1309. In at least one embodiment, by including the packed instruction set 1309, along with the associated circuitry for executing the instructions, in the instruction set of the general-purpose processor 1302, arithmetic used by many multimedia applications can be performed using the packed data of the general-purpose processor 1302. In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by performing arithmetic on packed data using the full width of the processor's data bus, thereby eliminating the need to transfer smaller units of data across the processor's data bus to perform one or more arithmetic operations on a single data element at a time.
[0177] In at least one embodiment, the execution unit 1308 may also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, the computer system 1300 may include, without limitation, memory 1320. In at least one embodiment, memory 1320 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. Memory 1320 may store instructions 1319 and / or data 1321, which may be represented by data signals executed by the processor 1302.
[0178] In at least one embodiment, a system logic chip may be coupled to a processor bus 1310 and memory 1320. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub ("MCH") 1316, and the processor 1302 may communicate with the MCH 1316 via the processor bus 1310. In at least one embodiment, the MCH 1316 may provide a high-bandwidth memory path 1318 to memory 1320 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1316 may lead data signals between the processor 1302, memory 1320, and other components of the computer system 1300, and may bridge data signals between the processor bus 1310, memory 1320, and system I / O 1322. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH1316 may be coupled to memory 1320 via a high-bandwidth memory path 1318, and the graphics / video card 1312 may be coupled to the MCH1316 via an Accelerated Graphics Port ("AGP") interconnect 1314.
[0179] In at least one embodiment, the computer system 1300 may use a system I / O 1322, which is a proprietary hub interface bus, to connect the MCH 1316 to the I / O controller hub ("ICH") 1330. In at least one embodiment, the ICH 1330 may provide direct connectivity to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1320, the chipset, and the processor 1302. Examples may include, but are not limited to, an audio controller 1329, a firmware hub ("Flash BIOS") 1328, a wireless transceiver 1326, data storage 1324, a legacy I / O controller 1323 including user input and keyboard interfaces, a serial expansion port 1327 such as a Universal Serial Bus ("USB"), and a network controller 1334. The data storage 1324 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0180] In at least one embodiment, Figure 13 shows a system including interconnected hardware devices or “chips,” while in other embodiments, Figure 13 may show an exemplary system-on-a-chip (“SoC”). In at least one embodiment, the devices shown in Figure cc may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of system 1300 are interconnected using compute express link (CXL) interconnects.
[0181] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 13 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0182] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system in Figure 13 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0183] Figure 14 is a block diagram showing an electronic device 1400 for utilizing a processor 1410, according to at least one embodiment. In at least one embodiment, the electronic device 1400 may be, for example, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device, without limiting them.
[0184] In at least one embodiment, the system 1400 may include, without limitation, a processor 1410 communicatively coupled to any number or type of preferred components, peripherals, modules, or devices. In at least one embodiment, the processor 1410 is coupled using a bus or interface such as a 1°C bus, a System Management Bus ("SMBus"), a Low Pin Count (LPC) bus, a Serial Peripheral Interface ("SPI"), a High Definition Audio ("HDA") bus, a Serial Advance Technology Attachment ("SATA") bus, a Universal Serial Bus ("USB") (versions 1, 2, or 3), or a Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 14 shows a system including interconnected hardware devices or “chips,” while in other embodiments, Figure 14 may show an exemplary system-on-a-chip (“SoC”). In at least one embodiment, the devices shown in Figure 14 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of Figure 14 may be interconnected using a Compute Express Link (CXL) interconnect.
[0185] In at least one embodiment, Figure 14 shows a display 1424, a touch screen 1425, a touch pad 1430, a Near Field Communications unit ("NFC") 1445, a sensor hub 1440, a thermal sensor 1446, an Express Chipset ("EC") 1435, a Trusted Platform Module ("TPM") 1438, a BIOS / firmware / flash memory ("BIOS, FW flash") 1422, a DSP 1460, a drive such as a Solid State Disk ("SSD") or Hard Disk Drive ("HDD") 1420, a Wireless Local Area Network Unit ("WLAN") 1450, a Bluetooth unit 1452, and a Wireless Wide Area Network Unit ("WWAN"). The components may include a Wide Area Network unit (1456), a Global Positioning System (GPS) (1455), a camera such as a USB 3.0 camera ("USB 3.0 camera") (1454), or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") (1415) implemented, for example, according to the LPDDR3 standard. Each of these components may be implemented in any preferred manner.
[0186] In at least one embodiment, other components may be communicatively coupled to the processor 1410 via the components described above. In at least one embodiment, the accelerometer 1441, ambient light sensor ("ALS") 1442, compass 1443, and gyroscope 1444 may be communicatively coupled to the sensor hub 1440. In at least one embodiment, the thermal sensor 1439, fan 1437, keyboard 1446, and touchpad 1430 may be communicatively coupled to the EC 1435. In at least one embodiment, the speaker 1463, headphones 1464, and microphone ("mic") 1465 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1464, which may be communicatively coupled to the DSP 1460. In at least one embodiment, the audio unit 1464 may include, for example, an audio coder / decoder ("codec") and a class D amplifier, without limitation. In at least one embodiment, the SIM card ("SIM") 1457 may be communicatively coupled to the WWAN unit 1456. In at least one embodiment, components such as the WLAN unit 1450 and the Bluetooth unit 1452, as well as the WWAN unit 1456, may be implemented in a Next Generation Form Factor ("NGFF").
[0187] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 14 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0188] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system in Figure 14 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0189] Figure 15 shows a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 is configured to carry out various processes and methods described throughout this disclosure.
[0190] In at least one embodiment, the computer system 1500 includes, but is not limited to, at least one central processing unit ("CPU") 1502, which is connected to a communications bus 1510 implemented using any preferred protocol, such as PCI:Peripheral Component Interconnect ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP:Accelerated Graphics Port ("Accelerated Graphics Port"), Hypertransport, or any other bus or point-to-point communications protocol. In at least one embodiment, the computer system 1500 includes, but is not limited to, main memory 1504 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in the main memory 1504, which may take the form of random access memory ("RAM"). In at least one embodiment, the network interface subsystem ("Network Interface") 1522 provides an interface with other computing devices and networks for receiving data from other systems and transmitting data from the computer system 1500 to other systems.
[0191] In at least one embodiment, the computer system 1500 includes, in at least one embodiment without limitation, an input device 1508, a parallel processing system 1512, and a display device 1506, the display device which can be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from the input device 1508, such as a keyboard, mouse, touchpad, or microphone. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.
[0192] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 15 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0193] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system in Figure 15 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0194] Figure 16 shows a computer system 1600 according to at least one embodiment. In at least one embodiment, the computer system 1600 includes, but is not limited to, a computer 1610 and a USB stick 1620. In at least one embodiment, the computer system 1610 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1610 includes, but is not limited to, a server, a cloud instance, a laptop, and a desktop computer.
[0195] In at least one embodiment, the USB stick 1620 includes, but is not limited to, a processing unit 1630, a USB interface 1640, and a USB interface logic 1650. In at least one embodiment, the processing unit 1630 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1630 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1630 comprises an application-specific integrated circuit ("ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1630 is a tensor processing unit ("TPC") optimized to perform machine learning inference operations. In at least one embodiment, the processing core 1630 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.
[0196] In at least one embodiment, the USB interface 1640 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1640 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1640 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1650 may include any amount and type of logic that enables the processing unit 1630 to interface with or to a device (e.g., a computer 1610) via the USB connector 1640.
[0197] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of Figure 16 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0198] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with the system in Figure 16 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0199] Figure 17A shows an exemplary architecture in which multiple GPUs 1710-1713 are communicably coupled to multiple multi-core processors 1705-1706 via high-speed links 1740-1743 (e.g., bus, point-to-point interconnect). In one embodiment, high-speed links 1740-1743 support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. Various interconnection protocols may be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.
[0200] Furthermore, in one embodiment, two or more of the GPUs 1710-1713 may be interconnected via high-speed links 1729-1730, which may be implemented using the same or different protocols / links as those used for high-speed links 1740-1743. Similarly, two or more of the multi-core processors 1705-1706 may be connected via high-speed link 1728, which can be a symmetric multiprocessor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various system components shown in Figure 17A may be implemented using the same protocol / link (for example, via a common interconnection fabric).
[0201] In one embodiment, each multi-core processor 1705-1706 is communicatively coupled to processor memory 1701-1702 via memory interconnects 1726-1727, and each GPU 1710-1713 is communicatively coupled to GPU memory 1720-1723 via GPU memory interconnects 1750-1753. The memory interconnects 1726-1727 and 1750-1753 may utilize the same or different memory access technologies. For example, but not limited to, the processor memories 1701-1702 and GPU memories 1720-1723 may be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or non-volatile memory such as 3D XPoint or Nano-Ram. In one embodiment, (for example, using a two-level memory (2LM) hierarchy), some portions of processor memory 1701-1702 may be volatile memory and other portions may be non-volatile memory.
[0202] As described herein, various processors 1705-1706 and GPUs 1710-1713 may be physically coupled to specific memories 1701-1702 and 1720-1723, respectively, and an integrated memory architecture may be implemented in which the same virtual system address space (also called the “effective address” space) is distributed among the various physical memories. For example, processor memories 1701-1702 may each have a 64GB system memory address space, and GPU memories 1720-1723 may each have a 32GB system memory address space (in this example, a total of 256GB of addressable memory is obtained).
[0203] Figure 17B shows further details of the interconnection between a multi-core processor 1707 and a graphics acceleration module 1746 in one exemplary embodiment. The graphics acceleration module 1746 may include one or more GPU chips integrated on a line card coupled to the processor 1707 via a high-speed link 1740. Alternatively, the graphics acceleration module 1746 may be integrated on the same package or chip as the processor 1707.
[0204] In at least one embodiment, the illustrated processor 1707 includes a plurality of cores 1760A to 1760D, each having a translation lookaside buffer 1761A to 1761D and one or more caches 1762A to 1762D. In at least one embodiment, cores 1760A to 1760D may include various other components (not shown) for executing instructions and processing data. Caches 1762A to 1762D may have Level 1 (L1) and Level 2 (L2) caches. Furthermore, one or more shared caches 1756 may be included in caches 1762A to 1762D and shared by the set of cores 1760A to 1760D. For example, one embodiment of processor 1707 includes 24 cores, each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. The processor 1707 and the graphics acceleration module 1746 are connected to system memory 1714, which may include processor memories 1701-1702 in Figure 17A.
[0205] Coherence is maintained for data and instructions stored in various caches 1762A-1762D, 1756, and system memory 1714 through inter-core communication via coherence bus 1764. For example, each cache may have its own associated cache coherence logic / circuitry to communicate via coherence bus 1764 in response to detecting a read or write to a particular cache line. In one embodiment, a cache snooping protocol is implemented via coherence bus 1764 to monitor cache access.
[0206] In one embodiment, the proxy circuit 1725 connects the graphics acceleration module 1746 to the coherence bus 1764 in a communicative manner, allowing the graphics acceleration module 1746 to participate in the cache coherence protocol as a peer of cores 1760A-1760D. Specifically, interface 1735 provides a connection to the proxy circuit 1725 via a high-speed link 1740 (e.g., PCIe bus, NVLink, etc.), and interface 1737 connects the graphics acceleration module 1746 to link 1740.
[0207] In one embodiment, the accelerator integration circuit 1736 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 1731, 1732, N of the graphics acceleration module 1746. Each of the graphics processing engines 1731, 1732, N may comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1731, 1732, N may comprise different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a bullet engine. In at least one embodiment, the graphics acceleration module 1746 may be a GPU having multiple graphics processing engines 1731, 1732, N, or the graphics processing engines 1731, 1732, N may be individual GPUs integrated on a common package, line card, or chip.
[0208] In one embodiment, the accelerator integration circuit 1736 includes a memory management unit (MMU) 1739 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 1714. The MMU 1739 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In one embodiment, a cache 1738 stores commands and data so that they can be efficiently accessed by graphics processing engines 1731-1732, N. In one embodiment, data stored in cache 1738 and graphics memory 1733-1734, M is kept coherent with core caches 1762A-1762D, 1756, and system memory 1714. As stated above, this may be achieved via a proxy circuit 1725 instead of cache 1738 and memory 1733-1734, M (for example, by sending updates regarding cache line modifications / access in processor caches 1762A-1762D, 1756 to cache 1738 and receiving updates from cache 1738).
[0209] A set of registers 1745 stores context data for threads executed by graphics processing engines 1731-1732, N, and the context management circuit 1748 manages the thread contexts. For example, the context management circuit 1748 may perform save and restore operations to save and restore the contexts of various threads during a context switch (for example, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1748 may store the current register values in a designated area of memory (for example, identified by a context pointer). Then, when returning to the context, the context management circuit 1748 may restore the register values. In one embodiment, the interrupt management circuit 1747 receives and processes interrupts received from system devices.
[0210] In one embodiment, virtual / effective addresses from the graphics processing engine 1731 are translated by the MMU 1739 to real / physical addresses in the system memory 1714. One embodiment of the accelerator integration circuit 1736 supports multiple (e.g., 4, 8, or 16) graphics accelerator modules 1746 and / or other accelerator devices. The graphics accelerator modules 1746 may be dedicated to a single application running on the processor 1707, or they may be shared among multiple applications. In one embodiment, there exists a virtualized graphics execution environment in which resources of graphics processing engines 1731-1732, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices," which are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0211] In at least one embodiment, the accelerator integration circuit 1736 functions as a bridge to the system for the graphics acceleration module 1746 and provides address translation and system memory caching services. Furthermore, the accelerator integration circuit 1736 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1731-1732.
[0212] The hardware resources of the graphics processing engines 1731-1732, N are explicitly mapped to the real address space seen by the host processor 1707, so that any host processor can directly address these resources using effective address values. In one embodiment, one function of the accelerator integration circuit 1736 is to physically isolate the graphics processing engines 1731-1732, N so that they appear as independent units to the system.
[0213] In at least one embodiment, one or more graphics memories 1733-1734, M are each coupled to one of the graphics processing engines 1731-1732, N. The graphics memories 1733-1734, M store instructions and data processed by each of the graphics processing engines 1731-1732, N. The graphics memories 1733-1734, M may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memory such as 3D XPoint or Nano-Ram.
[0214] In one embodiment, a biasing technique is used to reduce data traffic over link 1740, such that the data stored in graphics memory 1733-1734, M is the data that will be most frequently used by the graphics processing engines 1731-1732, N, and preferably the data that will not be used (or at least not frequently used) by cores 1760A-1760D. Similarly, the biasing mechanism attempts to keep the data that the cores need (and therefore preferably not needed by the graphics processing engines 1731-1732, N) in the core caches 1762A-1762D, 1756, and system memory 1714.
[0215] Figure 17C shows another exemplary embodiment in which the accelerator integration circuit 1736 is integrated within the processor 1707. In this embodiment at least, the graphics processing engines 1731-1732, N communicate directly with the accelerator integration circuit 1736 via the high-speed link 1740 through interfaces 1737 and 1735 (in this case, any form of bus or interface protocol can be used). The accelerator integration circuit 1736 may perform the same operations as described with respect to Figure 17B, but may potentially operate at higher throughput given its proximity to the coherence bus 1764 and caches 1762A-1762D, 1756. In one embodiment, the accelerator integration circuit supports different programming models, including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1736 and a programming model controlled by the graphics acceleration module 1746.
[0216] In at least one embodiment, the graphics processing engines 1731-1732,N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can achieve virtualization within a VM / partition by directing other application requests to the graphics processing engines 1731-1732,N.
[0217] In at least one embodiment, the graphics processing engines 1731-1732,N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 1731-1732,N to allow access by each operating system. In a single-partition system without a hypervisor, the graphics processing engines 1731-1732,N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1731-1732,N to provide access to each process or application.
[0218] In at least one embodiment, the graphics acceleration module 1746 or the individual graphics processing engines 1731-1732, N select a process element using a process handle. In one embodiment, the process element is stored in system memory 1714 and is addressable using the effective address-to-actual address translation technique described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the host process context with the graphics processing engines 1731-1732, N (i.e., calling system software to add the process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element link list.
[0219] Figure 17D shows an exemplary accelerator integration slice 1790. As used herein, a “slice” comprises a designated portion of the processing resources of the accelerator integration circuit 1736. The application effective address space 1782 in system memory 1714 stores a process element 1783. In at least one embodiment, the process element 1783 is stored in response to a GPU call 1781 from an application 1780 running on processor 1707. The process element 1783 contains the process state of the corresponding application 1780. A work descriptor (WD) 1784 contained in the process element 1783 may be a single job requested by the application, or it may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1784 is a pointer to a job request queue in the application's address space 1782.
[0220] The graphics acceleration module 1746 and / or individual graphics processing engines 1731-1732, N can be shared by all or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD1784 to the graphics acceleration module 1746 to start jobs in a virtualized environment.
[0221] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns either the graphics acceleration module 1746 or the individual graphics processing engines 1731. Since the graphics acceleration module 1746 is owned by a single process, when the graphics acceleration module 1746 is allocated, the hypervisor initializes the accelerator integration circuit 1736 for the owning partition, and the operating system initializes the accelerator integration circuit 1736 for the owning process.
[0222] During operation, the WD fetch unit 1791 within the accelerator integration slice 1790 fetches the following WD 1784, which includes the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1746. As shown, the data from the WD 1784 is stored in the register 1745 and may be used by the MMU 1739, the interrupt management circuit 1747, and / or the context management circuit 1748. For example, one embodiment of the MMU 1739 includes a segment / page walk circuit for accessing the segment / page table 1786 within the OS virtual address space 1785. The interrupt management circuit 1747 may process the interrupt event 1792 received from the graphics acceleration module 1746. When executing a graphics operation, the effective address 1793 generated by the graphics processing engines 1731 - 1732, N, is translated to a physical address by the MMU 1739.
[0223] In one embodiment, the same set of registers 1745 is replicated for each of the graphics processing engines 1731 - 1732, N, and / or the graphics acceleration module 1746 and may be initialized by the hypervisor or the operating system. Each of these replicated registers may be included in the accelerator integration slice 1790. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0224] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0225] In one embodiment, each WD1784 is specific to a particular graphics acceleration module 1746 and / or graphics processing engines 1731 - 1732, and is specific to N. The WD1784 can accommodate all the information required for the graphics processing engines 1731 - 1732, N to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0226] FIG. 17E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor physical address space 1798 in which a process element list 1799 is stored. The hypervisor physical address space 1798 is accessible via a hypervisor 1796 that virtualizes the graphics acceleration module engine of the operating system 1795.
[0227] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1746. There are two programming models in which the graphics acceleration module 1746 is shared by multiple processes and partitions: time slice sharing and graphics-directed sharing.
[0228] In this model, the system hypervisor 1796 owns the graphics acceleration module 1746 and makes its functionality available to all operating systems 1795. In order for the graphics acceleration module 1746 to support virtualization by the system hypervisor 1796, the graphics acceleration module 1746 may comply with the following: 1) Application job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 1746 must provide a mechanism for saving and restoring context. 2) Application job requests must be guaranteed by the graphics acceleration module 1746 to be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1746 must provide a function to preempt job processing. 3) When the graphics acceleration module 1746 is operating under a specified shared programming model, fairness between processes must be guaranteed.
[0229] In at least one embodiment, application 1780 needs to make a system call to operating system 1795 with the type of graphics acceleration module 1746, a work descriptor (WD), an authorization mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1746 describes the acceleration function desired in the system call. In at least one embodiment, the type of graphics acceleration module 1746 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1746 and can be in the form of a command for graphics acceleration module 1746, an effective address pointer to a user-defined structure, an effective address pointer to a queue of commands, or any other data structure for describing the work performed by graphics acceleration module 1746. In one embodiment, the AMR value is the AMR state for use in the current process. In at least one embodiment, the value passed to the operating system is the same as that of the application setting the AMR. If embodiments of the accelerator integration circuit 1736 and graphics acceleration module 1746 do not support a user privilege mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR to the hypervisor call. The hypervisor 1796 may optionally apply the current privilege mask override register (AMOR) value before placing the AMR into process element 1783. In at least one embodiment, CSRP is one of the registers 1745 that contains the effective address of an area in the application's effective address space 1782 for the graphics acceleration module 1746 to save and restore context state. This pointer is optional if no state needs to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0230] Upon receiving the system call, operating system 1795 may verify that application 1780 is registered and authorized to use graphics acceleration module 1746. Operating system 1795 then calls hypervisor 1796 with the information shown in Table 3. [Table 3]
[0231] Upon receiving a hypervisor call, hypervisor 1796 verifies that operating system 1795 is registered and authorized to use graphics acceleration module 1746. Hypervisor 1796 then places process element 1783 into a process element link list of the corresponding graphics acceleration module 1746 type. The process element may contain the information shown in Table 4. [Table 4]
[0232] In at least one embodiment, the hypervisor initializes the registers 1745 of multiple accelerator integration slices 1790.
[0233] As shown in Figure 17F, in at least one embodiment, integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1701-1702 and GPU memories 1720-1723. In this embodiment, operations performed on GPUs 1710-1713 utilize the same virtual / effective memory address space as accessing processor memories 1701-1702, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1701, a second portion to a second processor memory 1702, a third portion to GPU memory 1720, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes called the effective address space) is distributed across processor memory 1701-1702 and GPU memory 1720-1723, respectively, so that either processor or GPU can access either physical memory, with virtual addresses mapped to physical memory.
[0234] In one embodiment, bias / coherence management circuits 1794A-1794E in one or more of the MMUs 1739A-1739E ensure cache coherence between the cache of one or more host processors (e.g., 1705) and the caches of GPUs 1710-1713, and implement bias techniques to indicate the physical memory where a particular type of data should be stored. Multiple instances of the bias / coherence management circuits 1794A-1794E are shown in Figure 17F, but the bias / coherence circuits may be implemented within the MMU of one or more host processors 1705 and / or within the accelerator integration circuit 1736.
[0235] One embodiment allows GPU-enabled memory 1720-1723 to be mapped as part of system memory and made accessible using shared virtual memory (SVM) techniques without the performance degradation associated with full system cache coherence. In at least one embodiment, the accessibility of GPU-enabled memory 1720-1723 as system memory without cumbersome cache coherence overhead provides a beneficial operating environment for GPU offloading. This configuration allows host processor 1705 software to set operands and access computation results without the overhead of conventional I / O DMA data copying. Such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) access, all of which are less efficient than simple memory access. In at least one embodiment, the ability to access GPU-enabled memory 1720-1723 without cache coherence overhead may be essential for the execution time of offloaded computations. For example, in the case of significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 1710-1713. In at least one embodiment, the efficiency of operand configuration, the efficiency of accessing results, and the efficiency of GPU computation can be helpful in determining the effectiveness of GPU offloading.
[0236] In at least one embodiment, the selection between GPU bias and host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page-granular structure containing 1 or 2 bits per GPU-enabled memory page (i.e., controlled by memory page granularity). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU-enabled memories 1720-1723, with or without a bias cache (for example, to cache frequently used / recently used entries in the bias table) located in or without GPUs 1710-1713. Alternatively, the entire bias table may be maintained within the GPU.
[0237] In at least one embodiment, the bias table entries associated with each access to GPU-biased memory 1720-1723 are accessed before the actual access to GPU memory, resulting in the following behavior: Firstly, local requests from GPUs 1710-1713 to find their pages in the GPU bias are forwarded directly to the corresponding GPU memories 1720-1723. Local requests from GPUs to find their pages in the host bias are forwarded to processor 1705 (for example, via the high-speed link described above). In one embodiment, a request from processor 1705 to find the requested page in the host processor bias completes the request in the same way as a normal memory read. Alternatively, a request directed to a GPU-biased page may be forwarded to GPUs 1710-1713. In at least one embodiment, the GPU may then move the page to the host processor bias if the page is not currently in use. In at least one embodiment, the bias state of a page can be altered by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply by a hardware-based mechanism.
[0238] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL) which calls the GPU's device driver, which sends a message to the GPU (or queues a command descriptor) to change the bias state and, for some transitions, directs the GPU to perform a cache-flushing operation on the host. In at least one embodiment, the cache-flushing operation is used for transitions from a host processor 1705 bias to a GPU bias, but not for transitions in the opposite direction.
[0239] In one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 1705. To access these pages, processor 1705 may request access from GPU 1710, and GPU 1710 may immediately grant access or not. Therefore, to reduce communication between processor 1705 and GPU 1710, it is beneficial to have GPU-biased pages requested by the GPU but not by the host processor 1705, or vice versa.
[0240] Hardware structure 915 is used to carry out one or more embodiments. Details relating to hardware structure (x)915 are provided herein in conjunction with Figures 9A and / or 9B.
[0241] Figure 18 shows exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to those shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0242] Figure 18 is a block diagram showing an exemplary system-on-chip integrated circuit 1800 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1800 includes one or more application processors 1805 (e.g., CPUs), at least one graphics processor 1810, and may further include an image processor 1815 and / or a video processor 1820, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 1800 includes peripherals or bus logic, including a USB controller 1825, a UART controller 1830, an SPI / SDIO controller 1835, and an I.sup.2S / I.sup.2C controller 1840. In at least one embodiment, the integrated circuit 1800 may include a display device 1845 coupled to one or more of the following: a High-Definition Multimedia Interface (HDMI®) controller 1850 and a Mobile Industry Processor Interface (MIPI) display interface 1855. In at least one embodiment, storage may be provided by a flash memory subsystem 1860, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1865 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits may further include an embedded security engine 1870.
[0243] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in integrated circuit 1800 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0244] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in conjunction with integrated circuit diagram 1800 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0245] FIGS. 19A-19B illustrate exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuits may be included, including additional graphics processors / cores, peripheral device interface controllers, or general-purpose processor cores.
[0246] Figures 19A and 19B are block diagrams illustrating exemplary graphics processors for use in a SoC according to embodiments described herein. Figure 19A shows an exemplary graphics processor 1910, a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. Figure 19B shows a further exemplary graphics processor 1940, a system-on-chip integrated circuit that can be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1910 in Figure 19A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1940 in Figure 19B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1910 and 1940 can be a variation of the graphics processor 1810 in Figure 18.
[0247] In at least one embodiment, the graphics processor 1910 includes a vertex processor 1905 and one or more fragment processors 1915A-1915N (e.g., 1915A, 1915B, 1915C, 1915D-1915N-1, and 1915N). In at least one embodiment, the graphics processor 1910 can execute different shader programs via separate logic, thereby optimizing the vertex processor 1905 to perform operations for a vertex shader program, while one or more fragment processors 1915A-1915N perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. In at least one embodiment, the vertex processor 1905 executes the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, the fragment processors 1915A to 1915N use primitive and vertex data generated by the vertex processor 1905 to generate a frame buffer for display on a display device. In at least one embodiment, the fragment processors 1915A to 1915N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform similar operations to pixel shader programs provided in the Direct 3D API.
[0248] In at least one embodiment, the graphics processor 1910 further includes one or more memory management units (MMUs) 1920A-1920B, caches 1925A-1925B, and circuit interconnects 1930A-1930B. In at least one embodiment, one or more MMUs 1920A-1920B include vertex processors 1905 and / or fragment processors 1915A-1915N, providing virtual-to-physical address mappings for the graphics processor 1910, which may reference vertex or image / text data stored in memory, in addition to vertex or image / text data stored in one or more caches 1925A-1925B. In at least one embodiment, one or more MMUs 1920A-1920B may be synchronized with other MMUs in the system, including one or more MMUs associated with one or more application processors 1805, image processor 1815, and / or video processor 1820 in Figure 18, thereby enabling each processor 1805-1820 to participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1930A-1930B enable the graphics processor 1910 to interface with other IP cores in the SoC via the SoC's internal bus or via a direct connection.
[0249] In at least one embodiment, the graphics processor 1940 includes one or more MMUs 1920A-1920B, caches 1925A-1925B, and circuit interconnects 1930A-1930B of the graphics processor 1910 in Figure 19A. In at least one embodiment, the graphics processor 1940 includes one or more shader cores 1955A-1955N (e.g., 1955A, 1955B, 1955C, 1955D, 1955E, 1955F-1955N-1, and 1955N), which provide an integrated shader core architecture in which a single core, or type, or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1940 includes an intercore task manager 1945 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1955A-1955N, and a tiling unit 1958 for accelerating tiling operations for tile-based rendering, where the rendering operation of the scene is subdivided in image space, for example, to take advantage of local space coherence in the scene or to optimize the use of an internal cache.
[0250] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in integrated circuits 19A and / or 19B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0251] In at least one embodiment, the spatially adaptive separable convolutional layer 7 may be used in the integrated circuit 19A and / or 19B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operation, neural network functionality and / or architecture, or neural network use case described herein.
[0252] Figures 20A and 20B show further exemplary graphics processor logic according to the embodiments described herein. Figure 20A shows a graphics core 2000, which in at least one embodiment may be included in the graphics processor 1810 of Figure 18, and in at least one embodiment may be an integrated shader core 1955A to 1955N, as shown in Figure 19B. Figure 20B shows a highly parallel general-purpose graphics processing unit 2030 suitable for deployment in a multi-chip module in at least one embodiment.
[0253] In at least one embodiment, the graphics core 2000 includes a shared instruction cache 2002, a texture unit 2018, and a cache / shared memory 2020, which are common to the execution resources within the graphics core 2000. In at least one embodiment, the graphics core 2000 may include multiple slices 2001A-2001N, or per-core partitions, and the graphics processor may include multiple instances of the graphics core 2000. Slices 2001A-2001N may include support logic including local instruction caches 2004A-2004N, thread schedulers 2006A-2006N, thread dispatchers 2008A-2008N, and register sets 2010A-2010N. In at least one embodiment, slices 2001A to 2001N may include a set of additional function units (AFU2012A to 2012N), floating-point units (FPU2014A to 2014N), integer arithmetic and logical operation units (ALU2016 to 2016N), address calculation units (ACU2013A to 2013N), double-precision floating-point units (DPFPU2015A to 2015N), and matrix processing units (MPU2017A to 2017N).
[0254] In at least one embodiment, the FPU2014A-2014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPU2015A-2015N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU2016A-2016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured to perform mixed-precision operations. In at least one embodiment, the MPU2017A-2017N can also be configured to perform mixed-precision matrix operations, including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPU2017A-2017N can perform various matrix operations to accelerate machine learning application frameworks, including supporting General-Purpose Matrix Multiplication (GEMM) acceleration. In at least one embodiment, AFU2012A~2012N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric function operations (e.g., sine, cosine, etc.).
[0255] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics core 2000 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0256] In at least one embodiment, the spatially adaptive separable convolutional layer 7 may be used with the graphics core 2000 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0257] Figure 20B shows a General Purpose Processing Unit (GPGPU) 2030, which in at least one embodiment can be configured to enable highly parallel computational operations by an array of graphics processing units. In at least one embodiment, the GPGPU 2030 can be directly linked to other instances of the GPGPU 2030 to generate multiple GPU clusters to improve the training speed of deep neural networks. In at least one embodiment, the GPGPU 2030 includes a host interface 2032 to enable connectivity with a host processor. In at least one embodiment, the host interface 2032 is a PCI Express interface. In at least one embodiment, the host interface 2032 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, the GPGPU 2030 receives commands from the host processor and uses a global scheduler 2034 to distribute the execution threads associated with these commands to a set of compute clusters 2036A-2036H. In at least one embodiment, compute clusters 2036A to 2036H share cache memory 2038. In at least one embodiment, cache memory 2038 can act as a high-level cache for cache memory within compute clusters 2036A to 2036H.
[0258] In at least one embodiment, the GPGPU 2030 includes memory 2044A to 2044B coupled to compute clusters 2036A to 2036H via a set of memory controllers 2042A to 2042B. In at least one embodiment, memory 2044A to 2044B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.
[0259] In at least one embodiment, each compute cluster 2036A–2036H includes a set of graphics cores, such as the graphics core 2000 in Figure 20A, which may include multiple types of integer and floating-point logic units capable of performing computational operations at varying precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 2036A–2036H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of floating-point units may be configured to perform 64-bit floating-point operations.
[0260] In at least one embodiment, multiple instances of GPGPU2030 can be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 2036A-2036H differs across embodiments. In at least one embodiment, multiple instances of GPGPU2030 communicate via host interface 2032. In at least one embodiment, GPGPU2030 includes an I / O hub 2039, which couples GPGPU2030 to a GPU link 2040, enabling direct connections to other instances of GPGPU2030. In at least one embodiment, GPU link 2040 is coupled to a dedicated GPU-to-GPU bridge enabling communication and synchronization between multiple instances of GPGPU2030. In at least one embodiment, GPU link 2040 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU2030 are located in separate data processing systems and communicate via a network device accessible through the host interface 2032. In at least one embodiment, the GPU link 2040 can be configured to enable connection to a host processor in addition to, or instead of, the host interface 2032.
[0261] In at least one embodiment, the GPGPU2030 can be configured to train a neural network. In at least one embodiment, the GPGPU2030 can be used within an inference platform. In at least one embodiment, when the GPGPU2030 is used for inference, the GPGPU may include fewer compute clusters 2036A-2036H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with memory 2044A-2044B may differ between the inference configuration and the training configuration, with high-bandwidth memory technology being used in the training configuration. In at least one embodiment, the inference configuration of the GPGPU2030 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more 8-bit integer dot product instructions, which may be used during the inference operation of a deployed neural network.
[0262] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the GPGPU2030 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0263] In at least one embodiment, the spatially adaptive separable convolutional layer 7 may be used in GPGPU2030 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0264] Figure 21 is a block diagram showing a computing system 2100 according to at least one embodiment. In at least one embodiment, the computing system 2100 includes a processing subsystem 2101 having one or more processors 2102 and system memory 2104 communicating via an interconnection path which may include a memory hub 2105. In at least one embodiment, the memory hub 2105 may be a separate component within a chipset component or may be integrated within one or more processors 2102. In at least one embodiment, the memory hub 2105 is coupled to an I / O subsystem 2111 via a communication link 2106. In at least one embodiment, the I / O subsystem 2111 includes an I / O hub 2107 which can enable the computing system 2100 to receive input from one or more input devices 2108. In at least one embodiment, the I / O hub 2107 may enable a display controller, which may be included in one or more processors 2102 and provide output to one or more display devices 2110A. In at least one embodiment, one or more display devices 2110A coupled to the I / O hub 2107 may include local, internal, or embedded display devices.
[0265] In at least one embodiment, the processing subsystem 2101 includes one or more parallel processors 2112 coupled to a memory hub 2105 via a bus or other communication link 2113. In at least one embodiment, the communication link 2113 may be one of any number of communication link technologies or protocols based on standards such as PCI Express, or it may be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 2112 form a computation-intensive parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, one or more parallel processors 2112 form a graphics processing subsystem that can output pixels to one of one or more display devices 2110A coupled via an I / O hub 2107. In at least one embodiment, one or more parallel processors 2112 may also include a display controller and a display interface (not shown) that enable direct connection to one or more display devices 2110B.
[0266] In at least one embodiment, the system storage unit 2114 can be connected to the I / O hub 2107 to provide storage functionality for the computing system 2100. In at least one embodiment, an I / O switch 2116 can be used to provide an interface mechanism for enabling communication between the I / O hub 2107 and other components such as a network adapter 2118 and / or a wireless network adapter 2119, which may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 2120. In at least one embodiment, the network adapter 2118 may be an Ethernet® adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2119 may include one or more other network devices, including Wi-Fi, Bluetooth, Near Field Communication (NFC), or one or more wireless radios.
[0267] In at least one embodiment, the computing system 2100 may include other components not explicitly shown, such as USB or other port connections, optical storage drives, and video capture devices, which may also be connected to the I / O hub 2107. In at least one embodiment, the communication paths interconnecting the various components of Figure 21 may be implemented using any preferred protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link High-Speed Interconnect, or other interconnection protocols.
[0268] In at least one embodiment, one or more parallel processors 2112 incorporate circuits optimized for graphics and video processing, such as video output circuits, to constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2112 incorporate circuits optimized for general-purpose processing. In at least one embodiment, the components of the computing system 2100 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2112, a memory hub 2105, a processor 2102, and an I / O hub 2107 can be integrated into a system-on-a-chip (SoC) integrated circuit. In at least one embodiment, the components of the computing system 2100 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computing system 2100 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0269] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in System Figure 2100 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0270] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in system diagram 2100 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operation, neural network functionality and / or architecture, or neural network use case described herein.
[0271] Processor Figure 22A shows a parallel processor 2200 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2200 may be implemented using one or more integrated circuit devices such as a programmable processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 2200 is a variation of one or more parallel processors 2112 shown in Figure 21 according to an exemplary embodiment.
[0272] In at least one embodiment, the parallel processor 2200 includes a parallel processing unit 2202. In at least one embodiment, the parallel processing unit 2202 includes an I / O unit 2204 that enables communication with other devices, including other instances of the parallel processing unit 2202. In at least one embodiment, the I / O unit 2204 may be directly connected to other devices. In at least one embodiment, the I / O unit 2204 is connected to other devices via the use of a hub or switch interface, such as a memory hub 2105. In at least one embodiment, the connection between the memory hub 2105 and the I / O unit 2204 forms a communication link 2113. In at least one embodiment, the I / O unit 2204 is connected to a host interface 2206 and a memory crossbar 2216, where the host interface 2206 receives commands targeting the execution of processing operations and the memory crossbar 2216 receives commands targeting the execution of memory operations.
[0273] In at least one embodiment, when the host interface 2206 receives a command buffer via the I / O unit 2204, the host interface 2206 can direct work operations to the front end 2208 to execute these commands. In at least one embodiment, the front end 2208 is coupled to a scheduler 2210, which is configured to distribute commands or other work items to a processing cluster array 2212. In at least one embodiment, the scheduler 2210 ensures that the processing cluster array 2212 is properly configured and enabled before tasks are distributed to the processing cluster array 2212. In at least one embodiment, the scheduler 2210 is implemented via firmware logic running on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2210 can be configured to perform complex scheduling and work distribution operations with both coarse and fine granularity, enabling rapid preemption and context switching of threads running on the processing array 2212. In at least one embodiment, the host software can prove scheduling workloads on the processing array 2212 through one of several graphics processing doorbells. In at least one embodiment, the scheduler 2210 logic within the microcontroller, including the scheduler 2210, can then automatically distribute the workload across the processing cluster array 2212.
[0274] In at least one embodiment, the processing cluster array 2212 may contain up to "N" processing clusters (e.g., cluster 2214A, cluster 2214B to cluster 2214N). In at least one embodiment, each cluster 2214A to 2214N of the processing cluster array 2212 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 2210 may allocate work to clusters 2214A to 2214N of the processing cluster array 2212 using various scheduling and / or work distribution algorithms, which may differ depending on the workload arising for each type of program or computation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 2210 or partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 2212. In at least one embodiment, different clusters 2214A to 2214N of the processing cluster array 2212 may be allocated to process different types of programs or to execute different types of computations.
[0275] In at least one embodiment, the processing cluster array 2212 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2212 can be configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, the processing cluster array 2212 can include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations including physical operations, and performing data transformations.
[0276] In at least one embodiment, the processing cluster array 2212 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2212 may include additional logic to support the performance of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as mosaic logic and other vertex processing logic. In at least one embodiment, the processing cluster array 2212 can be configured to run graphics processing-related shader programs, including but not limited to vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2202 can transfer data from system memory via I / O unit 2204 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2222) during processing and then written back to system memory.
[0277] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2202, the scheduler 2210 can be configured to divide the processing workload into tasks of roughly equal size in order to better distribute the graphics processing operations among the multiple clusters 2214A to 2214N of the processing cluster array 2212. In at least one embodiment, parts of the processing cluster array 2212 can be configured to perform different types of processing. For example, in at least one embodiment, to generate and display a rendered image, a first part may be configured to perform vertex shading and topology generation, a second part may be configured to perform mosaic and geometry shading, and a third part may be configured to perform pixel shading or other screen-space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2214A to 2214N may be stored in a buffer so that the intermediate data can be transmitted between the clusters 2214A to 2214N for further processing.
[0278] In at least one embodiment, the processing cluster array 2212 may receive processing tasks to be executed via a scheduler 2210, which receives commands defining the processing tasks from a front-end 2208. In at least one embodiment, the processing task may include an index of the data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters, and commands defining how the data should be processed (e.g., which program to run). In at least one embodiment, the scheduler 2210 may be configured to fetch the index corresponding to the task, or to receive the index from the front-end 2208. In at least one embodiment, the front-end 2208 may be configured to ensure that the processing cluster array 2212 is configured to be in a valid state before the workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.
[0279] In at least one embodiment, each of one or more instances of the parallel processing unit 2202 can be coupled to a parallel processor memory 2222. In at least one embodiment, the parallel processor memory 2222 can be accessed via a memory crossbar 2216, which can receive memory requests from the processing cluster array 2212 and the I / O unit 2204. In at least one embodiment, the memory crossbar 2216 can access the parallel processor memory 2222 via a memory interface 2218. In at least one embodiment, the memory interface 2218 may include a plurality of partition units (for example, partition unit 2220A, partition units 2220B to 2220N), each of which can be coupled to a portion of the parallel processor memory 2222 (for example, a memory unit). In at least one embodiment, the number of partition units 2220A to 2220N is configured to be equal to the number of memory units, so that the first partition unit 2220A has a corresponding first memory unit 2224A, the second partition unit 2220B has a corresponding memory unit 2224B, and the Nth partition unit 2220N has a corresponding Nth memory unit 2224N. In at least one embodiment, the number of partition units 2220A to 2220N does not have to be equal to the number of memory devices.
[0280] In at least one embodiment, memory units 2224A to 2224N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2224A to 2224N may also include, but not limited to, high-bandwidth memory (HBM) and 3D stacked memory. In at least one embodiment, to efficiently use the available bandwidth of parallel processor memory 2222, render targets such as frame buffers or texture maps may be stored across memory units 2224A to 2224N, allowing partition units 2220A to 2220N to write each portion of the render target in parallel. In at least one embodiment, local instances of parallel processor memory 2222 may be excluded to favor an integrated memory design using system memory and local cache memory together.
[0281] In at least one embodiment, any one of the clusters 2214A to 2214N of the processing cluster array 2212 can process data that will be written to any of the memory units 2224A to 2224N in the parallel processor memory 2222. In at least one embodiment, the memory crossbar 2216 can be configured to forward the output of each cluster 2214A to 2214N to any partition unit 2220A to 2220N, or another cluster 2214A to 2214N, which can perform further processing operations on the output. In at least one embodiment, each cluster 2214A to 2214N can communicate with the memory interface 2218 through the memory crossbar 2216 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2216 has a connection to a memory interface 2218 for communicating with the I / O unit 2204, and a connection to a local instance of the parallel processor memory 2222, enabling processing units in different processing clusters 2214A to 2214N to communicate with system memory or other memory not local to the parallel processing unit 2202. In at least one embodiment, the memory crossbar 2216 can use virtual channels to separate traffic streams between the clusters 2214A to 2214N and the partition units 2220A to 2220N.
[0282] In at least one embodiment, multiple instances of the parallel processing unit 2202 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2202 can be configured to interact with each other, even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2202 may include higher-precision floating-point units than other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2202 or parallel processor 2200 can be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0283] Figure 22B is a block diagram of partition unit 2220 according to at least one embodiment. In at least one embodiment, partition unit 2220 is one instance of partition units 2220A to 2220N in Figure 22A. In at least one embodiment, partition unit 2220 includes an L2 cache 2221, a frame buffer interface 2225, and a ROP 2226 (raster operations unit). The L2 cache 2221 is a read / write cache configured to perform load and store operations received from the memory crossbar 2216 and the ROP 2226. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 2221 to the frame buffer interface 2225 for processing. In at least one embodiment, updates are also sent to the frame buffer via the frame buffer interface 2225 for processing. In at least one embodiment, the frame buffer interface 2225 interfaces with one of the memory units of the parallel processor memory, such as memory units 2224A to 2224N (for example, in the parallel processor memory 2222) shown in Figure 22.
[0284] In at least one embodiment, ROP2226 is a processing unit that performs raster operations such as stenciling, z-testing, and blending. In at least one embodiment, ROP2226 then outputs the processed graphics data stored in graphics memory. In at least one embodiment, ROP2226 includes compression logic for compressing depth or color data written to memory and for decompressing depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The type of compression performed by ROP2226 may be changed based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0285] In at least one embodiment, ROP2226 is located within each processing cluster (for example, clusters 2214A to 2214N in Figure 22) rather than within the partition unit 2220. In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via the memory crossbar 2216. In at least one embodiment, the processed graphics data may be displayed on a display device, such as one of the one or more display devices 2110 in Figure 21, routed for further processing by processor 2102, or routed for further processing by one of the processing entities in the parallel processor 2200 in Figure 22A.
[0286] Figure 22C is a block diagram of a processing cluster 2214 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 2214A to 2214N in Figure 22. In at least one embodiment, the processing cluster 2214 may be configured to run a large number of threads in parallel, where the term “thread” refers to an instance of a particular program running on a particular set of input data. In at least one embodiment, a single-instruction, multiple-data (SIMD) instruction issuing technique is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, a single-instruction, multiple-thread (SIMT) technique is used to support the parallel execution of a large number of threads in a globally synchronized manner, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0287] In at least one embodiment, the operation of the processing cluster 2214 can be controlled via a pipeline manager 2232 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2232 receives instructions from the scheduler 2210 in Figure 22 and manages the execution of these instructions via the graphics multiprocessor 2234 and / or texture unit 2236. In at least one embodiment, the graphics multiprocessor 2234 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2214. In at least one embodiment, one or more instances of the graphics multiprocessor 2234 may be included within the processing cluster 2214. In at least one embodiment, the graphics multiprocessor 2234 can process data, and a data crossbar 2240 may be used to distribute the processed data to one of several possible destinations, including other shader units. In at least one embodiment, the pipeline manager 2232 can facilitate the distribution of processed data by specifying the destination of the processed data to be distributed through the data crossbar 2240.
[0288] In at least one embodiment, each graphics multiprocessor 2234 within the processing cluster 2214 may include an identical set of function execution logic (e.g., arithmetic logic units, load / store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner, allowing new instructions to be issued before the previous instruction is completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be performed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0289] In at least one embodiment, instructions sent to the processing cluster 2214 constitute a thread. In at least one embodiment, a set of threads running across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program for different input data. In at least one embodiment, each thread in a thread group can be assigned to a different processing engine in the graphics multiprocessor 2234. In at least one embodiment, a thread group may contain fewer threads than the number of processing engines in the graphics multiprocessor 2234. In at least one embodiment, if a thread group contains fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycle in which the thread group is being processed. In at least one embodiment, a thread group may also contain more threads than the number of processing engines in the graphics multiprocessor 2234. In at least one embodiment, if a thread group contains more threads than the number of processing engines in the graphics multiprocessor 2234, processing can be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups can run simultaneously on the graphics multiprocessor 2234.
[0290] In at least one embodiment, the graphics multiprocessor 2234 includes internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2234 can abandon its internal cache and use cache memory in the processing cluster 2214 (e.g., L1 cache 2248). In at least one embodiment, each graphics multiprocessor 2234 may also access L2 cache in a partition unit (e.g., partition units 2220A-2220N in Figure 22), and these caches may be shared among all processing clusters 2214 and used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2234 may also access off-chip global memory, which may include one or more of the local parallel processor memory and / or system memory. In at least one embodiment, any memory outside the parallel processing unit 2202 may be used as global memory. In at least one embodiment, the processing cluster 2214 includes multiple instances of a graphics multiprocessor 2234, which can share common instructions and data, which may be stored in an L1 cache 2248.
[0291] In at least one embodiment, each processing cluster 2214 may include an MMU 2245 (memory management unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2245 may be located within the memory interface 2218 in Figure 22. In at least one embodiment, the MMU 2245 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles (tiling is described in detail) and optionally cache line indices. In at least one embodiment, the MMU 2245 may include an address translation lookaside buffer (TLB) or cache, which may be located within the graphics multiprocessor 2234 or the L1 cache, or within the processing cluster 2214. In at least one embodiment, physical addresses are processed to distribute surface data access locally, enabling efficient interleaving of requests between partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0292] In at least one embodiment, each graphics multiprocessor 2234 may be coupled to a texture unit 2236 to configure a processing cluster 2214 so that texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data, are performed. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2234 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2234 outputs processed tasks to a data crossbar 2240 to provide processed tasks to another processing cluster 2214 for further processing, or stores processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 2216. In at least one embodiment, a pre-ROP 2242 (pre-raster arithmetic unit) is configured to receive data from a graphics multiprocessor 2234 and direct the data to an ROP unit, which may be located within a partition unit (for example, partition units 2220A-2220N in Figure 22A) as described herein. In at least one embodiment, the pre-ROP 2242 unit can perform color blending optimization, organize pixel color data, and perform address translation.
[0293] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics processing cluster 2214 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0294] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in the graphics processing cluster 2214 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0295] Figure 22D shows a graphics multiprocessor 2234 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 2234 is coupled with a pipeline manager 2232 of a processing cluster 2214. In at least one embodiment, the graphics multiprocessor 2234 has an execution pipeline that includes, but is not limited to, an instruction cache 2252, an instruction unit 2254, an address mapping unit 2256, a register file 2258, one or more general-purpose graphics processing unit (GPGPU) cores 2262, and one or more load / store units 2266. The GPGPU cores 2262 and the load / store units 2266 are coupled to cache memory 2272 and shared memory 2270 via a memory and cache interconnect 2268.
[0296] In at least one embodiment, the instruction cache 2252 receives a stream of instructions to be executed from the pipeline manager 2232. In at least one embodiment, the instructions are cached in the instruction cache 2252 and dispatched to be executed by the instruction unit 2254. In at least one embodiment, the instruction unit 2254 may dispatch the instructions as a thread group (e.g., a warp), where each thread in the thread group is assigned to a different execution unit within the GPGPU core 2262. In at least one embodiment, the instructions can access either the local, shared, or global address space by specifying an address in the unified address space. In at least one embodiment, an address mapping unit 2256 can be used to translate an address in the unified address space to a separate memory address accessible by the load / store unit 2266.
[0297] In at least one embodiment, the register file 2258 provides a set of registers to the functional units of the graphics multiprocessor 2234. In at least one embodiment, the register file 2258 provides temporary storage for operands connected to the data paths of the functional units of the graphics multiprocessor 2234 (e.g., GPGPU core 2262, load / store unit 2266). In at least one embodiment, the register file 2258 is divided among the functional units such that each functional unit is allocated a dedicated portion of the register file 2258. In one embodiment, the register file 2258 is divided among different warps being executed by the graphics multiprocessor 2234.
[0298] In at least one embodiment, each GPGPU core 2262 may include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of the graphics multiprocessor 2234. The GPGPU cores 2262 may have similar architectures or different architectures. In at least one embodiment, a first part of the GPGPU core 2262 includes a single-precision FPU and an integer ALU, and a second part of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may perform IEEE 754-2008 standard floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 2234 may further include one or more fixed-function units or special-function units for performing specific functions such as rectangular copying or pixel blending operations. In at least one embodiment, one or more GPGPU cores may also include fixed or special-function logic.
[0299] In at least one embodiment, the GPGPU core 2262 includes SIMD logic that can execute a single instruction for multiple data sets. In at least one embodiment, the GPGPU core 2262 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core may be generated at compile time by the shader compiler, or they may be automatically generated when running a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logic unit.
[0300] In at least one embodiment, the memory and cache interconnect 2268 is an interconnect network connecting each functional unit of the graphics multiprocessor 2234 to the register file 2258 and shared memory 2270. In at least one embodiment, the memory and cache interconnect 2268 is a crossbar interconnect that allows the load / store unit 2266 to perform load and store operations between the shared memory 2270 and the register file 2258. In at least one embodiment, the register file 2258 can operate at the same frequency as the GPGPU core 2262, and therefore data transfer between the GPGPU core 2262 and the register file 2258 is very low latency. In at least one embodiment, the shared memory 2270 can be used to enable communication between threads running in the functional units within the graphics multiprocessor 2234. In at least one embodiment, the cache memory 2272 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2236. In at least one embodiment, the shared memory 2270 can also be used as a program-managed cache. In at least one embodiment, a thread running on the GPGPU core 2262 can programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 2272.
[0301] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated as a core in the same package or chip and communicatively coupled to the core via an internal (i.e., internal to the package or chip) processor bus / interconnection. In at least one embodiment, regardless of how the GPU is connected, the processor core may allocate work to such GPUs in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0302] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics multiprocessor 2234 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0303] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in a graphics multiprocessor 2234 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0304] Figure 23 shows a multi-GPU computing system 2300 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2300 may include a processor 2302 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2306A-D via a host interface switch 2304. In at least one embodiment, the host interface switch 2304 is a PCI Express switch device that couples the processor 2302 to a PCI Express bus, through which the processor 2302 can communicate with the GPGPUs 2306A-D. The GPGPUs 2306A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2316. In at least one embodiment, the GPU-to-GPU links 2316 are connected to each of the GPGPUs 2306A-D via dedicated GPU links. In at least one embodiment, the P2P GPU link 2316 enables direct communication between each of the GPGPUs 2306A-D without requiring communication via the host interface bus 2304 to which the processor 2302 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 2316, the host interface bus 2304 is kept available to allow access to system memory or to communicate with other instances of the multi-GPU computing system 2300, for example, via one or more network devices. In at least one embodiment, the GPGPUs 2306A-D are connected to the processor 2302 via the host interface switch 2304, and in at least one embodiment, the processor 2302 can be directly connected to the GPGPUs 2306A-D, including direct support for the P2P GPU link 2316.
[0305] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in a multi-GPU computing system 2300 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0306] In at least one embodiment, the spatially adaptive separable convolutional layer 7 may be used in a multi-GPU computing system 2300 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0307] Figure 24 is a block diagram of a graphics processor 2400 according to at least one embodiment. In at least one embodiment, the graphics processor 2400 includes a ring interconnect 2402, a pipeline front end 2404, a media engine 2437, and graphics cores 2480A to 2480N. In at least one embodiment, the ring interconnect 2402 connects the graphics processor 2400 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2400 is one of a number of processors integrated within a multi-core processing system.
[0308] In at least one embodiment, the graphics processor 2400 receives batches of commands via a ring interconnect 2402. In at least one embodiment, incoming commands are interpreted by a command streamer 2403 on a pipeline front end 2404. In at least one embodiment, the graphics processor 2400 includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 2480A-2480N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2403 feeds the commands to a geometry pipeline 2436. In at least one embodiment, for at least some media processing commands, the command streamer 2403 feeds the commands to a video front end 2434, which is coupled to a media engine 2437. In at least one embodiment, the media engine 2437 includes a Video Quality Engine (VQE) 2430 for post-processing of video and images, and a multi-format encoding / decoding (MFX) 2433 engine for encoding and decoding hardware-accelerated media data. In at least one embodiment, the geometry pipeline 2436 and the media engine 2437 each generate execution threads for thread execution resources provided by at least one graphics core 2480A.
[0309] In at least one embodiment, the graphics processor 2400 includes a scalable thread execution resource featuring modular cores 2480A-2480N (sometimes called core slices), each modular core 2480A-2480N having multiple sub-cores 2450A-550N, 2460A-2460N (sometimes called core sub-slices). In at least one embodiment, the graphics processor 2400 can have any number of graphics cores 2480A-2480N. In at least one embodiment, the graphics processor 2400 includes a graphics core 2480A having at least a first sub-core 2450A and a second sub-core 2460A. In at least one embodiment, the graphics processor 2400 is a low-power processor having a single sub-core (e.g., 2450A). In at least one embodiment, the graphics processor 2400 includes a plurality of graphics cores 2480A to 2480N, each of which includes a first set of subcores 2450A to 2450N and a second set of subcores 2460A to 2460N. In at least one embodiment, each of the first subcores 2450A to 2450N includes at least a first set of execution units 2452A to 2452N and media / texture samplers 2454A to 2454N. In at least one embodiment, each of the second subcores 2460A to 2460N includes at least a second set of execution units 2462A to 2462N and samplers 2464A to 2464N. In at least one embodiment, each sub-core 2450A-2450N, 2460A-2460N shares a set of shared resources 2470A-2470N. In at least one embodiment, the shared resources include shared cache memory and pixel operation logic.
[0310] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics processor 2400 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0311] In at least one embodiment, the spatially adaptive separable convolutional layer 7 may be used in the graphics processor 2400 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0312] Figure 25 is a block diagram showing the microarchitecture of a processor 2500, which may include logic circuits for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2500 may execute instructions including x86 instructions, AMR instructions, and special instructions for application-specific integrated circuits (ASICs). In at least one embodiment, the processor 2510 may include registers for storing packed data, such as 64-bit wide MMX® registers in a microprocessor enabled by MMX technology, provided by Intel Corporation, Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point formats, may operate with packed data elements accompanied by Single Instruction Multiple Data ("SIMD") and Streaming SIMD Extensions ("SSE") instructions. In at least one embodiment, 128-bit wide XMM registers relating to SSE2, SSE3, SSE4, AVX, or higher technologies (collectively referred to as "SSEx") may hold operands of such packed data. In at least one embodiment, the processor 2510 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0313] In at least one embodiment, the processor 2500 includes an in-order front-end ("front-end") 2501 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front-end 2501 may include several units. In at least one embodiment, an instruction prefetcher 2526 fetches instructions from memory and supplies them to an instruction decoder 2528, which decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2528 decodes the received instruction into one or more operations called "microinstructions" or "microoperations" (also called "microops" or "uops") that the machine can execute. In at least one embodiment, the instruction decoder 2528 may parse the instruction into opcodes and corresponding data, as well as control fields, which are used by the microarchitecture to perform the operation according to at least one embodiment. In at least one embodiment, the trace cache 2530 may assemble the decoded uops into a program-order sequence or trace in the uop queue 2534 so that they can be executed. In at least one embodiment, when the trace cache 2530 encounters a complex instruction, the microcode ROM 2532 provides the uops necessary to complete the operation.
[0314] In at least one embodiment, some instructions can be converted into a single micro-ops, while others require several micro-ops to complete the entire operation. In at least one embodiment, if five or more micro-ops are required to complete an instruction, the instruction decoder 2528 may access the microcode ROM 2532 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-ops so that it can be processed by the instruction decoder 2528. In at least one embodiment, if many micro-ops are required to complete the operation, the instruction may be stored in the microcode ROM 2532. In at least one embodiment, the trace cache 2530 refers to the entry-point programmable logic array ("PLA") to determine the correct microinstruction pointer for reading the microcode sequence in order to complete one or more instructions from the microcode ROM 2532 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2532 has finished sequencing microops for instructions, the machine's front-end 2501 may resume fetching microops from the trace cache 2530.
[0315] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2503 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has a large number of buffers to smooth the flow of instructions and change their order, optimizing performance as instructions are pipelined and scheduled for execution. The out-of-order execution engine 2503 includes, but is not limited to, an allocator / register renamer 2540, a memory uop queue 2542, an integer / floating-point uop queue 2544, a memory scheduler 2546, a fast scheduler 2502, a slow / general-purpose floating-point scheduler ("slow / general-purpose FP: floating-point scheduler") 2504, and a simple floating-point scheduler ("simple FP scheduler") 2506. In at least one embodiment, the fast scheduler 2502, the slow / general-purpose floating-point scheduler 2504, and the simple floating-point scheduler 2506 are collectively referred to herein as “uop schedulers 2502, 2504, and 2506”. The allocator / register renamer 2540 allocates the machine buffers and resources that each uop needs to run. In at least one embodiment, the allocator / register renamer 2540 renames logical registers upon entry into the register file. In at least one embodiment, the allocator / register renamer 2540 also allocates the entries of each uop to one of two uop queues, namely the memory uop queue 2542 for memory operations and the integer / floating-point uop queue 2544 for non-memory operations, prior to the memory scheduler 2546 and the uop schedulers 2502, 2504, 2506. In at least one embodiment, the uop schedulers 2502, 2504, 2506 determine when uops are ready to execute based on whether the sources for their dependent input register operands are ready and whether the execution resources required by the uop to complete their operations are available.In at least one embodiment, the high-speed scheduler 2502 of at least one embodiment may schedule every half of the main clock cycle, while the slow / general-purpose floating-point scheduler 2504 and the simple floating-point scheduler 2506 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2502, 2504, and 2506 arbitrate dispatch ports to schedule uops so that they can be executed.
[0316] In at least one embodiment, execution block b11 includes, without limitation, an integer register file / bypass network 2508, a floating-point register file / bypass network ("FP register file / bypass network") 2510, address generation units ("AGUs") 2512 and 2514, fast arithmetic logic units (ALUs) ("fast ALUs") 2516 and 2518, a slow arithmetic logic unit ("slow ALU") 2520, a floating-point ALU ("FP") 2522, and a floating-point move unit ("FP move") 2524. In at least one embodiment, the integer register file / bypass network 2508 and the floating-point register file / bypass network 2510 are also referred to herein as "register files 2508, 2510". In at least one embodiment, AGU2512 and 2514, high-speed ALU2516 and 2518, low-speed ALU2520, floating-point ALU2522, and floating-point movement unit 2524 are also referred to herein as “execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524”. In at least one embodiment, execution block b11 may include, without limitation, any number and type of register files (including zero), bypass networks, address generation units, and execution units in any combination.
[0317] In at least one embodiment, register files 2508, 2510 may be located between the uop schedulers 2502, 2504, 2506 and the execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524. In at least one embodiment, the integer register file / bypass network 2508 performs integer arithmetic. In at least one embodiment, the floating-point register file / bypass network 2510 performs floating-point arithmetic. In at least one embodiment, each of the register files 2508, 2510 may include, without limitation, a bypass network which may bypass or transfer newly completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, the register files 2508, 2510 may communicate data with each other. In at least one embodiment, the integer register file / bypass network 2508 may include, without limitation, two separate register files: one register file for low-order 32-bit data and a second register file for high-order 32-bit data. In at least one embodiment, since floating-point instructions typically have operands of 64 to 128 bits in width, the floating-point register file / bypass network 2510 may include, without limitation, 128-bit wide entries.
[0318] In at least one embodiment, execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524 may execute instructions. In at least one embodiment, register files 2508 and 2510 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2500 may include, without limitation, any number and combination of execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524. In at least one embodiment, floating-point ALU 2522 and floating-point movement unit 2524 may perform floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, the floating-point ALU2522 may include, without limitation, 64-bit floating-point dividers and perform division, square root, and the remaining micro-operations. In at least one embodiment, instructions involving floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to high-speed ALU2516, 2518. In at least one embodiment, high-speed ALU2516, 2518 may perform high-speed operations with an effective latency of half a clock cycle. In at least one embodiment, the slow ALU2520 may include, without limitation, integer execution hardware for long-latency types of operations such as multipliers, shifts, flag logic, and branching, so that most complex integer operations are passed to the slow ALU2520. In at least one embodiment, memory load / store operations may be performed by AGUS2512, 2514. In at least one embodiment, the high-speed ALU2516, high-speed ALU2518, and low-speed ALU2520 may perform integer arithmetic with 64-bit data operands. In at least one embodiment, the high-speed ALU2516, high-speed ALU2518, and low-speed ALU2520 may be implemented to support a variety of data bit sizes, including 16, 32, 128, 256, and so on. In at least one embodiment, the floating-point ALU2522 and floating-point movement unit 2524 may be implemented to support a wide range of operands with various bit widths.In at least one embodiment, the floating-point ALU 2522 and the floating-point movement unit 2524 may operate in conjunction with SIMD and multimedia instructions as 128-bit wide packed data operands.
[0319] In at least one embodiment, the uop schedulers 2502, 2504, and 2506 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, uops may be scheduled and executed speculatively in processor 2500, so processor 2500 may also include logic for handling memory misses. In at least one embodiment, if a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed the scheduler with temporarily inaccurate data. In at least one embodiment, a replay mechanism tracks and redelivers instructions that use inaccurate data. In at least one embodiment, dependent operations may need to be replayed, while independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0320] In at least one embodiment, the term “register” may refer to an onboard processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be accessible from outside the processor (from the programmer’s perspective). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuitry within the processor using any number of different techniques, such as dedicated physical registers, dynamically allocated physical registers using register renaming, or a combination of dedicated and dynamically allocated physical registers. In at least one embodiment, an integer register stores 32-bit integer data. The register file in at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0321] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into EXE block 2511 and other memories or registers, whether illustrated or not. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in EXE block 2511. Furthermore, weight parameters may be stored in on-chip or off-chip memories and / or registers (illustrated or not illustrated) that constitute the ALUs of EXE block 2511 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0322] Figure 26 shows a deep learning application processor 2600 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2600 uses instructions that cause the deep learning application processor 2600 to perform some or all of the processes and techniques described throughout this disclosure when executed by the deep learning application processor 2600. In at least one embodiment, the deep learning application processor 2600 is an application-specific integrated circuit (ASIC). In at least one embodiment, the application processor 2600 performs a matrix multiplication operation which is "hardwired" to hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2600 includes, without limitation, processing clusters 2610(1) to 2610(12), inter-chip links ("ICL") 2620(1) to 2620(12), inter-chip controllers ("ICC") 2630(1) to 2630(2), high-bandwidth memory second generation ("HBM2") 2640(1) to 2640(4), memory controllers ("Mem Ctrlrs") 2642(1) to 2642(4), and high-bandwidth memory physical layer ("HBM"). The processor includes PHYs 2644(1) to 2644(4), a Management-Controller Central Processing Unit ("Management-Controller CPU") 2650, serial peripheral interfaces, inter-integrated and general-purpose input / output blocks ("SPI, I2C, GPIO") 2660, peripheral component interconnect express controllers and direct memory access blocks ("PCIe controllers and DMA") 2670, and a 16-lane peripheral component interconnect express port ("PCI Express x16") 2680.
[0323] In at least one embodiment, the processing cluster 2610 may perform deep learning operations, including inference or prediction operations, based on weight parameters calculated using one or more training techniques, including the techniques described herein. In at least one embodiment, each processing cluster 2610 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2600 may include any number and type of processing clusters 2600. In at least one embodiment, the inter-chip link 2620 is bidirectional. In at least one embodiment, the inter-chip link 2620 and the inter-chip controller 2630 enable multiple deep learning application processors 2600 to exchange information, including activation information obtained as a result of running one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2600 may include any number and type of ICL2620 and ICC2630 (including zero).
[0324] In at least one embodiment, the HBM2 2640 provides a total of 32 gigabytes (GB) of memory. The HBM2 2640(i) is associated with both the memory controller 2642(i) and the HBM PHY2644(i). In at least one embodiment, any number of HBM2 2640s may provide any type and total amount of high-bandwidth memory and may be associated with any number and type of memory controllers 2642 and HBM PHY2644 (including zero). In at least one embodiment, SPI, I2C, GPIO2660, PCIe controller and DMA2670, and / or PCIe2680 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically feasible manner.
[0325] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2600. In at least one embodiment, the deep learning application processor 2600 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the deep learning application processor 2600. In at least one embodiment, the processor 2600 may be used to perform one or more use cases of neural networks described herein.
[0326] In at least one embodiment, a spatially adaptive separable convolutional layer 7 may be used in the deep learning application processor 2600 for inference or prediction operations, at least in part, based on weight parameters computed using the neural network training operations, neural network functions and / or architecture, or neural network use cases described herein.
[0327] Figure 27 is a block diagram of a neuromorphic processor 2700 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2700 may receive one or more inputs from an external source. In at least one embodiment, these inputs may be transmitted to one or more neurons 2702 within the neuromorphic processor 2700. In at least one embodiment, the neurons 2702 and their components may be implemented using circuits or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2700 may include thousands or millions of instances of neurons 2702, but any preferred number of neurons 2702 may be used. In at least one embodiment, each instance of a neuron 2702 may include a neuron input 2704 and a neuron output 2706. In at least one embodiment, neuron 2702 may generate an output, which may be sent to the input of another instance of neuron 2702. For example, in at least one embodiment, neuron input 2704 and neuron output 2706 may be interconnected via synapse 2708.
[0328] In at least one embodiment, neuron 2702 and synapse 2708 may be interconnected so that the neuromorphic processor 2700 can operate to process or analyze information received by the neuromorphic processor 2700. In at least one embodiment, neuron 2702 may transmit an output pulse (or “fire” or “spike”) when the input received via neuron input 2704 exceeds a threshold. In at least one embodiment, neuron 2702 may sum or integrate the signals received at neuron input 2704. For example, in at least one embodiment, neuron 2702 may be implemented as a leaky integrate-and-fire neuron, where neuron 2702 may generate an output (or “fire”) using a transfer function such as a sigmoid function or a threshold function when the sum (called “membrane potential”) exceeds a threshold. In at least one embodiment, the leakage integral firing neuron may sum the signals received at neuron input 2704 to obtain the membrane potential, or it may apply a collapse factor (or leak) to reduce the membrane potential. In at least one embodiment, the leakage integral firing neuron may fire if multiple input signals are received at neuron input 2704 quickly enough to exceed a threshold (i.e., before the membrane potential collapse is too small to fire). In at least one embodiment, neuron 2702 may be implemented using a circuit or logic that receives the input, integrates the input to obtain the membrane potential, and collapses the membrane potential. In at least one embodiment, the input may be averaged, or any other suitable transfer function may be used. Furthermore, in at least one embodiment, neuron 2702 may include, without limitation, a comparator circuit or logic that generates an output spike at neuron output 2706 when the result of applying the transfer function to neuron 2704 exceeds a threshold. In at least one embodiment, once neuron 2702 fires, it may ignore previously received input information, for example, by resetting the membrane potential to 0 or another suitable default value.In at least one embodiment, once the membrane potential is reset to 0, neuron 2702 may resume normal operation after a suitable period (or refractory period).
[0329] In at least one embodiment, neurons 2702 may be interconnected through synapses 2708. In at least one embodiment, synapses 2708 may operate to transmit a signal from the output of a first neuron 2702 to the input of a second neuron 2702. In at least one embodiment, neuron 2702 may transmit information through two or more instances of synapses 2708. In at least one embodiment, one or more instances of neuron output 2706 may be connected through an instance of synapse 2708 to an instance of neuron input 2704 of the same neuron 2702. In at least one embodiment, an instance of neuron 2702 that generates an output to be transmitted through an instance of synapse 2708 may be called a “presynaptic neuron” with respect to that instance of synapse 2708. In at least one embodiment, an instance of neuron 2702 that receives an input to be transmitted through an instance of synapse 2708 may be called a “postsynaptic neuron” with respect to that instance of synapse 2708. In at least one embodiment, an instance of neuron 2702 may receive input from one or more instances of synapse 2708 and transmit output through one or more instances of synapse 2708, so that a single instance of neuron 2702 may be both a "presynaptic neuron" and a "postsynaptic neuron" with respect to various instances of synapse 2708.
[0330] In at least one embodiment, the neurons 2702 may be organized into one or more layers. Each instance of neuron 2702 may have one neuron output 2706 that can fan out to one or more neuron inputs 2704 through one or more synapses 2708. In at least one embodiment, the neuron output 2706 of a neuron 2702 in a first layer 2710 may be connected to the neuron input 2704 of a neuron 2702 in a second layer 2712. In at least one embodiment, layer 2710 may be called a “feedforward layer”. In at least one embodiment, each instance of neuron 2702 in an instance of the first layer 2710 may fan out to each instance of neuron 2702 in the second layer 2712. In at least one embodiment, the first layer 2710 may be called a “fully connected feedforward layer”. In at least one embodiment, each instance of neuron 2702 in an instance of the second layer 2712 may be fanned out to fewer instances of neuron 2702 in the third layer 2714 than the total number of instances of neuron 2702 in the third layer 2714. In at least one embodiment, the second layer 2712 may be called a “loosely connected feedforward layer”. In at least one embodiment, neurons 2702 in the second layer 2712 may be fanned out to neurons 2702 in multiple other layers, including neurons 2702 in the (same) second layer 2712. In at least one embodiment, the second layer 2712 may be called a “regression layer”. The neuromorphic processor 2700 may include, without limitation, any preferred combination of regression layers and feedforward layers, including, without limitation, both loosely connected feedforward layers and fully connected feedforward layers.
[0331] In at least one embodiment, the neuromorphic processor 2700 may include, without limitation, a reconfigurable interconnect architecture or dedicated hardwired interconnect for connecting synapses 2708 to neurons 2702. In at least one embodiment, the neuromorphic processor 2700 may include, without limitation, circuits or logic that, based on the neural network topology and the fan-in / fan-out of neurons, allow synapses to be allocated to different neurons 2702 as needed. For example, in at least one embodiment, synapse 2708 may be connected to neuron 2702 using an interconnect fabric such as a network-on-a-chip or using a dedicated connection. In at least one embodiment, the synaptic interconnect and its components may be implemented using circuits or logic.
[0332] Figure 28 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2800 includes one or more processors 2802 and one or more graphics processors 2808, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 2802 or processor cores 2807. In at least one embodiment, system 2800 is a processing platform embedded in a system-on-a-chip (SoC) integrated circuit for use in a mobile device, portable device, or embedded device.
[0333] In at least one embodiment, system 2800 may include, or be incorporated into, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable game console, or an online game console. In at least one embodiment, system 2800 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2800 may also include, be coupled to, or be integrated into wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2800 is a television or set-top box device having one or more processors 2802 and a graphical interface produced by one or more graphics processors 2808.
[0334] In at least one embodiment, each of the one or more processors 2802 includes one or more processor cores 2807 for processing instructions that, when executed, perform actions for the system and user software. In at least one embodiment, each of the one or more processor cores 2807 is configured to process a particular instruction set 2809. In at least one embodiment, the instruction set 2809 may facilitate computing via composite instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction words (VLIW). In at least one embodiment, each of the processor cores 2807 may process a different instruction set 2809, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, the processor cores 2807 may also include other processing devices, such as a digital signal processor (DSP).
[0335] In at least one embodiment, the processor 2802 includes a cache memory 2804. In at least one embodiment, the processor 2802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of the processor 2802. In at least one embodiment, the processor 2802 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which may be shared among processor cores 2807 using known cache coherence techniques. In at least one embodiment, the processor 2802 further includes a register file 2806, which may contain different types of registers for storing different types of data (e.g., integer registers, floating-point registers, state registers, and instruction pointer registers). In at least one embodiment, the register file 2806 may contain general-purpose registers or other registers.
[0336] In at least one embodiment, one or more processors 2802 are coupled to one or more interface buses 2810 to transmit communication signals, such as addresses, data, or control signals, between the processors 2802 and other components in the system 2800. In at least one embodiment, the interface bus 2810 may be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface 2810 is not limited to a DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2802 includes an integrated memory controller 2816 and a platform controller hub 2830. In at least one embodiment, the memory controller 2816 facilitates communication between memory devices and other components of the system 2800, while the platform controller hub (PCH) 2830 provides connectivity to I / O devices via a local I / O bus.
[0337] In at least one embodiment, the memory device 2820 may be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase-change memory device, or any other memory device having performance suitable for acting as process memory. In at least one embodiment, the memory device 2820 may operate as system memory for system 2800 and store data 2822 and instructions 2821 for use by one or more processors 2802 when executing applications or processes. In at least one embodiment, the memory controller 2816 may also be coupled to an optional external graphics processor 2812, which may communicate with one or more graphics processors 2808 within processor 2802 to perform graphics and media operations. In at least one embodiment, a display device 2811 may be connected to processor 2802. In at least one embodiment, the display device 2811 may include one or more internal display devices, such as a mobile electronic device or laptop device, or external display devices that are attached via a display interface (e.g., a DisplayPort). In at least one embodiment, the display device 2811 may include a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.
[0338] In at least one embodiment, the platform controller hub 2830 allows peripheral devices to connect to the memory device 2820 and processor 2802 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2846, a network controller 2834, a firmware interface 2828, a wireless transceiver 2826, a touch sensor 2825, and a data storage device 2824 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2824 may be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a peripheral component interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2825 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2826 may be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2828 may enable communication with system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2834 may enable network connectivity to a wired network. In at least one embodiment, a high-performance network controller (not shown) may be coupled to the interface bus 2810. In at least one embodiment, the audio controller 2846 is a multi-channel high-definition audio controller.In at least one embodiment, system 2800 includes an optional legacy I / O controller 2840 for connecting legacy (e.g., Personal System 2 (PS / 2)) devices to the system. In at least one embodiment, platform controller hub 2830 can also connect to one or more connectable input devices of Universal Serial Bus (USB) controller 2842, such as a keyboard and mouse combination 2843, a camera 2844, or other USB input devices.
[0339] In at least one embodiment, instances of the memory controller 2816 and the platform controller hub 2830 may be integrated with a separate external graphics processor, such as an external graphics processor 2812. In at least one embodiment, the platform controller hub 2830 and / or the memory controller 2816 may be external to one or more processors 2802. For example, in at least one embodiment, the system 2800 may include an external memory controller 2816 and a platform controller hub 2830, which may be configured as a memory controller hub and peripheral controller hub within a system chipset that communicate with the processor 2802.
[0340] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the graphics processor 2800. For example, in at least one embodiment, the training and / or inference technique described herein may use one or more of the ALUs embodied in the 3D pipeline 2812. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or 9B. In at least one embodiment, weight parameters may be stored in on-chip or off-chip memory and / or registers (illustrated or not illustrated) that constitute the ALUs of the graphics processor 2800 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0341] Figure 29 is a block diagram of a processor 2900 having one or more processor cores 2902A-2902N, an integrated memory controller 2914, and an integrated graphics processor 2908, according to at least one embodiment. In at least one embodiment, the processor 2900 may include no more than a number of additional cores, including additional cores 2902N represented by dashed rectangles. In at least one embodiment, each of the processor cores 2902A-2902N includes one or more internal cache units 2904A-2904N. In at least one embodiment, each processor core also has access to one or more shared cache units 2906.
[0342] In at least one embodiment, the internal cache units 2904A-2904N and the shared cache unit 2906 represent a cache memory hierarchy within the processor 2900. In at least one embodiment, the cache memory units 2904A-2904N may include at least one level of instruction and data cache within each processor core, as well as one or more levels of shared intermediate level caches such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, where the highest level cache prior to external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2906 and 2904A-2904N.
[0343] In at least one embodiment, the processor 2900 may also include one or more bus controller units 2916 and a set of system agent cores 2910. In at least one embodiment, one or more bus controller units 2916 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2910 provides management functions for various processor components. In at least one embodiment, the system agent core 2910 includes one or more integrated memory controllers 2914 for managing access to various external memory devices (not shown).
[0344] In at least one embodiment, one or more of the processor cores 2902A to 2902N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2910 includes components for coordinating and operating the cores 2902A to 2902N during multithreaded processing. In at least one embodiment, the system agent core 2910 may further include a power control unit (PCU) which includes logic and components for coordinating the power states of one or more of the processor cores 2902A to 2902N and the graphics processor 2908.
[0345] In at least one embodiment, the processor 2900 further includes a graphics processor 2908 for performing graphics processing operations. In at least one embodiment, the graphics processor 2908 is coupled to a system agent core 2910 which includes a shared cache unit 2906 and one or more integrated memory controllers 2914. In at least one embodiment, the system agent core 2910 also includes a display controller 2911 for directing the output of the graphics processor to one or more coupled displays. In at least one embodiment, the display controller 2911 may also be a separate module coupled to the graphics processor 2908 via at least one interconnection, or it may be integrated within the graphics processor 2908.
[0346] In at least one embodiment, a ring-based interconnection unit 2912 is used to connect the internal components of the processor 2900. In at least one embodiment, alternative interconnection units such as point-to-point interconnection, switch interconnection, or other techniques may be used. In at least one embodiment, the graphics processor 2908 is connected to the ring interconnection 2912 via an I / O link 2913.
[0347] In at least one embodiment, I / O link 2913 represents at least one of a variety of I / O interconnects, including on-package I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 2918, such as eDRAM modules. In at least one embodiment, each of the processor cores 2902A to 2902N and the graphics processor 2908 use the embedded memory module 2918 as a shared last-level cache.
[0348] In at least one embodiment, the processor cores 2902A to 2902N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 2902A to 2902N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 2902A to 2902N execute a common instruction set, while one or more other cores of the processor cores 2902A to 2902N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, the processor cores 2902A to 2902N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more cores with lower power consumption. In at least one embodiment, the processor 2900 can be implemented on one or more chips or as an SoC integrated circuit.
[0349] The inference and / or training logic 915 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 915 are provided herein in conjunction with Figures 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the graphics processor 2910. For example, in at least one embodiment, the training and / or inference technique described herein may use one or more ALUs embodied in the 3D pipeline 2812, graphics core 2915A, shared function logic 2916, graphics core 2915B, shared function logic 2920, or other logics in Figure 29. Furthermore, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in Figure 9A or Figure 9B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that constitute the ALU of the graphics processor 2910 for performing one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0350] Figure 30 is a block diagram of a graphics processor 3000, which may be a separate graphics processing unit or a graphics processor integrated with multiple processing cores. In at least one embodiment, the graphics processor 3000 communicates with its registers using commands placed in memory via an I / O interface mapped to memory. In at least one embodiment, the graphics processor 3000 includes a memory interface 3014 for accessing memory. In at least one embodiment, the memory interface 3014 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.
[0351] In at least one embodiment, the graphics processor 3000 also includes a display controller 3002 for driving display output data toward the display device 3020. In at least one embodiment, the display controller 3002 includes one or more overlapping planes for the display device 3020 and hardware for compositing multilayer video or user interface elements. In at least one embodiment, the display device 3020 may be an internal or external display device. In at least one embodiment, the display device 3020 is a head-mounted display device such as a virtual reality (VR) display device or an augmented reality (AR) display device. In at least one embodiment, the graphics processor 3000 includes a video codec engine 3006 for encoding, decoding, or transcoding media to, from, or between, one or more media encoding formats, including but not limited to, Video Expert Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4AVC, and Joint Photographic Expert Group (JPEG) formats such as SMPTE 421M / VC-1 and JPEG, and Motion JPEG (MJPEG) format.
[0352] In at least one embodiment, the graphics processor 3000 includes a block image transfer (BLIT) engine 3004 for performing two-dimensional (2D) rasterizer operations, such as bit boundary block transfers. However, in at least one embodiment, 2D graphics operations are performed using one or more components of the graphics processing engine (GPE) 3010. In at least one embodiment, the GPE 3010 is a compute engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0353] In at least one embodiment, the GPE 3010 includes a 3D pipeline 3012 for performing 3D operations, such as rendering 3D images and scenes using processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 3012 includes programmable, fixed function elements that perform various tasks on the 3D / media subsystem 3015 and / or spawn execution threads. While media operations can be performed using the 3D pipeline 3012, in at least one embodiment, the GPE 3010 also includes a media pipeline 3016 used to perform media operations such as video post-processing and image enhancement.
[0354] In at least one embodiment, the media pipeline 3016 includes a fixed function or programmable logic unit for performing one or more special media operations, such as video decoding acceleration, video deinterlacing, and encoding acceleration, instead of or on behalf of the video codec engine 3006. In at least one embodiment, the media pipeline 3016 further i...
Claims
1. A processor comprising one or more circuits, wherein the circuit is Obtain first information used in one or more convolutional layers of one or more neural networks, The one or more neural networks are made to generate one or more feature maps in which different parts of the feature map represent different features in one or more images, using the one or more convolutional layers and the first information. The one or more neural networks are then made to generate second information using the one or more feature maps that have been generated. This is a circuit for that purpose. The aforementioned circuit is A processor that generates one or more feature maps, at least in part, based on combining values from two or more filters in different ways in different parts of the feature maps.
2. The processor according to claim 1, wherein the first information includes a matrix corresponding to one or more aspects, features, or objects of the one or more images.
3. The processor according to claim 1, wherein the second information includes the probability of classification corresponding to the different features.
4. The processor according to claim 1, wherein the one or more feature maps are generated at least partially based on two or more filters used in part of the feature maps.
5. The processor according to claim 1, wherein the one or more feature maps are generated at least partially based on the arrangement of a plurality of filters in each part of the feature map.
6. The processor according to claim 1, wherein the one or more feature maps are generated at least in part on applying filters of different sizes to the first information.
7. The processor according to claim 1, wherein the one or more feature maps are generated at least in part on the basis of aggregating a first value from a first filter and a second value from a second filter to generate a third value.
8. A step of obtaining first information to be used in one or more convolutional layers of one or more neural networks, The steps include causing one or more neural networks to generate one or more feature maps in which different parts of the feature maps represent different features in one or more images, using the one or more convolutional layers and the first information, The steps include causing one or more neural networks to generate second information using the generated one or more feature maps, and In a method including, A method further comprising generating one or more feature maps, at least in part, based on combining values from two or more filters in different ways in different parts of the feature maps.
9. The method according to claim 8, wherein the step of causing one or more neural networks to generate a second information using the generated one or more feature maps includes performing depth-based convolution and point-based convolution.
10. The method according to claim 8, wherein the step of causing one or more neural networks to generate second information using the generated one or more feature maps includes calculating a first value of a first filter and a second value of a second filter for a portion of the feature map, the first value and the second value indicating whether the first filter or the second filter detected a feature in the portion.
11. The method according to claim 8, further comprising aligning two or more filters on different parts of the feature map to generate a selection weight map.
12. The method according to claim 8, further comprising generating a selection weight map using the one or more feature maps in order to generate the second information.
13. A system having one or more processors comprising one or more circuits, the circuit comprises, Obtain first information used in one or more convolutional layers of one or more neural networks, The one or more neural networks are made to generate one or more feature maps in which different parts of the feature map represent different features in one or more images, using the one or more convolutional layers and the first information. The one or more neural networks are then made to generate second information using the one or more feature maps that have been generated. This is a circuit for that purpose. The aforementioned circuit is A system for generating one or more feature maps, at least partially based on combining values from two or more filters in different ways in different parts of the feature maps.
14. The system according to claim 13, wherein the one or more convolutional layers include depth-based convolution and point-based convolution.
15. The system according to claim 13, wherein the one or more neural networks include a convolutional neural network having layers for feature learning and classification.
16. The system according to claim 13, wherein the one or more feature maps are at least partially generated based on depth-unit convolution to obtain information about the features of the one or more images.
17. The system according to claim 13, wherein the one or more feature maps are generated at least in part on point-level convolutions that combine information obtained from depth-level convolutions.
18. The system according to claim 13, wherein the one or more feature maps are used to generate a selection weight map.
19. The system according to claim 13, wherein the one or more feature maps are generated at least in part on a softmax function applied after a convolution in units of depth.
Citation Information
Patent Citations
Image multi-scale information extraction method capable of being integrated into neural network architecture and application
CN109934241A
Authenticating objects using machine learning from microscopic differences
JP2017520864A
Material characteristic estimation device and material characteristic estimation method
JP2019012037A
Serial Convolutional Neural Network
JP2019515376A
Cascaded convolutional neural network
US20190042892A1