5g-NR multi-cell software framework
The software PHY library optimizes 5G-NR network processing by implementing a PHY pipeline that leverages parallel processing units to efficiently manage resource usage, addressing the challenge of increased user and cell densities in 5G-NR networks.
Patent Information
- Application Number
- JP2025035346
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-25
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-10
AI Technical Summary
The increasing demand for 5G-NR network processing resources is strained by the substantial memory and time resources required for multi-cell physical layer (PHY) processing in fifth generation (5G) new radio (NR) networks, especially as more users and cells are added to a 5G-NR base station.
A software PHY library implements a PHY pipeline using novel techniques, leveraging parallel processing units (PPUs) like GPUs to perform parallelized multi-user and multi-cell 5G-NR PHY operations, optimizing resource usage through hierarchical and temporal data organization, and batched parameter compilation.
This approach enhances the efficiency of 5G-NR network processing by reducing resource consumption while supporting increased user and cell densities, thereby improving the scalability and performance of 5G-NR networks.
Smart Images

Figure 2025087829000001_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to processing resources used to perform and facilitate multi-cell physical layer (PHY) processing in a fifth generation (5G) new radio (NR) network. For example, at least one embodiment uses a software PHY library that implements a PHY pipeline according to various novel techniques described herein, and relates to a processor or computing system used to perform parallelized multi-user and / or multi-cell 5G-NR PHY operations.
Background Art
[0002] Processing of the operations of the physical layer (PHY) in a fifth generation (5G) new radio (NR) communication network may use substantial memory resources, time resources, or other computing resources. This resource usage increases with the addition of further users or computing cells to a 5G-NR base station in a 5G-NR network. The increasing prevalence of wireless communication devices and further the implementation of 5G-NR network infrastructure are increasing the demand for 5G-NR network processing resources.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 14C
Figure 14D
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19A
Figure 19B
Figure 19C
Figure 19D
Figure 19E
Figure 19F
Figure 20
Figure 21A
Figure 21B
Figure 22A
Figure 22B
Figure 23
Figure 24A
Figure 24B
Figure 24C
Figure 24D
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33A
Figure 33B
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Best Mode for Carrying Out the Invention
[0004] FIG. 1 is a block diagram showing a fifth generation (5G) new radio (NR) physical layer (PHY) pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software fifth generation (5G) new radio (NR) library according to at least one embodiment. In at least one embodiment, 5G-NR is a network communication standard for wireless access technology, 5G indicates that it is the fifth generation of wireless technology, and new radio indicates a new radio interface and wireless access technology for cellular communication networks. In at least one embodiment, a 5G-NR network includes a base station that processes communication information from cells such as a tower having a plurality of connected user equipment (UE) such as mobile phones. To process information from multiple cells, in one embodiment, each base station performs various processing operations further described herein. In at least one embodiment, the processing operations in a 5G-NR network are classified into a hierarchy including different layers such as a first layer (L1) 106 or physical layer (PHY) that performs lower level operations and a second layer (L2) 102 that performs higher level operations.
[0005] In at least one embodiment, the second layer (L2) 102 is a logical organization of hardware operations, software operations, and higher level operations performed by a base station comprising that hardware and software. In at least one embodiment, the higher level operations are 5G-NR computing operations that depend on or otherwise require interaction with the lower level operations performed at L1 / PHY 106 on the base station. In at least one embodiment, L2 102 includes one or more computing operations that facilitate 5G-NR network communication. In at least one embodiment, the operations of L2 102 prepare data and / or other information for the computing operations performed by L1 / PHY 106. L2 102 uses the L2-L1 interface 104 to call or otherwise interact with one or more computing operations performed by L1 / PHY 106.
[0006] In at least one embodiment, the L2-L1 interface 104 is hardware and / or software instructions that, when executed, provide an interface between the L2 102 and the L1 106 in a 5G-NR network. In at least one embodiment, the L2-L1 interface 104 is an application programming interface (API). In at least one embodiment, the L2-L1 interface 104 is a hardware interface. In at least one embodiment, the L2-L1 interface 104 is another interface that facilitates the interaction between the L2 102 and the L1 106 in a 5G-NR network and the transfer of data and / or other information.
[0007] In at least one embodiment, the first layer (L1) 106 or the physical layer (PHY) is a logical organization of hardware operations, software operations, and low-level operations executed by a base station comprising such hardware and software. In at least one embodiment, the L1 / PHY 106 is implemented in hardware. In at least one embodiment, the L1 / PHY 106 is implemented by one or more software libraries. In at least one embodiment, the L1 / PHY 106 is implemented by one or more software libraries to achieve acceleration of the operations of the L1 / PHY 106 using one or more parallel processing units (PPUs) such as a graphics processing unit (GPU).
[0008] In at least one embodiment, L1 / PHY106 is mapped to physical channels, such as uplink and downlink. In at least one embodiment, each channel performs functions for transmitting and receiving data. In at least one embodiment, each channel performs functions for transmitting and receiving control information, cell discovery, and initial access. In at least one embodiment, uplink and downlink signal processing components for L1 / PHY106, such as software operations implemented in the physical layer (PHY) library 116, provide a signal processing pipeline having signal processing blocks for channel-specific operations of each L1 / PHY106. In at least one embodiment, for a downlink channel in which the baseband unit (BBU) is performing transmitter communication, the signal processing blocks are determined by the NR standard specifications of the 3rd Generation Partnership Project (3GPP (registered trademark)). In at least one embodiment, for an uplink channel in which the BBU is performing receiver communication, the signal processing blocks are implementation-specific and may include various components that perform operations further described herein.
[0009] In at least one embodiment, between L1 / PHY106 and L2 102 of a 5G-NR communication network or any other type of communication network, the L2-L1 interface 104 provides an interface between signal processing operations of the L1 / PHY106 layer, such as operations implemented by the software PHY library 116, such as cuPHY, cuBB, or any other software 5th generation 5G-NR library, to upper layers such as the L2 102 of the BBU. In at least one embodiment, the L2-L1 interface 104 operates as an interface between components of L1 / PHY106, such as the PHY library 116 that performs signal processing operations, and upper layers (such as L2 102) of the 3rd Generation Partnership Project (3GPP) protocol stack.
[0010] In at least one embodiment, the L2-L1 interface interacts with or otherwise communicates with the physical layer (PHY) library driver 112. In at least one embodiment, the PHY library driver is software instructions that, when executed, orchestrate and / or call one or more signal processing operations of the one or more L1 / PHY 106 performed by the physical layer (PHY) library 116. In one embodiment, to execute or cause the execution of the signal processing operations of the PHY library driver 112, the above PHY library driver 112 implements a physical layer (PHY) library driver interface 110. In at least one embodiment, the PHY library driver interface 110 is software instructions that, when executed, provide an application programming interface (API) for calling one or more signal processing operations of the one or more L1 / PHY 106 that are executed by the PHY library driver 112 and at least partially performed by the PHY library 116.
[0011] In at least one embodiment, the physical layer (PHY) library 116 is software instructions that, when executed, perform various signal processing operations according to a 5G-NR protocol stack such as a 3GPP protocol stack. In at least one embodiment, the PHY library 116 comprises or otherwise provides a physical layer (PHY) library interface 114. In at least one embodiment, the PHY library interface 114 is software instructions that, when executed, provide an API to the PHY library 116 for performing various signal processing operations performed by the above PHY library 116. In at least one embodiment, the PHY library interface 114 is an API. In at least one embodiment, the PHY library interface 114 provides an API compliant with a standard such as the FAPI interface of the Small Cell Forum. In at least one embodiment, the PHY library interface 114 provides an API that is proprietary.
[0012] In at least one embodiment, a PHY library 116, such as a cuPHY, cuBB, or any other software fifth generation 5G-NR library, manages software executed on one or more parallel processing units (PPUs), such as a graphics processing unit (GPU) further described herein. In at least one embodiment, the PHY library 116 manages a software kernel or segment of software instructions that perform one or more specific operations that implement signal processing operations for the L1 / PHY 106 of a 5G-NR wireless communication system, and the software kernel is executed by one or more PPUs, such as a GPU, as further described herein.
[0013] In at least one embodiment, one or more PPUs, such as a GPU, perform all functions, such as all operations of the L1 / PHY 106 in a signal processing pipeline. In at least one embodiment, one or more PPUs, such as a GPU, accelerate specific operations of the L1 / PHY 106 or blocks of L1 / PHY 106 operations in a signal processing pipeline. In at least one embodiment, the PHY library driver 112 and / or the PHY library 106 provide software, such as one or more interfaces 110, 114, or other APIs, to manage the interaction of the PPUs. In at least one embodiment, the PHY library 116 transmits or otherwise provides one or more parameters and / or descriptors, as described below, to one or more software kernels executed by one or more PPUs, such as a GPU, to perform one or more signal processing operations of the L1 / PHY 106. In at least one embodiment, the PHY library 116 and / or the PHY library driver 112 manage the output from one or more software kernels executed by one or more PPUs, such as a GPU.
[0014] In at least one embodiment, the PHY library 116 performs and / or executes signal processing operations for the transmission of data and / or other information in the L1 / PHY 106 of a 5G-NR network. In at least one embodiment, the signal processing operations performed and / or executed by the PHY library 116 include a physical uplink shared channel (PUSCH). In at least one embodiment, the PUSCH in 5G-NR is designated to carry multiplexed control information and user application data, as further described herein. In at least one embodiment, the signal processing operations performed and / or executed by the PHY library 116 include a physical downlink shared channel (PDSCH). In at least one embodiment, the PDSCH carries user data and upper layer signaling, as further described herein.
[0015] In at least one embodiment, the signal processing operations performed and / or executed by the PHY library 116 include components for the transmission of control information. In at least one embodiment, the control information transmission components include a physical downlink control channel (PDCCH) and a physical uplink control channel (PUCCH). In at least one embodiment, the PDCCH and PUCCH carry information regarding the transmission format and resource allocation related to the PDSCH and PUSCH channels, as further described below.
[0016] In at least one embodiment, the control information transmission component includes an L1 / PHY106 reference signal. In at least one embodiment, the reference signal of L1 / PHY106 in the control information transmission component is a demodulation reference signal (DMRS), a phase-tracking reference signal (PTRS), a sounding reference signal (SRS), and a channel-state information reference signal (CSI-RS). In at least one embodiment, the DMRS is used to estimate a radio channel for demodulation, as further described herein. In at least one embodiment, the PTRS is utilized to enable compensation for oscillator phase noise, as further described herein. In at least one embodiment, the SRS and CSI-RS are utilized to perform channel-state information (CSI) measurements for scheduling, beamforming, and / or link adaptation, as further described herein.
[0017] In at least one embodiment, the signal processing operations implemented and / or executed by the PHY library 116 provide components for initial access and cell discovery. In at least one embodiment, cell discovery includes at least a physical random access channel (PRACH) and a physical broadcast channel (PBCH), as further described herein. In at least one embodiment, a synchronization signal block (SS Block) may be broadcast to select a serving cell, as further described herein.
[0018] In at least one embodiment, the signal processing operations performed and / or executed by the PHY library 116 include Low PHY functions that perform basic operations on 5G-NR signals, as further described herein. In at least one embodiment, the Low PHY functions include fast fourier transform (FFT) and inverse fast fourier transform (IFFT). In at least one embodiment, the FFT and IFFT convert frequency-based signal information into time-based data for processing and vice versa, as further described herein. In at least one embodiment, the Low PHY functions include cyclic prefix (CP) insertion and removal. In at least one embodiment, the CP insertion and removal facilitate the execution of FFT and IFFT operations that perform convolution, as further described herein. In at least one embodiment, the Low PHY functions include transmit beamforming (Tx Beamforming) and receive beamforming (Rx Beamforming). In at least one embodiment, beamforming is a signal filtering technique used in 5G-NR and other wireless networks, as further described herein. In at least one embodiment, one or more antennas in one or more radio units, such as a cell, transmit and receive signal data from one or more user equipment (UEs), such as a mobile phone and / or other wireless communication-enabled devices.
[0019] In at least one embodiment, the signal processing operations performed and / or executed by the PHY library 116 include other L1 / PHY106 operations as further described herein and / or required by 3GPP standards or other 5G-NR standard documents. To interact with the L2 102 or transfer data in other ways, the above L1 / PHY106 operations use the L2-L1 interface 104 between the L2 102 and the L1 106, as described above.
[0020] FIG. 2 is a block diagram showing a function call 202 to a physical layer (PHY) pipeline implemented by a PHY library 210 to execute a PHY operation including the operations described above in conjunction with FIG. 1. In at least one embodiment, as described above, one or more components implementing a fifth generation (5G) new radio (NR) network protocol stack, such as a layer 2 (L2) or layer 1 (L1) PHY driver, execute one or more function calls 202 to a PHY library interface 208. In at least one embodiment, the PHY library interface 208 is software instructions that, when executed, provide an application programming interface (API) to the PHY library 210. In at least one embodiment, the PHY library 210 is software instructions that, when executed, perform one or more L1 operations to execute one or more PHY functions of the 5G-NR protocol stack described above in conjunction with FIG. 1 and further described herein. In at least one embodiment, the PHY library 210 is a software-implemented 5G-NR library such as cuPHY, cuBB, or any other software 5G-NR library.
[0021] In at least one embodiment, one or more function calls 202 are software instructions that, when executed, call or otherwise invoke one or more functions provided by the API of the PHY library interface 208. In at least one embodiment, as further described below in conjunction with FIGS. 3A and 3B, one or more function calls include one or more descriptors 204 as input to the one or more function calls. In at least one embodiment, descriptor 204 is a data structure, such as a software container, for parameters 206 for one or more components of the PHY pipeline implemented by PHY library 210. In at least one embodiment, descriptor 204 occurs in the context of one or more kernel interfaces. In at least one embodiment, descriptor 204 occurs in the context of other interfaces of the 5G-NR platform. In at least one embodiment, as further described below, parameter 206 is a data value that indicates or includes information provided to one or more operations performed by one or more components of the PHY pipeline implemented by PHY library 210. In at least one embodiment, parameter 206 includes attributes of one or more PHY operations. In at least one embodiment, an attribute is a data value that indicates one or more properties of one or more PHY operations. For example, in one embodiment, an attribute indicates one or more cells that transmit information from one or more user equipment (UE) devices that are at least processed by one or more PHY operations implemented by PHY library 210. In other embodiments, an attribute indicates an identifier that is unique to one UE device or shared among multiple UE devices.
[0022] In at least one embodiment, as further described below, one or more function calls 202 call one or more functions provided by a PHY library interface 208 to a PHY library 210, and provide one or more descriptors 204 including one or more parameters 206 as input to the one or more functions provided by the PHY library interface 208. In at least one embodiment, the PHY library 210 performs one or more signal processing operations. In at least one embodiment, one or more descriptors 204 including one or more parameters 206 indicate one or more configurations of one or more signal processing operations performed by the PHY library 210 to execute a batch 214 of one or more signal processing operations called by one or more function calls 202.
[0023] In at least one embodiment, the batch 214 is a logical compilation of one or more signal processing operations or descriptors 204 and / or parameters 206 that make up one or more signal processing operations performed by the PHY library 210. In at least one embodiment, the PHY library 210 executes the batch 214 according to one or more characteristics of one or more function calls 202. For example, in one embodiment, the PHY library 210 executes a batch 214 of one or more function calls 202 corresponding to a single or multiple cell sites or other groupings of one or more user equipment (UEs) in a 5G-NR network. In other embodiments, the PHY library 210 executes the batch 214 according to other logical compilations of members and / or components of the 5G-NR network. To execute the batch 214 or otherwise support it, in one embodiment, as further described below, the PHY library 210 includes a structured data compilation 212. In at least one embodiment, the data compilation 212 is a logical compilation of data that facilitates the batch 214 by the PHY library 210, such as through a tree or other linked data relationship between data containers of the PHY library 210 described above.
[0024] FIG. 3A is a block diagram showing a physical layer (PHY) descriptor 302 according to at least one embodiment. In at least one embodiment, the PHY descriptor 302 is a data container that includes component parameters 304, 306, 308. In at least one embodiment, the component parameters 304, 306, 308 are one or more values or other data containers that describe the arrangement of control and / or data information for the PHY processing pipeline, as described above in conjunction with FIG. 1, and / or data that includes components that process information within the PHY processing pipeline.
[0025] In at least one embodiment, the PHY descriptor 302 includes data values that indicate common parameters available across one or more processing components in the PHY pipeline. In at least one embodiment, as described below, the PHY descriptor 302 includes data values that indicate kernel arguments 316 for one or more processing components in the PHY pipeline. In at least one embodiment, the PHY descriptor 302 includes data values that indicate launch geometry, such as which computing unit among one or more parallel processing units (PPUs), such as a graphics processing unit (GPU), executes different kernels to execute the PHY processing pipeline. In at least one embodiment, the PHY descriptor 302 includes data values that indicate kernel selection parameters that determine which kernels executed by a PPU, such as a GPU, execute the various computing components of the PHY processing pipeline. In at least one embodiment, the PHY descriptor 302 includes data values that indicate other information that configures control and / or data processing by a PPU, such as a central processing unit (CPU) and / or a GPU.
[0026] In at least one embodiment, component parameters 304, 306, 308 are data containers that include one or more data values that can be used to configure a PHY processing pipeline or processing components of a PHY processing pipeline. In at least one embodiment, component parameters 304, 306, 308 include data values that control data placement, operation timing, data size, and / or data refresh rate in memory. In at least one embodiment, component parameters 304, 306, 308 include data values that control the placement of data values of component descriptor 310 in memory. In at least one embodiment, component parameters 304, 306, 308 include data values that indicate the update time and / or update rate of parameters to a PHY processing pipeline and / or its processing components. For example, in one embodiment, component parameters 304, 306, 308 include data values that indicate that kernel arguments are updated early in a start time window, such as being updated at setup time or runtime by a driver as described above in conjunction with FIG. 1. In at least one embodiment, component parameters 304, 306, 308 include data values that indicate the movement of parameters in memory as a bulk data transfer, such as by moving all parameters corresponding to one kernel of a PPU such as a GPU.
[0027] In at least one embodiment, component parameters 304, 306, 308 include a data container such as component descriptor 310. In at least one embodiment, component descriptor 310 is a container of data values that includes data values that can be used to configure one or more processing components of a PHY pipeline. In at least one embodiment, component descriptor 310 facilitates the configuration of processing operations by a slot processing engine in a 5G-NR baseband device or other computing device to facilitate 5G-NR network operation. In at least one embodiment, the slot processing engine schedules computations to be executed during a slot. In at least one embodiment, a slot is a time window for execution by a central processing unit (CPU) or a processing element such as a GPU (PPU).
[0028] In at least one embodiment, the component descriptor 310 includes one or more flags 312. In at least one embodiment, the flag 312 is a data value indicating one or more binary or other data values corresponding to processing components of the PHY pipeline. For example, in one embodiment, the flag 312 includes a binary data value indicating whether the processing component of the PHY pipeline is enabled. In at least one embodiment, the component descriptor 310 includes a data value of the configuration 314. In at least one embodiment, the data value of the configuration 314 is a data value that can be used to configure the processing components of the PHY pipeline. For example, in one embodiment, the data value of the configuration 314 indicates one or more kernels executed by a PPU such as a GPU to execute the processing components of the PHY pipeline indicated by the component descriptor 310. In at least one embodiment, the component descriptor 310 includes kernel arguments 316. In at least one embodiment, the kernel argument is a data value indicating one or more data values indicating the configuration of a kernel such as a software kernel to perform and execute processing component operations in the PHY pipeline. In at least one embodiment, the component descriptor 310 includes other data values that can be used to configure the processing components of the PHY pipeline.
[0029] FIG. 3B is a block diagram showing an example PUSCH pipeline descriptor 318 in a physical layer (PHY) pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software fifth generation (5G) new radio (NR) library according to at least one embodiment, as described above in conjunction with FIG. 1. In at least one embodiment, the PUSCH pipeline descriptor 318 is a PHY descriptor, as described above in conjunction with FIG. 3A. That is, in one embodiment, the PUSCH pipeline descriptor 318 is an example data container that includes data values and / or additional data containers that can be used to configure one or more operations for executing PUSCH in the PHY pipeline, as described above in conjunction with FIG. 1.
[0030] In at least one embodiment, the example PUSCH pipeline descriptor 318 includes containers 322, 324, 326, 328, 330, 332, 334, 336, and each container 322, 324, 326, 328, 330, 332, 334, 336 includes parameters specific to low-level PHY processing operations that can be used to perform PUSCH in a fifth-generation (5G) new radio (NR) network. In at least one embodiment, the PUSCH pipeline descriptor 318 includes common parameters 320 that are data values indicating one or more configurations or other options shared among one or more of the containers 322, 324, 326, 328, 330, 332, 334, 336 within that PUSCH pipeline descriptor 318.
[0031] In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the PUSCH pipeline descriptor 318 include individual containers specific to each low-level computing operation that is executed as part of the PUSCH computing pipeline represented by the PUSCH pipeline descriptor 318. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the exemplary PUSCH pipeline descriptor 318 include channel estimation parameter 322. In at least one embodiment, the channel estimation parameter 322 is a data value that includes information that can be used to configure the channel estimation operation performed in the PUSCH pipeline, as further described herein. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the exemplary PUSCH pipeline descriptor 318 include equalizer parameter 324. In at least one embodiment, the equalizer parameter 324 is a data value that includes information that can be used to configure one or more equalization operations performed in the PUSCH pipeline. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the exemplary PUSCH pipeline descriptor 318 include soft demap parameter 326. In at least one embodiment, the soft demap parameter 326 is a data value that includes information that can be used to configure the soft demapping operation performed in the PUSCH pipeline.
[0032] In one example, the PUSCH pipeline descriptor 318 includes containers 322, 324, 326, 328, 330, 332, 334, 336, which include scrambling analysis parameter 328 in one embodiment. In at least one embodiment, the scrambling analysis parameter 328 is a data value including information that can be used to configure a scrambling analysis operation executed in the PUSCH pipeline. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the example PUSCH pipeline descriptor 318 include rate matching 330. In at least one embodiment, the rate matching parameter 330 is a data value including information that can be used to configure one or more rate matching operations executed in the PUSCH pipeline. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the example PUSCH pipeline descriptor 318 include low density parity check (LDPC) decoding parameter 332. In at least one embodiment, the LDPC decoding parameter 332 is a data value including information that can be used to configure one or more LDPC decoding operations executed as part of the PUSCH pipeline.
[0033] In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the exemplary PUSCH pipeline descriptor 318 include cyclic redundancy check (CRC) parameters 334 for the code block. The code block CRC parameter 334 is a data value that includes information that can be used to configure one or more code block CRC operations, as further described herein, that are performed as part of the PUSCH pipeline corresponding to the PUSCH pipeline descriptor 318. In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the exemplary PUSCH pipeline descriptor 318 include transport block CRC parameters 336. In at least one embodiment, the transport block CRC parameter 336 is a data value that includes information that can be used to configure one or more transport block CRC operations that are performed as part of the PUSCH pipeline.
[0034] In at least one embodiment, the containers 322, 324, 326, 328, 330, 332, 334, 336 of the PUSCH pipeline descriptor 318 of one or more examples at least include a PUSCH component descriptor 338. In at least one embodiment, the PUSCH component descriptor 338 is a data container that includes data values indicating one or more configuration options for the computational components of the PUSCH pipeline. For example, in one embodiment, the PUSCH component descriptor 338 includes an enable flag 340 that is a data value indicating whether a given PUSCH component corresponding to the PUSCH component descriptor 338 is enabled or executed in the PUSCH pipeline. In at least one embodiment, the PUSCH component descriptor 338 includes a kernel count 342. In at least one embodiment, the kernel count 342 is a data value and / or data structure that selects kernels and supplies arguments to those selected kernels. For example, in one embodiment, the kernel count 342 is a bitmap that is a data item that at least includes a one-dimensional or two-dimensional array of binary values, and each binary value indicates whether a particular software kernel executes PUSCH component operations on a parallel processing unit (PPU) such as a graphics processing unit (GPU). In at least one embodiment, the PUSCH component descriptor 338 includes one or more kernel arguments 344, 346, 348 for each kernel selected by the kernel selection bitmap 342. In at least one embodiment, the kernel arguments 344, 346, 348 are data values indicating one or more arguments or parameters provided to each selected kernel to execute PUSH component operations on a PPU such as a GPU.
[0035] FIG. 4A is a block diagram showing a hierarchical data organization for a physical layer (PHY) pipeline implemented by a PHY library according to at least one embodiment. In at least one embodiment, the search and access of one or more data values in a tree are fast computational operations, and since the tree structure has a low storage overhead, the hierarchical data organization improves the efficiency of data access and storage. In at least one embodiment, at the root of the tree structure, the cell parameter 402 includes data values indicating information such as cell-specific configuration information of a fifth generation (5G) new radio (NR) network in order to organize data as shown in FIG. 4A. By including all cell-specific information in the tree structure having the cell-specific parameter 402 at the root of the above tree structure, in one embodiment, information sharing between cells is eliminated, data dependency is reduced, and cells can be added to the 5G-NR network without modifying the existing cell configuration. In at least one embodiment, the cell parameter 402 includes data values including other cell-specific information together with versioning information, device-specific information, and the number of cells represented. In at least one embodiment, the number of cells represented is an attribute of a higher level of abstraction in a 5G-NR implementation. In at least one embodiment, the cell information indicated by the cell parameter 402 is visible to all other elements of the tree structure corresponding to the cell represented by the tree structure.
[0036] In at least one embodiment, the children of the nodes of the parent cell parameter 402 in the tree structure are the pipeline-specific parameters 404, 406, 408, 410. In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include pipeline-level information such as information that can be shared across different pipelines 404, 406, 408, 410, and the pipeline-level information is included within each pipeline rather than being propagated back to the parent cell parameter 402. In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include information that is visible from the nodes of each pipeline-specific parameter 404, 406, 408, 410 of the tree structure to all children and descendant nodes (descending node) of that tree structure.
[0037] In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include PHY channel parameters such as the PUCCH reception parameter 404. In at least one embodiment, the PUCCH reception parameter 404 is a container, as described above in conjunction with FIGS. 3A and 3B, and includes parameters and / or other information specific to the PUCCH reception operation in the PHY pipeline, as described above in conjunction with FIG. 1. In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include the PUSCH reception parameter 406. In at least one embodiment, the PUSCH reception parameter 406 is a container, as described above in conjunction with FIGS. 3A and 3B, and includes parameters and / or other information specific to the PUSCH reception operation in the PHY pipeline, as described above in conjunction with FIG. 1. In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include the PDSCH transmission parameter 408. In at least one embodiment, the PDSCH transmission parameter 408 is a container, as described above in conjunction with FIGS. 3A and 3B, and includes parameters and / or other information specific to the PDSCH transmission operation in the PHY pipeline, as described above in conjunction with FIG. 1. In at least one embodiment, the pipeline-specific parameters 404, 406, 408, 410 include the PDCCH transmission parameter 410. In at least one embodiment, the PDCCH transmission parameter 410 is a container, as described above in conjunction with FIGS. 3A and 3B, and includes parameters and / or other information specific to the PDCCH transmission operation in the PHY pipeline, as described above in conjunction with FIG. 1.
[0038] In at least one embodiment, the containers of each pipeline-specific parameter 404, 406, 408, 410 include pipeline-specific operation parameters 412, 414, 416, 418. In at least one embodiment, the pipeline-specific operation parameters 412, 414, 416, 418 are component descriptors as described above in conjunction with FIGS. 3A and 3B. In at least one embodiment, the pipeline-specific operation parameters 412, 414, 416, 418 include common parameters as described above in conjunction with FIG. 3B. In at least one embodiment, the pipeline-specific operation parameters 412, 414, 416, 418 include other component parameters not explicitly shown in FIG. 4, such as cyclic redundancy check (CRC) parameters, together with channel estimation parameter 414, rate matching parameter 416, and low density parity check (LDPC) parameter 418. In at least one embodiment, the pipeline-specific operation parameters 412, 414, 416, 418 include any other parameters corresponding to one or more PHY pipeline operations executed as part of a PHY pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein.
[0039] FIG. 4B is a block diagram showing the temporal data compilation for a PHY pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software fifth generation (5G) new radio (NR) library further described herein, according to at least one embodiment. In at least one embodiment, the parameters on a parallel processing unit (PPU) such as a central processing unit (CPU) and / or a graphics processing unit (GPU) are compiled by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein, according to elements to be considered temporally such as access rate and variability. In at least one embodiment, the temporally compiled parameters include static parameter 422, semi-static parameter 424, and / or dynamic parameter 426.
[0040] In at least one embodiment, the static parameter 422 is a parameter that is invariant during execution, such as the parameters described above in conjunction with FIGS. 3A and 3B. In at least one embodiment, the static parameter 422 is initialized by a PHY library, such as cuPHY, cuBB, or any other software 5G-NR library further described herein, at pipeline construction time and / or configuration time, stored in persistent memory, or backed up. In at least one embodiment, the semi-static parameter 424 is a parameter that changes over a relatively small number of slots or computational windows during 5G-NR pipeline execution, such as the parameters described above in conjunction with FIGS. 3A and 3B. In at least one embodiment, the semi-static parameter 424 is initialized when a specific event occurs, such as a configuration message from a higher layer (e.g., the second layer, as described above in conjunction with FIG. 1). In at least one embodiment, the dynamic parameter 426 is a parameter that is updated at each PHY pipeline execution slot, at a slot rate, and / or during execution slot setup, such as the parameters described above in conjunction with FIGS. 3A and 3B. In at least one embodiment, the dynamic parameter 426 is a parameter with a value that is updated frequently or a parameter whose value changes frequently.
[0041] In at least one embodiment, the temporally orchestrated parameters have increasing flexibility 428. That is, in one embodiment, the static parameter 422 has low flexibility or low ability to change, while the semi-static parameter 424 has increasing flexibility and the dynamic parameter 426 has maximum flexibility and can be updated or changed. In at least one embodiment, the temporally orchestrated parameters have increasing performance 430 that is inversely correlated with flexibility 428. That is, in one embodiment, the dynamic parameter 426 has low performance due to frequent updates, while the semi-static parameter 424 has increasing performance due to less frequent updates and / or changes, and the static parameter 422 has maximum performance due to its immutability.
[0042] FIG. 5 is a block diagram showing an example PUSCH pipeline data structure for a physical layer (PHY) pipeline implemented by a software PHY library such as cuPHY, cuBB, or any other software 5th generation (5G) new radio (NR) library further described herein. In at least one embodiment, example PUSCH reception parameter 502 is a PHY descriptor or a PHY component descriptor, as described above in conjunction with FIG. 3A. In at least one embodiment, PUSCH reception parameter 502 includes a pointer to a parent 504, such as a pointer to the parent of a tree structure as described above in conjunction with FIG. 4A, for hierarchical data organization.
[0043] In at least one embodiment, PUSCH reception parameter 502 includes common parameter 506, where the common parameter is a data value indicating one or more configuration options or other information shared among one or more component descriptors 510, 512, 514 corresponding to the descriptor of the above PUSCH reception parameter 502. In at least one embodiment, PUSCH reception parameter 502 includes a pointer to a child 508 in hierarchical organization, as described above in conjunction with FIG. 4.
[0044] In at least one embodiment, the pointer to the child 508 points to child component descriptors 510, 512, 514. In at least one embodiment, child component descriptors 510, 512, 514 include channel estimation parameters 510, rate matching parameters 512, low density parity check (LDPC) parameters, and / or any other parameters corresponding to components that perform one or more computational operations to execute the PUSCH reception pipeline, as described above in conjunction with FIGS. 1 and 3B.
[0045] In at least one embodiment, as described above in conjunction with FIG. 4B, the common parameters 506, 516 in the PUSCH reception parameter 502 descriptor are compiled into static parameters 518, semi-static parameters 520, and / or dynamic parameters 522 by a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein. In at least one embodiment, the static parameters 518, 524 include i parameters 526, 528, where, as described above in conjunction with FIG. 4B, the i parameters 526, 528 are invariant. In at least one embodiment, the semi-static parameters 520, 530 include j parameters 532, 534, where, as described above in conjunction with FIG. 4B, the j parameters 532, 534 vary according to a slot frequency or other execution scheduling metric for one or more slots in order to execute a PHY pipeline such as the PUSCH reception pipeline. In at least one embodiment, the dynamic parameters 522, 536 include k parameters 538, 540, where, as described above in conjunction with FIG. 4B, the k parameters 538, 540 vary and / or are updated frequently.
[0046] In at least one embodiment, a parallel processing unit (PPU) 542, such as a graphics processing unit (GPU), stores the static parameters 524, 544 in a memory that can be used to store invariant data values. In at least one embodiment, a PPU 542, such as a GPU, stores the semi-static parameters 530, 546 in a memory that can be used to store periodically updated data values. In at least one embodiment, a PPU 542, such as a GPU, stores the dynamic parameters 536, 548 in a memory that can be used for data values having frequent changes and / or updates.
[0047] FIG. 6 is a block diagram showing physical layer (PHY) descriptor buffering according to at least one embodiment. In at least one embodiment, prior to the slot execution time for the PHY pipeline corresponding to the static PHY descriptor 604, the static PHY descriptor 604 is assembled in the memory of the central processing unit (CPU) 602 as described above in conjunction with FIGS. 3A and 4B by a software PHY library such as cuPHY, cuBB, or any other software fifth generation (5G) new radio (NR) library further described herein, and is copied to the memory of a parallel processing unit (PPU) 610 such as the memory of a graphics processing unit (GPU). In at least one embodiment, the static PHY descriptor 603 is assembled by a software PHY library during the setup of the PHY pipeline corresponding to that static PHY descriptor. In at least one embodiment, a PPU 610 such as a GPU stores the copied static PHY descriptor 612 as described above in conjunction with FIG. 5.
[0048] In at least one embodiment, a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein buffers the quasi-static PHY descriptor 606 and the dynamic PHY descriptor 608 on the CPU 602. In at least one embodiment, buffering the quasi-static PHY descriptor 606 and the dynamic PHY descriptor 608 facilitates slot processing of the PHY pipeline corresponding to that quasi-static PHY descriptor 606 and that dynamic PHY descriptor 608. In at least one embodiment, a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein buffers the quasi-static PHY descriptor 606 and the dynamic PHY descriptor 608 on the CPU 602 and copies that quasi-static PHY descriptor 606 and that dynamic PHY descriptor 608 to one or more PPUs 610. In at least one embodiment, one or more PPUs 610 such as a GPU store the copied quasi-static PHY descriptor 614 and the copied dynamic PHY descriptor 616 as described above in conjunction with FIG. 5.
[0049] In at least one embodiment, as illustrated in FIG. 6, the buffering of the time-classified PHY descriptors is used only when necessary. For example, in one embodiment, the pipeline level may require the static, semi-static, and dynamic parameters included in the static PHY descriptors 604, 612, the buffered semi-static PHY descriptors 606, 614, and the buffered dynamic PHY descriptors 608, 616, but the component may require only the static PHY descriptors 604, 612, and / or the buffered dynamic PHY parameters 608, 616. In at least one embodiment, the depth of many PHY channel processing pipelines and the corresponding PHY descriptor buffers is adjusted by a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein to compensate for processing latency. For example, up to N semi-static PHY descriptors 606, 614 and up to M dynamic PHY descriptors 608, 616 may be buffered by a software PHY library in one embodiment to compensate for the processing latency for execution slots during 5G-NR processing.
[0050] FIG. 7 is a block diagram showing batched parameter compilation during a physical layer (PHY) operation batch according to at least one embodiment. In at least one embodiment, a batch is a logical compilation of computational PHY operations in a PHY pipeline, or a combination thereof, such that PHY operations are computed by one or more kernels on a parallel processing unit (PPU) such as a graphics processing unit (GPU). In at least one embodiment, a software PHY library such as cuPHY, cuBB, or any other software fifth generation (5G) new radio (NR) library as further described herein batches PHY operations according to different workload configurations. In at least one embodiment, an example workload configuration is a number of cell sites with a small number of connected user equipment (UE) such as mobile phones, processed by a 5G-NR baseband unit (BBU). In at least one embodiment, another example workload configuration is a small number of cell sites with a large number of connected UE, processed by a 5G-NR BBU.
[0051] In at least one embodiment, a software PHY library, such as cuPHY, cuBB, or any other software 5G-NR library further described herein, batches parameters corresponding to PHY pipeline operations based on a workload. For example, in one embodiment, the software PHY library batches parameters corresponding to PHY pipeline operations in accordance with the arrival of the workload. In at least one embodiment, the batch according to the arrival of the workload arranges or groups the parameters according to spatial characteristics, where the software PHY library batches the parameters across the simultaneous workloads available in a given slot, processing time slot, such as parameters that configure operations on information received from devices within a cell or across multiple cells. In other embodiments, the batch according to the arrival of the workload groups the parameters according to temporal characteristics, where the software PHY library batches the parameter-workload over time intervals, such as by processing multiple symbols within an execution slot, processing multiple cells sequentially for small per-cell workloads, or processing across multiple PHY channels, to sequentially perform operations such as PUSCH and PDSCH.
[0052] Other examples of batches of PHY parameters corresponding to PHY pipeline operation based on workloads are, in one embodiment, software PHY libraries such as cuPHY, cuBB, or any other software 5G-NR library that batches according to the workload configuration. In at least one embodiment, the batch according to the workload configuration arranges or groups the parameters according to the homogeneous characteristics of the PHY operation parameters to be batched. In at least one embodiment, with a homogeneous batch, a single kernel can process multiple identically configured workloads. In at least one embodiment, the batch according to the homogeneous characteristics includes batching within a specific kernel specialization dimension by aggregating parameters that arrived simultaneously at the software PHY library or parameters that arrived over time. In at least one embodiment, the batch according to the workload configuration arranges or groups the parameters according to the heterogeneous characteristics of the PHY operation parameters to be batched. In at least one embodiment, with a heterogeneous batch, several heterogeneous workloads can be set up and processed by a single component. In at least one embodiment, the batch according to the heterogeneous characteristics includes batching across the kernel specialization dimension to combine the workloads into a single computational graph.
[0053] In at least one embodiment, the kernel specialization dimension is a kernel executed by a PPU such as a GPU, and the kernel is customized by the software PHY library for each workload configuration to fit the problem size, leading to better execution time and / or throughput at the expense of an increase in startup overhead. In at least one embodiment, the kernel generalization dimension is a kernel executed by a PPU such as a GPU, and the kernel is customized by the software PHY library to support multiple workloads, which can thereby reduce the efficiency of the kernel.
[0054] FIG. 7 is a block diagram showing an example of a batch of input parameters by a software PHY library. The PUSCH batch configuration parameter 702 is, in one embodiment, a data container including parameters 704, 706, and 708 that are batched for each PHY pipeline operation, such as channel estimation batch parameter 704, channel equalization batch parameter 706, and low density parity check (LDPC) batch parameter 708, to execute PHY PUSCH. In at least one embodiment, the batch configuration parameter 702 specifies how the batch is performed, such as how to group the parameters of the UE group super set. In at least one embodiment, the batch configuration parameter 702 is a part of each component. In at least one embodiment, the batch configuration parameter 702 is a part of the PHY pipeline. In at least one embodiment, the software PHY library executes heterogeneous batches according to the workload type. The first LDPC batch parameter 710 indicates, in one embodiment, a number of different workload types batched by the software PHY library. For example, in FIG. 7, the first LDPC batch parameter 710 indicates three workload types that require three software kernels 712, 726, and 736 to execute the LDPC operation. In at least one embodiment, the software PHY library batches or groups the parameters to the number of kernels 712, 726, and 736 indicated by the first LDPC batch parameter 710. For the first type of LDPC batch parameters 714, 716, 718, 720, 722, and 724, in one embodiment, the software PHY library batches its batch parameters to the first LDPC kernel 712.
[0055] For the first type of LDPC batch parameters 714, 716, 718, 720, 722, 724, in one embodiment, the software PHY library heterogeneously groups or batches the LDPC batch parameters 714, 716, 718, 720, 722, 724 processed or executed by the first LDPC kernel 712 according to the type of workload. In at least one embodiment, for the second type of LDPC batch parameters 728, 730, 732, 734, the software PHY library heterogeneously groups or batches the LDPC batch parameters 728, 730, 732, 734 processed or executed by the second LDPC kernel 726 according to the type of workload. In at least one embodiment, for the third type of LDPC batch parameters 738, 740, 742, 744, 746, the software PHY library heterogeneously groups or batches the LDPC batch parameters 738, 740, 742, 744, 746 processed or executed by the third LDPC kernel 736 according to the type of workload.
[0056] In at least one embodiment, batch parameters such as LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 are encoded in a type-length-value format for flexibility and efficient memory usage, as shown in FIG. 7. In at least one embodiment, batch parameters such as LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 are encoded by a 5G-NR PHY library for using an array of a fixed maximum length for each group or batch type. In at least one embodiment, for each group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746, the first LDPC batch parameters 714, 728, 738 indicate the type associated with each group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 processed by each LDPC kernel 712, 726, 736. In at least one embodiment, for each group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746, the second LDPC batch parameters 716, 730, 740 indicate the length or number of parameters in that group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746. In at least one embodiment, the remaining LDPC batch parameters 718, 720, 722, 724, 732, 734, 742, 744, 746 of the group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 include parameter data values such as indicators to a user equipment (UE) group super-set 748, as described below.In other embodiments, the remaining LDPC batch parameters 718, 720, 722, 724, 732, 734, 742, 744, 746 of the group or batch of LDPC batch parameters 714, 716, 718, 720, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 include any other parameter - data values that facilitate the configuration of one or more PHY pipeline operations.
[0057] In at least one embodiment, the UE group - super - set 748 is a data container that includes batched PUSCH kernel parameters 750, 752, 754 for each kernel 712, 726, 736. In at least one embodiment, the software PHY library batches the PUSCH kernel parameters 750, 752, 754 heterogeneously to the UE group - super - set 748 according to one or more characteristics of the PUSCH kernel parameters 750, 752, 754, such as arrival time or slot execution time requirements. In at least one embodiment, the software PHY library batches the parameters 714, 716, 718, 720, 722, 724, 728, 730, 732, 734, 738, 740, 742, 744, 746 for each computational operation, such as channel estimation 704, channel equalization 706, LDPC, and / or any other low - level PHY operations, heterogeneously to one or more kernels 712, 726, 736 according to the type of the parameters. In at least one embodiment, the software PHY library batches the parameters 750, 752, 754 heterogeneously to the UE group - super - set 748 according to other parameter characteristics, such as arrival time or slot execution time requirements.
[0058] FIG. 8 is a block diagram showing a pipeline topology of an example for executing a batched PHY operation workload according to at least one embodiment. In at least one embodiment, one or more parallel processing units (PPUs), such as a graphics processing unit (GPU), execute software kernels 802, 810, 834, where each software kernel performs one or more PHY calculation operations. In at least one embodiment, each software kernel 802, 810, 834 performs one or more PHY calculation operations using the batched parameters as described above in conjunction with FIG. 7. In at least one embodiment, each software kernel 802, 810, 834 performs one or more PHY calculation operations using the batched parameters, where the batched parameters are grouped according to a homogeneous workload configuration batch as described above in conjunction with FIG. 7. In at least one embodiment, each software kernel 802, 810, 834 performs one or more PHY calculation operations using the batched parameters, where the batched parameters are grouped according to a heterogeneous workload configuration batch as described above in conjunction with FIG. 7. In at least one embodiment, each software kernel 802, 810, 834 performs one or more PHY calculation operations using the batched parameters, where the batched parameters are grouped according to a spatial grouping based on workload arrival as described above in conjunction with FIG. 7. In at least one embodiment, each software kernel 802, 810, 834 performs one or more PHY calculation operations using the batched parameters, where the batched parameters are grouped according to a temporal grouping based on workload arrival as described above in conjunction with FIG. 7.
[0059] In at least one embodiment, the software kernels 802, 810, 834 execute one or more PHY computing operations in parallel with other software kernels 802, 810, 834. In at least one embodiment, each software kernel 802, 814, 834 executes one or more PHY computing operations configured based on parameters grouped or batched according to type by heterogeneous batches, as described above. In at least one embodiment, each software kernel 802, 810, 834 executes one or more PHY computing operations for each of the pipeline stages 818, 820, 822, 826, 828. For each configuration specified by individually batched parameters, in one embodiment, the software kernels 802, 810, 834 execute one or more PHY computing operations. Between the PHY computing operations, one or more of the pipeline stages 818, 820, 822, 826, 828 store data calculated as a result of each PHY computing operation configured according to the batched parameters.
[0060] In at least one embodiment, one or more of the pipeline stages 818, 820, 822, 826, 828 are memories such as registers that store one or more values received by the one or more of the pipeline stages 818, 820, 822, 826, 828 as output data from one or more parallel PHY computing operations performed by one or more of the kernels 802, 810, 834. In at least one embodiment, one or more of the pipeline stages 818, 820, 822, 826, 828 are shared across one or more of the kernels 802, 810, 834. In at least one embodiment, each of one or more of the kernels 802, 810, 834 includes an individual pipeline stage that stores intermediate data results of one or more PHY computing operations performed by the corresponding kernel among the one or more of the kernels 802, 810, 834.
[0061] Between each of the pipeline stages 818, 820, 822, 826, 828, each of one or more kernels 802, 810, 834 performs one or more PHY computing operations, where each computing operation performed by each kernel is configured by batched or grouped parameters specific to the workload, as described above in conjunction with FIG. 7. In at least one embodiment, the one or more PHY computing operations include channel estimations 804, 812, 836, as further described herein. In at least one embodiment, the operation of each channel estimation 804, 812, 836 is configured by a batch of parameters corresponding to the type of workload, as described above in conjunction with FIG. 7. In at least one embodiment, the one or more PHY computing operations include channel estimations 806, 814, 838, as further described herein. In at least one embodiment, the operation of each channel estimation 806, 814, 838 is configured by a batch of parameters corresponding to the type of workload, as described above in conjunction with FIG. 7.
[0062] In at least one embodiment, the one or more PHY computing operations are shared among sets of configuration parameters grouped by batch. In at least one embodiment, the shared PHY computing operations include rate matching and scramble resolution 824, code block cyclic redundancy check (CB CRC) and aggregation 830, transport block (TB) CRC 832, and / or any other PHY computing operations shareable among batches of configuration parameters. In at least one embodiment, the one or more PHY computing operations not shared among one or more kernels 802, 810, 834 include low density parity check (LDPC) decoding and / or encoding 808, 816, 840, as further described herein. In at least one embodiment, the operation of each LDPC decoding and / or encoding 808, 816, 840 is configured by a batch of parameters corresponding to the type of workload, as described above in conjunction with FIG. 7.
[0063] FIG. 9 is a block diagram showing an example of a physical layer (PHY) batching topology based on workload time slots according to at least one embodiment. In at least one embodiment, one or more kernels 906, 926, 946 perform one or more PHY computing operations, such as segmentation and code block cyclic redundancy check (Seg+CB CRC) 908, 928, 948, low density parity check (LDPC) encoding / decoding 910, 930, 950, rate matching 912, 932, 952, scrambling 914, 934, 954, modulation 916, 936, 956, layer mapping 918, 938, 958, precoding 920, 940, 960, mapping 922, 942, 962, and / or any other one or more PHY computing operations such as any other fifth generation (5G) new radio (NR) PHY pipeline operations, as further described herein. In at least one embodiment, kernels 906, 926, 946 perform one or more PHY computing operations configured using parameters based on grouping according to time characteristics such as slot execution time, as further described above in conjunction with FIG. 7.
[0064] In at least one embodiment, one or more kernels 906, 926, 946 perform one or more PHY computing operations configured by parameters batched according to execution time slots. In at least one embodiment, slot execution start point 902 is the point in time at which one or more kernels 906, 926, 946 are scheduled for execution by a scheduler provided by a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library described herein. In at least one embodiment, time slots 904, 924, 944 are windows of slot execution time that occur after slot execution start point 902. From slot execution start point 902, in one embodiment, one or more software kernels 906 perform one or more PHY pipeline computing operations configured using parameters batched according to time slot t 0 904. In at least one embodiment, one or more software kernels 926 perform in a later time slot t1 Execute one or more PHY pipeline calculation operations configured using other parameters batched according to 924. Delay t from slot execution start point 902 n After 944, in one embodiment, one or more kernels 946 have the above time delay t to the execution slot n Execute one or more PHY pipeline calculation operations configured using parameters batched by the software PHY library according to. In at least one embodiment, each time slot t 1 924 ··· t n For 944, t n ≧t 1 +t proc where X is the time to complete a process such as a PHY pipeline and / or one or more PHY pipeline operations.
[0065] FIG. 10 is a block diagram showing the arrangement of batched physical layer (PHY) descriptors according to at least one embodiment. In at least one embodiment, for a PUSCH pipeline as further described herein, a PUSCH pipeline batch descriptor 1002 is a container including one or more pipeline descriptors 1006, 1008, 1010 for a number of pipeline instances 1004 that execute batched PUSCH pipeline operations by a software PHY library such as cuPHY, cuBB, or any other software 5th generation (5G) new radio (NR) library described herein. In at least one embodiment, a PUSCH pipeline batch descriptor 1002 or any other PHY pipeline batch descriptor includes one or more pointers to one or more pipeline descriptors 1006, 1008, 1010, along with data indicating the number of pipeline instances 1004. In at least one embodiment, each pointer to a pipeline descriptor 1006, 1008, 1010 is data including a memory address indicating a storage location for a pipeline PHY descriptor such as a PUSCH pipeline PHY descriptor 1012.
[0066] In at least one embodiment, a pipeline PHY descriptor, such as the PUSCH pipeline PHY descriptor 1012, is a data container. In at least one embodiment, a pipeline PHY descriptor, such as the PUSCH pipeline PHY descriptor 1012, is a data container that includes a common parameter 1014 and one or more pointers to component descriptors 1016, 1018, 1020, as described above in conjunction with FIGS. 3 and 5. In at least one embodiment, one or more pointers to component descriptors 1016, 1018, 1020 are data that includes a memory address indicating a storage location for one or more component descriptors, such as the PUSCH component batch descriptor 1022. In at least one embodiment, a component is one or more PHY computing operations that are executed by one or more kernels using one or more parallel processing units (PPUs), such as a graphics processing unit (GPU), as described above.
[0067] In at least one embodiment, a component descriptor, such as a PUSCH component batch descriptor 1022, is a data container. In at least one embodiment, a component descriptor, such as a PUSCH component batch descriptor 1022, includes data indicating the number of component instances to be executed by individual kernels each executing a different configuration indicated by component parameters 1026, 1028, 1030, 1032. In at least one embodiment, a component descriptor, such as a PUSCH component batch descriptor 1022, includes parameters grouped according to heterogeneous batches within a PHY component by a software PHY library, such as cuPHY, cuBB, or any other software 5G-NR library described herein, as described above in conjunction with FIG. 7. In at least one embodiment, a component descriptor, such as a PUSCH component batch descriptor 1022, includes a pointer to a batched component descriptor and / or parameters for a heterogeneous configuration, where N3 kernels each execute a different configuration batched as component parameters 1026, 1028, 1030, 1032. In at least one embodiment, a component descriptor, such as a PUSCH component batch descriptor 1022, includes the number of component instances 1024. In at least one embodiment, the number of component instances 1024 is a data value indicating the number N3 of groups or batches of its component parameters 1026, 1028, 1030, 1032 to be executed by N3 kernels each executing a different component configuration indicated by component parameters 1026, 1028, 1030, 1032.
[0068] In at least one embodiment, one or more component parameters 1026, 1028, 1030, 1032 of a component descriptor such as a PUSCH component batch descriptor 1022 are compiled or batched by a software PHY library according to temporal characteristics or homogeneous characteristics such as parameter update frequency, as described above in conjunction with FIGS. 4B and 7. In at least one embodiment, the software PHY library compiles the component parameters 1026, 1028, 1030, 1032 into component static parameters such as PUSCH component static parameters 1034. In at least one embodiment, the software PHY library compiles other component parameters 1026, 1028, 1030, 1032 into component semi-static parameters such as PUSCH component semi-static parameters 1034. In at least one embodiment, the software PHY library compiles the component parameters 1026, 1028, 1030, 1032 into component dynamic parameters such as PUSCH component dynamic parameters 1038. In at least one embodiment, component static parameters such as PUSCH component static parameters 1034, component semi-static parameters such as PUSCH component semi-static parameters 1036, and component dynamic parameters such as PUSCH component dynamic parameters 1038 include pointers to batched component descriptors and / or parameters for homogeneous configurations, where each kernel batch processes a plurality of workloads having the same configuration.
[0069] FIG. 11 is a block diagram showing an example application programming interface (API) 1110 to a physical layer (PHY) pipeline implemented by a software PHY library to execute a pipeline configuration and / or batch as described above, according to at least one embodiment. In at least one embodiment, a software PHY library such as cuPHY, cuBB, or any other software 5th generation (5G) new radio (NR) library implements the API 1110 to configure and execute PHY pipeline operations using a configuration defined by parameters included in a descriptor, as described above in conjunction with FIGS. 2 and 3. In at least one embodiment, the PHY pipeline API 1110 is a software instruction that provides an interface callable to execute one or more PHY pipeline operations when executed. In at least one embodiment, the software PHY library providing the PHY pipeline API batches the parameters received in the descriptor as a result of one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110.
[0070] In at least one embodiment, the PHY pipeline API 1110 receives one or more descriptors including one or more PHY operations to execute one or more PHY operations and / or one or more parameters to configure one or more components, as described above in conjunction with FIGS. 2 and 3, as a result of one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110. In at least one embodiment, one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110 are software instructions that call one or more functions provided by the PHY pipeline API 1110 when executed.
[0071] In at least one embodiment, one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110 call an initialization (init) or de-initialization (deinit) function 1102 provided by that PHY pipeline API 1110. In at least one embodiment, the init 1102 function, when executed, is a logical compilation of software instructions that perform pipeline construction and / or configuration-time operations for a PHY pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library described herein. For example, the init 1102 function, when executed, performs object instantiation and / or memory allocation for a PHY library such as a compute uniform device architecture (CUDA) or other parallel computing library and / or any other software library further described herein. In at least one embodiment, the deinit function 1102, when executed, is a logical compilation of software instructions that tear down or otherwise stop pipeline execution and / or free resources such as memory used by the pipeline.
[0072] In at least one embodiment, as described above in conjunction with FIG. 4B, the initialization function, or the create 1102 function, updates static parameters when executed. In at least one embodiment, the create 1102 function asynchronously updates one or more static parameters in relation to slot execution. In at least one embodiment, the create 1102 function is executed by one or more parallel processing units (PPUs), such as a central processing unit (CPU) and / or a graphics processing unit (GPU), to initialize resources available for use by the software PHY library. In at least one embodiment, the create 1102 function is called less frequently compared to other functions of the PHY pipeline API 1110. In at least one embodiment, the create 1102 function has a time budget of about a few seconds. In at least one embodiment, the create 1102 function is executed as a result of one or more calls to the PHY pipeline API 1110 implemented by a PHY library, such as cuPHY, cuBB, or any other software 5G-NR library, to process cell information, such as sector-carrier information.
[0073] In at least one embodiment, one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110 call a configuration (config) or reconfiguration (reconfig) function 1104 provided by that PHY pipeline API 1110. In at least one embodiment, the config 1104 function is a logical compilation of software instructions that, when executed, perform a pipeline configuration update. In at least one embodiment, the config 1104 function is a logical compilation of software instructions that, when executed, perform a pipeline configuration update using parameters as described above in conjunction with FIGS. 2 and 3A, and where the update frequency is less than the slot rate. For example, the config 1104 function, when executed, updates the configuration using new parameters received as a result of a call to the above-described config 1104 function of one or more PHY pipeline operations executed by a PHY library such as cuPHY, cuBB, or a compute uniform device architecture (CUDA) or any other parallel computing or any other software library such as the 5G-NR library further described herein. In at least one embodiment, the reconfig function 1104 is a logical compilation of software instructions that, when executed, adjust the configuration of one or more PHY pipeline calculation operations during execution or between execution slots, as described above.
[0074] In at least one embodiment, as described above in conjunction with FIG. 4B, the config and / or reconfig1104 functions update static parameters when executed. In at least one embodiment, as described above in conjunction with FIG. 4B, the config and / or reconfig1104 functions update quasi-static parameters when executed. In at least one embodiment, the config and / or reconfig1104 functions update one or more static parameters asynchronously. In at least one embodiment, the config and / or reconfig1104 functions update one or more static parameters asynchronously before a slot boundary. In at least one embodiment, the config and / or reconfig1104 functions are implemented by a software PHY library and executed by one or more PPU such as a CPU and / or GPU to configure one or more PHY operations executed by a CPU and / or one or more PPU. In at least one embodiment, the config and / or reconfig1104 functions are called less frequently compared to other functions of the PHY pipeline API1110. In at least one embodiment, the config and / or reconfig1104 functions are called at a similar frequency compared to other functions of the PHY pipeline API1110. In at least one embodiment, the config and / or reconfig1104 functions have a time budget of tens to hundreds of milliseconds. In at least one embodiment, the config and / or reconfig1104 functions have a time budget of hundreds of microseconds. In at least one embodiment, the config and / or reconfig1104 functions are executed as a result of one or more calls to the PHY pipeline API1110 implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library to process signaling information such as region updates.In at least one embodiment, the config and / or reconfig1104 functions are executed as a result of one or more calls to the PHY pipeline API1110 implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library to process user equipment (UE) information such as whether the UE is connected or inactive.
[0075] In at least one embodiment, one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API1110 call the setup1106 function provided by that PHY pipeline API1110. In at least one embodiment, the setup1106 function, when executed, is a logical compilation of software instructions that perform PHY descriptor setup using the slot structure information required to execute one or more PHY pipelines implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library described herein. For example, the setup1106 function, when executed, configures and executes batches using descriptors containing parameters, as described above, by a PHY library and / or other software library such as compute uniform device architecture (CUDA) or any other parallel computing library further described herein.
[0076] In at least one embodiment, as described above in conjunction with FIG. 4B, the setup 1106 function updates dynamic parameters when executed. In at least one embodiment, the setup 1106 function synchronously updates one or more dynamic parameters before the slot execution boundary. In at least one embodiment, the setup 1106 function is executed by one or more processing units (PPUs), such as a CPU and / or GPU, to configure and / or batch one or more PHY pipeline operations implemented by a software PHY library. In at least one embodiment, the setup 1106 function is called more frequently than other functions of the PHY pipeline API 1110. In at least one embodiment, the setup 1106 function has a time budget of 125 microseconds or less. In at least one embodiment, the setup 1106 function is executed as a result of one or more calls to the PHY pipeline API 1110 implemented by a PHY library, such as cuPHY, cuBB, or any other software 5G-NR library, to process slot allocation information such as downlink allocation and uplink grant.
[0077] In at least one embodiment, one or more function calls 1102, 1104, 1106, 1108 to the PHY pipeline API 1110 call the execution 1108 function provided by that PHY pipeline API 1110. In at least one embodiment, the execution 1108 function, when executed, is a logical compilation of software instructions that perform a pipeline startup for one or more PHY pipelines implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library described herein. For example, the execution 1108 function, when executed, triggers the start of one or more pipelines implemented by a PHY library and / or any other software library such as a compute uniform device architecture (CUDA) or any other parallel computing library further described herein, which are executed by one or more PPUs such as a CPU and / or GPU.
[0078] In at least one embodiment, the execution 1108 function, when executed, does not update any of the parameters described above in conjunction with FIG. 4B. In at least one embodiment, the execution 1108 function is executed synchronously during slot execution and / or symbol reception. In at least one embodiment, the execution 1108 function is executed by one or more PPUs such as a CPU and / or GPU to start the execution of one or more PHY pipelines implemented by a software PHY library. In at least one embodiment, the execution 1108 function is called more frequently than other functions of the PHY pipeline API 1110. In at least one embodiment, the execution 1108 function has an immediate time budget because it is a trigger to start slot execution. In at least one embodiment, the execution 1108 function operates as a slot processing trigger that causes the startup of one or more PPU kernels and / or the startup of one or more compute graphs, and is thus executed as a result of one or more calls to the PHY pipeline API 1110 implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library.
[0079] Figure 12 is a diagram showing a process 1200 for performing PHY operations in a fifth generation (5G) New Radio (NR) physical layer (PHY) pipeline implemented by a PHY library such as cuPHY, cuBB, or any other software 5G-NR library further described herein. In at least one embodiment, process 1200 begins 1202 by constructing 1204 one or more PHY pipelines to perform PHY operations. During pipeline construction 1204, in one embodiment, one or more data structures are allocated and initialized in a memory corresponding to one or more parallel processing units (PPUs) such as a central processing unit (CPU) and / or a graphics processing unit (GPU) as described above in conjunction with FIG. 11.
[0080] In at least one embodiment, when a software PHY library constructs 1204 one or more pipelines, the software PHY library configures 1206 the one or more pipelines according to configuration parameters received as a result of one or more function calls as described above in conjunction with FIGS. 2 and 3A. After configuration 1206, in one embodiment, a software PHY library such as cuPHY, cuBB, or any other software 5G-NR library performs a setup 1208 operation to set up PHY pipeline operations for slot execution according to configuration information provided by one or more descriptors as described above in conjunction with FIGS. 3A and 5. In at least one embodiment, setup 1208 includes a batch of one or more PHY operations based on parameters provided by one or more PHY descriptors as described above in conjunction with FIGS. 7 through 9.
[0081] In at least one embodiment, when a software PHY library sets up one or more PHY pipelines 1208 according to one or more parameters included in one or more descriptors received as a result of one or more function calls to a software PHY library interface, as described above in conjunction with FIGS. 1, 2, and 11, the software PHY library then launches the one or more PHY pipelines 1210. In one embodiment, the software PHY library launches one or more PHY pipelines 1210 that are executed in one or more slots by one or more PPUs such as a GPU. In other embodiments, the software PHY library launches one or more PHY pipelines 1210 that are executed in one or more slots by a CPU, as described above in conjunction with FIG. 11.
[0082] In at least one embodiment, when launching one or more PHY pipelines 1210, the software PHY library may need to reconfigure 1212 some or all of the one or more PHY pipelines during execution. In at least one embodiment, when one or more PHY pipelines or operations that execute the one or more PHY pipelines are reconfigured 1212 as a result of one or more function calls to a PHY library interface that includes updated parameters and / or descriptors, in one embodiment, the PHY library reconfigures 1206 the one or more PHY pipelines and / or operations that execute the one or more PHY pipelines.
[0083] In at least one embodiment, the PHY library determines whether the slot execution of one or more of the above PHY pipelines is completed 1212. In one embodiment, when the slot execution of one or more PHY pipelines is completed 1212, the process 1200 determines whether reconfiguration is required 1214. In at least one embodiment, when reconfiguration is required 1214, the process 1200 reconfigures the pipeline 1206. In at least one embodiment, when reconfiguration 1214 is not required, the process 1200 determines whether an additional pipeline is to be executed 1216. In at least one embodiment, when an additional pipeline is to be executed 1216, the process 1200 continues slot execution by setting up the PHY descriptor 1208. In at least one embodiment, when an additional pipeline is not to be executed 1216, or when the execution is not completed, the process 1200 ends 1218.
[0084] The techniques described and proposed herein enable, in one embodiment, 5th generation (5G) New Radio (NR) operations, such as physical layer (PHY) operations of a PHY pipeline, to be executed in parallel using computing resources such as one or more parallel processing units (PPUs), as described above in conjunction with FIG. 1 and further described herein. In other embodiments, the techniques described and proposed herein enable 5G-NR operations to be executed in parallel using other computing resources such as one or more software kernels. As described above, in one embodiment, one or more computing operations, such as 5G-NR PHY operations, are grouped according to computing resources such as one or more kernels and / or one or more PPUs. In at least one embodiment, one or more computing operations, such as 5G-NR PHY operations, are grouped according to attributes indicative of other computing resources such as a 5G-NR cell and / or a user equipment (UE) connected to the 5G-NR cell.
[0085] In at least one embodiment, as described above, software libraries such as 5G-NR PHY libraries group one or more computing operations so that their computing operations can be executed in parallel using computing resources such as software kernels and / or PPU. In at least one embodiment, the techniques described and proposed herein for enabling 5G-NR operations to be executed in parallel according to one or more computing resources are implemented using one or more circuits to cause the 5G-NR operations to be executed in parallel according to the techniques described above. In at least one embodiment, the techniques described and proposed herein are implemented in one or more systems comprising one or more processors including, but not limited to, a PPU such as a central processing unit and / or a graphics processing unit. In at least one embodiment, the techniques described and proposed herein for executing 5G-NR operations in parallel are implemented using a software library to perform one or more parallelization methods further described herein. In at least one embodiment, the techniques described and proposed herein for executing 5G-NR operations in parallel are implemented as one or more instructions on a machine-readable or computer-readable medium that group the 5G-NR operations according to attributes indicative of computing resources as described above.
[0086] Data center FIG. 13 shows an exemplary data center 1300 in which at least one embodiment may be used. In at least one embodiment, the data center 1300 includes a data center infrastructure layer 1310, a framework layer 1320, a software layer 1330, and an application layer 1340.
[0087] In at least one embodiment, as shown in FIG. 13, the data center infrastructure layer 1310 may include a resource orchestrator 1312, grouped computing resources 1314, and node computing resources (referred to as "node C.R.") 1316(1) to 1316(N), where "N" represents any positive integer. In at least one embodiment, the node C.R.s 1316(1) to 1316(N) may include any number of central processing units (referred to as "CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic random access memory), storage devices (e.g., semiconductor drives or disk drives), network input / output (referred to as "NW I / O") devices, network switches, virtual machines (referred to as "VMs"), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R.s 1316(1) to 1316(N) may be servers having one or more of the computing resources described above.
[0088] In at least one embodiment, the grouped computing resources 1314 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). In at least one embodiment, the separate groups of node C.R.s within the grouped computing resources 1314 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s including a CPU or processor may be grouped within one or more racks to provide compute resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0089] In at least one embodiment, the resource orchestrator 1312 may configure or otherwise control one or more node C.R.s 1316(1) - 1316(N) and / or the grouped computing resources 1314. In at least one embodiment, the resource orchestrator 1312 may include a software design infrastructure ("SDI") management entity for the data center 1300. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.
[0090] In at least one embodiment shown in FIG. 13, the framework layer 1320 includes a job scheduler 1332, a configuration manager 1334, a resource manager 1336, and a distributed file system 1338. In at least one embodiment, the framework layer 1320 may include a framework for supporting software 1332 of the software layer 1330 and / or one or more applications 1342 of the application layer 1340. In at least one embodiment, the software 1332 or the application 1342 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1320 may be a kind of free and open-source software web application framework, such as Apache Spark (trademark) (hereinafter referred to as "Spark") that can use the distributed file system 1338 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 1332 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1300. In at least one embodiment, the configuration manager 1334 may be able to configure different layers, such as the software layer 1330 and the framework layer 1320 including Spark and the distributed file system 1338 for supporting large-scale data processing. In at least one embodiment, the resource manager 1336 may be able to manage clustered or grouped computing resources mapped or allocated to support the distributed file system 1338 and the job scheduler 1332. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1314 in the data center infrastructure layer 1310.In at least one embodiment, the resource manager 1336 may manage these mapped or allocated computing resources in cooperation with the resource orchestrator 1312.
[0091] In at least one embodiment, the software 1332 included in the software layer 1330 may include software used by at least a portion of the node C.R. 1316(1) - 1316(N), the grouped computing resources 1314, and / or the distribution file system 1338 of the framework layer 1320. In at least one embodiment, one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0092] In at least one embodiment, the application 1342 included in the application layer 1340 may include one or more types of applications used by at least a portion of the node C.R. 1316(1) - 1316(N), the grouped computing resources 1314, and / or the distribution file system 1338 of the framework layer 1320. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning applications including machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0093] In at least one embodiment, any of configuration manager 1334, resource manager 1336, and resource orchestrator 1312 may implement any number and type of self-correcting actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting actions may enable data center operators of data center 1300 to avoid determining potentially poor configurations and eliminate underutilized and / or underperforming portions of the data center.
[0094] In at least one embodiment, data center 1300 may include tools, services, software, or other resources for training one or more machine learning models or predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 1300. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1300 by using weight parameters calculated by one or more techniques described herein.
[0095] In at least one embodiment, the data center may use a CPU, application specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above may be configured as a service to enable a user to train or perform inference on information, such as image recognition, speech recognition, or other artificial intelligence services.
[0096] 14A illustrates an example of an autonomous vehicle 1400 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1400 (alternatively referred to herein as “vehicle 1400”) may be a passenger vehicle, such as, without limitation, a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, the vehicle 1400 may be a semi-tractor trailer truck for hauling cargo. In at least one embodiment, the vehicle 1400 may be an aircraft, a robotic vehicle, or other type of vehicle.
[0097] Autonomous vehicles may be described in terms of levels of automation as defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers' ("SAE") "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, issued June 15, 2018, Standard No. J3016-201609, issued September 30, 2016, and previous and new versions of this standard). In one or more embodiments, vehicle 1400 may be capable of functionality according to one or more of autonomous driving levels 1 through 5. For example, in at least one embodiment, vehicle 1400 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0098] In at least one embodiment, vehicle 1400 may include components such as, without limitation, a chassis, a vehicle body, wheels (e.g., two, four, six, eight, eighteen, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, vehicle 1400 may include a propulsion system 1450 such as, without limitation, an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, propulsion system 1450 may be coupled to the drive train of vehicle 1400, and the drive train may include, without limitation, a transmission for enabling propulsion of vehicle 1400. In at least one embodiment, propulsion system 1450 may be controlled in response to receiving a signal from throttle / accelerator 1452.
[0099] In at least one embodiment, a steering system 1454 that may include, without limitation, a steering wheel is used to steer vehicle 1400 (e.g., along a desired path or route) when propulsion system 1450 is operating (e.g., when the vehicle is moving). In at least one embodiment, steering system 1454 may receive a signal from a steering actuator 1456. In at least one embodiment, the steering wheel may be optional with respect to fully automated (level 5) functionality. In at least one embodiment, a brake sensor system 1446 may be used to operate vehicle brakes in response to receiving a signal from a brake actuator 1448 and / or a brake sensor.
[0100] Controller 1436, which in at least one embodiment may include without limitation one or more system on chip ("SoC") (not shown in FIG. 14A) and / or a graphics processing unit ("GPU"), provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 1400. For example, in at least one embodiment, controller 1436 may send signals to operate vehicle brakes via brake actuator 1448, to operate steering system 1454 via steering actuator 1456, and to operate propulsion system 1450 via throttle / accelerator 1452. In at least one embodiment, controller 1436 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 1400. In at least one embodiment, the controllers 1436 may include a first controller 1436 for autonomous driving functions, a second controller 1436 for functional safety functions, a third controller 1436 for artificial intelligence functions (e.g., computer vision), a fourth controller 1436 for infotainment functions, a fifth controller 1436 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 1436 may handle two or more of the above functionalities, two or more controllers 1436 may handle a single functionality, and / or some combination thereof.
[0101] In at least one embodiment, controller 1436 provides signals to control one or more components and / or systems of vehicle 1400 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, sensor data may be received from, for example, without limitation, global navigation satellite system ("GNSS") sensors 1458 (e.g., global positioning system sensors), RADAR sensors 1460, ultrasonic sensors 1462, LIDAR sensors 1464, inertial measurement units ("IMUs"), or other sensors. 14A ), a long-range camera (not shown in FIG. 14A ), a mid-range camera (not shown in FIG. 14A ), a speed sensor 1444 (e.g., for measuring the speed of the vehicle 1400), a vibration sensor 1442, a steering sensor 1440, a brake sensor (e.g., as part of a brake sensor system 1446), and / or other types of sensors.
[0102] In at least one embodiment, one or more of the controllers 1436 may receive input (e.g., represented by input data) from the instrument cluster 1432 of the vehicle 1400 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface ("HMI") display 1434, an audible annunciator, a loudspeaker, and / or via other components of the vehicle 1400. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high definition map (not shown in FIG. 14A ), location data (e.g., the location of the vehicle 1400 on a map, etc.), direction, locations of other vehicles (e.g., occupancy grid), information about objects and object conditions sensed by the controller 1436, etc. For example, in at least one embodiment, the HMI display 1434 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about a driving maneuver that the vehicle has made, is making, or will make (e.g., currently changing lanes, leaving exit 34B in 2 miles, etc.).
[0103] In at least one embodiment, vehicle 1400 further includes a network interface 1424, which may use a wireless antenna 1426 and / or a modem for communicating over one or more networks. For example, in at least one embodiment, network interface 1424 may be capable of communicating over Long Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. Additionally, in at least one embodiment, the wireless antenna 1426 may enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low power wide-area networks ("LPWAN") such as LoRaWAN, SigFox, etc.
[0104] In at least one embodiment, the software physical layer (PHY) library 116 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low power wide area networks ("LPWAN") such as LoRaWAN, SigFox, etc.
[0105] 14B illustrates example camera locations and fields of view for the autonomous vehicle 1400 of FIG. 14A according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are an example implementation and are not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or cameras may be positioned at different locations on the vehicle 1400.
[0106] In at least one embodiment, the camera type may include, but is not limited to, a digital camera that may be adapted for use with components and / or systems of vehicle 1400. In at least one embodiment, the camera may operate at automotive safety integrity level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera type may be capable of any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red, clear, clear, clear ("RCCC") color filter array, a red, clear, clear, blue ("RCCB") color filter array, a red, blue, green, clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, a clear pixel camera may be used, such as a camera having a RCCC, RCCB, and / or RBGC color filter array, to increase light sensitivity.
[0107] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance systems (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0108] In at least one embodiment, one or more of the cameras may be attached to a mounting assembly such as a custom-designed (three-dimensional (3D) printed) assembly to eliminate stray light and reflections from inside the vehicle (e.g., reflections reflected from the dashboard to the windshield) that may interfere with the camera's image data capture performance. Referring to a door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed such that the camera mounting plate fits the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. For a side view camera, in at least one embodiment, the camera may also be integrated into the four pillars at each corner of the vehicle in this case.
[0109] In at least one embodiment, a camera (e.g., a front camera) having a field of view that includes a portion of the environment in front of the vehicle 1400 is used for the surrounding view to facilitate identification of the front path and obstacles, and may be used with one or more of the controller 1436 and / or the control SoC to assist in providing information essential for the generation of an occupancy grid and / or the determination of a preferred vehicle path. In at least one embodiment, many of the same ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance, may be performed using the front camera. In at least one embodiment, the front camera may also be used for ADAS functions and systems including, without limitation, other functions such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0110] In at least one embodiment, various cameras including a platform of a monocular camera including, for example, a CMOS:complementary metal oxide semiconductor ("complementary metal oxide semiconductor") color imaging device may be used in a front configuration. In at least one embodiment, a wide-angle camera 1470 may be used to sense objects (e.g., pedestrians, cross traffic, or bicycles) entering the view from the surroundings. Although only one wide-angle camera 1470 is shown in FIG. 14B, in other embodiments, the vehicle 1400 may have any number (including zero) of wide-angle cameras 1470. In at least one embodiment, any number of long-range cameras 1498 (e.g., a pair of stereo cameras for a long-range view) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained thereon. In at least one embodiment, the long-range camera 1498 may also be used for object detection and classification, and basic object tracking.
[0111] In at least one embodiment, any number of stereo cameras 1468 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 1468 may include an integrated control unit with an expandable processing unit, which may provide an integrated Controller Area Network (CAN) or an Ethernet® interface, programmable logic (FPGA) on a single chip, and a multi-core microprocessor. In at least one embodiment, such units may be used to generate a 3D map of the environment of the vehicle 1400, including distance estimation for all points within the image. In at least one embodiment, one or more of the stereo cameras 1468 may include, without limitation, a compact stereo vision sensor that measures the distance from the vehicle 1400 to a target object and can activate functions such as autonomous emergency braking and lane departure warning using the generated information (e.g., metadata), and may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip. In at least one embodiment, in addition to or instead of those described herein, other types of stereo cameras 1468 may be used.
[0112] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view that includes a portion of the environment to the side of the vehicle 1400 may be used for the surrounding view to provide information for creating and updating an occupancy grid and for generating a side collision warning. For example, in at least one embodiment, surround cameras 1474 (e.g., four surround cameras 1474 as shown in FIG. 14B) may be disposed on the vehicle 1400. In at least one embodiment, the surround cameras 1474 may include, without limitation, any number and combination of wide-angle cameras 1470, fisheye cameras, and / or 360-degree cameras, etc. For example, in at least one embodiment, four fisheye cameras may be disposed in front of, behind, and to the sides of the vehicle 1400. In at least one embodiment, the vehicle 1400 may use three surround cameras 1474 (e.g., left, right, and rear), and as a fourth surround camera, one or more other cameras (e.g., a front camera) may be utilized.
[0113] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment behind the vehicle 1400 may be used for parking assistance, the surrounding view, and rear collision warning, and an occupancy grid may be created and updated. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as the front cameras described herein (e.g., a long-range camera 1498, and / or a mid-range camera 1476, a stereo camera 1468), an infrared camera 1472, etc.
[0114] FIG. 14C is a block diagram showing an exemplary system architecture of the autonomous vehicle 1400 of FIG. 14A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1400 of FIG. 14C is shown as being connected via a bus 1402. In at least one embodiment, the bus 1402 may include, without limitation, a CAN data interface (or referred to herein as the (CAN bus)). In at least one embodiment, CAN may be an in-vehicle network used to assist in controlling various features and functions of the vehicle 1400, such as brake actuation, acceleration, brake control, steering, windshield wipers, etc. In at least one embodiment, the bus 1402 may be configured to have dozens or even hundreds of nodes, each having its own unique identifier (e.g., CAN ID). In at least one embodiment, the bus 1402 may be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle state indicators. In at least one embodiment, the bus 1402 may be a CAN bus compliant with ASIL B.
[0115] In at least one embodiment, in addition to or instead of CAN, FlexRay and / or Ethernet® may be used. In at least one embodiment, any number of buses 1402 may be present, including, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 1402 may be used to perform different functions and / or to provide redundancy. For example, a first bus 1402 may be used for a collision avoidance function and a second bus 1402 may be used for actuation control. In at least one embodiment, each bus 1402 may communicate with any of the components of vehicle 1400, and two or more buses 1402 may communicate with the same component. In at least one embodiment, each of any number of system-on-chips (“SoCs”) 1404, each of the controllers 1436, and / or each computer in the vehicle may be accessible to the same input data (e.g., input from sensors of vehicle 1400) and may be connected to a common bus such as a CAN bus.
[0116] In at least one embodiment, vehicle 1400 may include one or more controllers 1436, such as those described herein with respect to FIG. 14A. In at least one embodiment, the controller 1436 may be used for various functions. In at least one embodiment, the controller 1436 may be coupled to any of the various other components and systems of vehicle 1400 and may be used for control of vehicle 1400, the artificial intelligence of vehicle 1400, and / or the infotainment of vehicle 1400.
[0117] In at least one embodiment, vehicle 1400 may include any number of SoCs 1404. Each of the SoCs 1404 may include, without limitation, a central processing unit (“CPU”) 1406, a graphics processing unit (“GPU”) 1408, a processor 1410, a cache 1412, an accelerator 1414, a data store 1416, and / or other components and features not shown. In at least one embodiment, the SoC 1404 may be used to control the vehicle 1400 in various platforms and systems. For example, in at least one embodiment, the SoC 1404 may be incorporated into a system (such as the system of vehicle 1400) having a high definition (“HD”) map 1422 that can obtain map refreshes and / or updates via a network interface 1424 from one or more servers (not shown in FIG. 14C).
[0118] In at least one embodiment, the CPU 1406 may include a CPU cluster, or a CPU complex (or referred to herein as “CCPLEX”). In at least one embodiment, the CPU 1406 may include multiple cores and / or a level 2 (“L2”) cache. For example, in at least one embodiment, the CPU 1406 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 1406 may include four dual-core clusters, where each cluster has a dedicated L2 cache (e.g., a 2MB L2 cache). In at least one embodiment, the CPU 1406 (e.g., CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of the clusters of the CPU 1406 to be activated at any given time.
[0119] In at least one embodiment, one or more of the CPUs 1406 may implement a power management function, which may include, without limitation, one or more of the following features: individual hardware blocks can be automatically clock-gated during idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to the execution of a wait-for-interrupt ("WFI") / wait-for-event ("WFE") instruction; each core can be independently power-gated; when all cores are clock-gated or power-gated, each core cluster can be independently clock-gated; and / or when all cores are power-gated, each core cluster can be independently power-gated. In at least one embodiment, the CPU 1406 may further implement an extended algorithm for managing power states, where the allowed power states and the expected wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. In at least one embodiment, the processing cores may be supported in software with a simple sequence to enter a power state while the work is offloaded to microcode.
[0120] In at least one embodiment, GPU 1408 may include an integrated GPU (or, as referred to herein, an "iGPU"). In at least one embodiment, GPU 1408 may be programmable and may be efficient for parallel workloads. In at least one embodiment, GPU 1408 may use an extended tensor instruction set. In one embodiment, GPU 1408 may include one or more streaming microprocessors, where each streaming microprocessor may include a level 1 ("L1") cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache having a storage capacity of 512 KB). In at least one embodiment, GPU 1408 may include at least eight streaming microprocessors. In at least one embodiment, GPU 1408 may use a compute application programming interface (API). In at least one embodiment, GPU 1408 may use one or more parallel computing platforms and / or programming modules (e.g., NVIDIA's CUDA).
[0121] In at least one embodiment, one or more of the GPUs 1408 may be power-optimized to achieve best performance in automotive and embedded use cases. For example, in one embodiment, the GPU 1408 can be fabricated on a fin field-effect transistor (「FinFET」). In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores partitioned into multiple blocks. For example, without limitation, 64 PF32 cores and 32 PF64 cores can be partitioned into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR cores for deep learning matrix operations, a level zero (「L0」) instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor includes independent parallel data paths for integers and floating points, and realizes efficient execution of the workload by mixing computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor includes an independent thread scheduling function, which may enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance and simplify programming.
[0122] In at least one embodiment, one or more of the GPUs 1408 include high bandwidth memory ("HBM") and / or a 16 GB HBM2 memory subsystem, and in some examples may provide a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to or instead of HBM memory, synchronous graphics random-access memory ("SGRAM"), such as five synchronous random access memories of the graphics double data rate type five ("GDDR5"), may be used.
[0123] In at least one embodiment, the GPU 1408 may include integrated memory technology. In at least one embodiment, address translation services ("ATS") support may be used to enable the GPU 1408 to directly access the page table of the CPU 1406. In at least one embodiment, when the GPU 1408 memory management unit ("MMU") encounters a miss, an address translation request may be sent to the CPU 1406. In at least one embodiment, in response, the CPU 1406 may search its page table for a virtual-to-physical address mapping and send the translation back to the GPU 1408. In at least one embodiment, the integrated memory technology enables a single integrated virtual address space for the memories of both the CPU 1406 and the GPU 1408, thereby simplifying the programming of the GPU 1408 and the porting of applications to the GPU 1408.
[0124] In at least one embodiment, GPU 1408 may include any number of access counters that can record the access frequency of GPU 1408 to the memory of other processors. In at least one embodiment, the access counter may assist in ensuring that memory pages are moved to the physical memory of the processor that most frequently accesses the pages, thereby improving the efficiency of the memory range shared among processors.
[0125] In at least one embodiment, one or more of SoC 1404 may include any number of caches 1412, including those described herein. For example, in at least one embodiment, cache 1412 may include a level 3 (“L3”) cache that is available to both CPU 1406 and GPU 1408 (e.g., connected to both CPU 1406 and GPU 1408). In at least one embodiment, cache 1412 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4 MB or more, depending on the embodiment, although smaller cache sizes may be used.
[0126] In at least one embodiment, one or more of the SoCs 1404 may include one or more accelerators 1414 (e.g., hardware accelerators, software accelerators, or combinations thereof). In at least one embodiment, the SoC 1404 may include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memories. In at least one embodiment, a large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU 1408 and offload some of the tasks of the GPU 1408 (e.g., free up more cycles of the GPU 1408 to execute other tasks). In at least one embodiment, the accelerators 1414 can be used for workloads (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.) that are stable enough to accept acceleration. In at least one embodiment, the CNN may include region-based, i.e., regional convolutional neural networks (“RCNNs”), and (e.g., used for object detection) fast RCNNs, or other types of CNNs.
[0127] In at least one embodiment, the accelerator 1414 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include, without limitation, one or more Tensor processing units (TPUs), which may be further configured to provide 10 trillion operations per second for deep learning applications and inferences. In at least one embodiment, the TPU may be an accelerator configured and optimized to execute image processing functions (e.g., CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating point operations, as well as for inferences. In at least one embodiment, the design of the DLA can improve the performance per millimeter compared to typical general-purpose GPUs, typically far exceeding the performance of CPUs. In at least one embodiment, the TPU may execute several functions, including, for example, a single instance of a convolution function and a post-processing function that support INT8, INT16, and FP16 data types for both features and weights. In at least one embodiment, the DLA may execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including, without limitation, CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection, identification, and detection using data from a microphone 1496, CNNs for face recognition and vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events.
[0128] In at least one embodiment, the DLA may execute any function of the GPU 1408. For example, by using an inference accelerator, the designer may target either the DLA or the GPU 1408 for any function. For example, in at least one embodiment, the designer may concentrate the processing of CNNs and floating-point operations on the DLA and leave other functions to the GPU 1408 and / or other accelerators 1414.
[0129] In at least one embodiment, the accelerator 1414 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (referred to herein alternatively as a computer vision accelerator). In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS) 1438, autonomous driving, augmented reality (AR) applications, and / or virtual reality (VR) applications. The PVA may maintain a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors, without limitation.
[0130] In at least one embodiment, the RISC core may interact with an image sensor (e.g., the image sensor of any of the cameras described herein), and / or an image signal processor, etc. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core can include an instruction cache and / or tightly coupled RAM.
[0131] In at least one embodiment, the DMA may enable the components of the PVA to access system memory independently of the CPU 1406. In at least one embodiment, the DMA may support any number of features used to optimize the PVA, including but not limited to multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0132] In at least one embodiment, the vector processor can be a programmable processor that may be designed to efficiently and flexibly execute programming for computer vision algorithms, and provides a signal processing function. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit (“VPU”), an instruction cache, and / or vector memory (e.g., “VMEM”). In at least one embodiment, the VPU may include a digital signal processor such as, for example, a single instruction multiple data (“SIMD”), a very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, throughput and speed may be improved by a combination of SIMD and VLIW.
[0133] In at least one embodiment, each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to execute independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on consecutive images or on portions of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster and any number of vector processors may be included in each of the PVAs. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to enhance the overall security of the system.
[0134] In at least one embodiment, the accelerator 1414 (e.g., a hardware acceleration cluster) includes an on-chip computer vision network and static random access memory (“SRAM”), and may provide high-bandwidth, low-latency SRAM for the accelerator 1414. In at least one embodiment, the on-chip memory may include at least 4MB of SRAM, for example, but not limited to, consisting of eight field-configurable memory blocks, which may be accessible from either the PVA or the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and the DLA may access the memory via a backbone that provides high-speed access to the memory for the PVA and the DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and the DLA to the memory (e.g., using the APB).
[0135] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide a ready signal and a valid signal before transmitting any control signal / address / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signal / address / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, the interface may comply with the standards of the International Organization for Standardization (“ISO”) 26262 or the International Electrotechnical Commission (“IEC”) 61508, although other standards and protocols may be used.
[0136] In at least one embodiment, one or more of the SoCs 1404 may include a hardware accelerator for real-time ray tracing. In at least one embodiment, the hardware accelerator for real-time ray tracing is used to quickly and efficiently determine the position and extent of an object (e.g., within a world model) for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of a SONAR system, for simulation of general waveform propagation, for comparison with LIDAR data for localization and / or other functions, and / or for real-time visualization simulation for other uses.
[0137] In at least one embodiment, the accelerator 1414 (e.g., a hardware accelerator cluster) has various uses for autonomous driving. In at least one embodiment, the PVA may be a programmable vision accelerator that can be used in the main processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well-suited to algorithm domains that require predictable processing with low power and low latency. In other words, the PVA functions well for semi-dense or dense regular calculations that require a predictable runtime with low latency and low power, even with a small data set. In at least one embodiment, in an autonomous vehicle such as the vehicle 1400, the PVA is designed to execute conventional computer vision algorithms because they are effective for object detection and integer arithmetic operations.
[0138] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using PVA. In at least one embodiment, in some examples, an algorithm based on semi-global matching may be used, but this is not limiting. In at least one embodiment, applications for level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may perform computer stereo vision functions on inputs from two monocular cameras.
[0139] In at least one embodiment, high-density optical flow may be performed using PVA. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and for example, processed time-of-flight data is provided by processing raw time-of-flight data.
[0140] In at least one embodiment, for example, without limitation, a DLA may be used to execute any type of network for enhancing control and driving safety, including a neural network that outputs a measure of reliability for each object detection. In at least one embodiment, reliability may be expressed or interpreted as the probability of each detection compared to other detections, or as providing its relative "weight". In at least one embodiment, reliability enables the system to make further decisions regarding which detections should be considered positive detections rather than false detections. In at least one embodiment, the system may set a threshold for reliability and consider only detections that exceed the threshold as positive detections. In embodiments where automatic emergency braking ("AEB") is used, false detections would cause the vehicle to automatically apply the emergency brake, which is clearly undesirable. In at least one embodiment, very highly reliable detections may be considered as triggers for AEB. In at least one embodiment, the DLA may execute a neural network to regress reliability values. In at least one embodiment, the neural network may take as its inputs at least some subset of parameters such as, among other things, the dimensions of the bounding box, ground estimation obtained (e.g., from another subsystem), the output from the IMU sensor 1466 correlated with the orientation of the vehicle 1400, distance, and the 3D location estimation of the object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 1464 or the RADAR sensor 1460).
[0141] In at least one embodiment, one or more of the SoCs 1404 may include a data store 1416 (e.g., memory). In at least one embodiment, the data store 1416 may be on-chip memory of the SoC 1404, and this memory may store neural networks executed on the GPU 1408 and / or DLA. In at least one embodiment, the capacity of the data store 1416 may be large enough to store multiple instances of the neural network for redundancy and security. In at least one embodiment, the data store 1412 may include an L2 or L3 cache.
[0142] In at least one embodiment, one or more of the SoCs 1404 may include any number of processors 1410 (e.g., embedded processors). In at least one embodiment, the processor 1410 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 1404 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance in transitioning the system to a low power state, management of the thermal and temperature sensors of the SoC 1404, and / or management of the power state of the SoC 1404. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1404 may use the ring oscillator to detect the temperature of the CPU 1406, GPU 1408, and / or accelerator 1414. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine, put the SoC 1404 in a low power state, and / or put the vehicle 1400 in a driver-safety stop mode (e.g., safely stop the vehicle 1400).
[0143] In at least one embodiment, the processor 1410 may further include a set of embedded processors that can serve as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via a multi-interface and a wide variety of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0144] In at least one embodiment, the processor 1410 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0145] In at least one embodiment, the processor 1410 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for handling safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, support peripherals (such as timers and interrupt controllers, etc.), and / or routing logic. In the safety mode, in at least one embodiment, two or more cores may operate in a lockstep mode and function as a single core having comparison logic for detecting any differences between these operations. In at least one embodiment, the processor 1410 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, the processor 1410 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of the camera processing pipeline.
[0146] In at least one embodiment, the processor 1410 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image in the window of the playback device. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1470, the surround camera 1474, and / or the in-cabin monitoring camera / sensor. In at least one embodiment, the in-cabin monitoring camera / sensor is preferably monitored by a neural network running on another instance of the SoC 1404, which is configured to identify events in the cabin and respond accordingly. In at least one embodiment, the in-cabin system may perform lip reading, among other things, to activate cellular service, make a call, write an email, change the destination of the vehicle, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable otherwise.
[0147] In at least one embodiment, the video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when there is motion in the video, the noise reduction appropriately weights the spatial information and downweights the information provided by adjacent frames. In at least one embodiment, when an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from the previous image to reduce the noise in the current image.
[0148] In at least one embodiment, the video image synthesizer may also be configured to perform stereo parallelization on the input stereo lens frames. In at least one embodiment, the video image synthesizer may further be used to synthesize a user interface when the operating system desktop is in use, and the GPU 1408 need not continuously render new surfaces. In at least one embodiment, when the GPU 1408 is powered on and performing active 3D rendering, the video image synthesizer may be used to offload the GPU 1408 to improve performance and responsiveness.
[0149] In at least one embodiment, one or more of the SoCs 1404 may further include a camera serial interface of a mobile industry processor interface ("MIPI") for receiving inputs from video and cameras, a high-speed interface, and / or a video input block that may be used for input functions of cameras and related pixels. In at least one embodiment, one or more of the SoCs 1404 may further include an input / output controller, which may be controlled by software and may be used to receive I / O signals that are not tied to a specific role.
[0150] In at least one embodiment, one or more of the SoCs 1404 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. The SoC 1404 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet®), sensors (e.g., LIDAR sensor 1464, RADAR sensor 1460, etc., which may be connected via Ethernet®), data from bus 1402 (e.g., the speed, steering wheel position, etc., of vehicle 1400), data from GNSS sensor 1458 (e.g., connected via Ethernet® or CAN bus), etc. In at least one embodiment, one or more of the SoCs 1404 may further include a dedicated high-performance large-capacity storage controller, which may include its own DMA engine and may be used to free the CPU 1406 from routine data management tasks.
[0151] In at least one embodiment, the SoC 1404 may be an end-to-end platform with a flexible architecture spanning automation levels 3 to 5, thereby providing a comprehensive functional safety architecture that leverages computer vision and ADAS techniques to obtain diversity and redundancy and use them efficiently, and a flexible and reliable driving software stack is provided along with deep learning tools. In at least one embodiment, the SoC 1404 is faster, more reliable, and more energy- and space-efficient than conventional systems. For example, in at least one embodiment, when the accelerator 1414 is combined with the CPU 1406, GPU 1408, and data store 1416, a fast and efficient platform for level 3 to 5 autonomous vehicles can be realized.
[0152] In at least one embodiment, the computer vision algorithm may be executed on a CPU, and this algorithm may be configured using a high-level programming language such as the C programming language to execute various processing algorithms over various visual data. However, in at least one embodiment, the CPU often cannot meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs cannot execute in real time complex object detection algorithms used in ADAS applications within a vehicle and in realistic level 3-5 autonomous vehicles.
[0153] The embodiments described herein can enable level 3-5 autonomous driving functions by allowing multiple neural networks to be executed simultaneously and / or sequentially and combining the results. For example, in at least one embodiment, the CNN running on the DLA or an individual GPU (e.g., GPU 1420) may include text and word recognition and enable a supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs and provide a semantic understanding of the signs, and can pass that semantic understanding to a path planning module running on the CPU complex.
[0154] In at least one embodiment, for level 3, 4, or 5 operation, multiple neural networks may be executed simultaneously. For example, in at least one embodiment, a warning sign that reads "Caution: Frozen when flashing" in conjunction with the electro-optical may be interpreted separately or collectively by several neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), the text "Frozen when flashing" may be interpreted by a second introduced neural network, and when a flashing light is detected, this neural network notifies the vehicle's (preferably running on the CPU complex) route planning software that a frozen state exists. In at least one embodiment, the flashing light may be identified by operating a third introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may be executed simultaneously, such as within the DLA and / or on the GPU 1408.
[0155] In at least one embodiment, a CNN for face recognition and vehicle owner identification may use data from a camera sensor to identify the presence of an approved driver and / or owner of the vehicle 1400. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle, turn on the lights when the owner approaches the driver's door, and make the vehicle inoperable in security mode when the owner leaves the vehicle. In this way, the SoC 1404 realizes security against theft and / or carjacking.
[0156] In at least one embodiment, the CNN for detecting and identifying emergency vehicles may detect and identify the sirens of emergency vehicles using data from microphone 1496. In at least one embodiment, SoC 1404 classifies environmental and urban sounds and uses a CNN to classify visual data. In at least one embodiment, the CNN executed on the DLA is trained to identify the relative speed at which an emergency vehicle is approaching (e.g., by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area where the vehicle is operating, identified by GNSS sensor 1458. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and when in the United States, attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine is used to reduce the speed of the vehicle, pull over to the side of the road, stop the vehicle, and / or use ultrasonic sensor 1462 to idle the vehicle until the emergency vehicle has passed.
[0157] In at least one embodiment, vehicle 1400 may include a CPU 1418 (e.g., a discrete CPU or dCPU), which may be coupled to SoC 1104 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU 1418 may include, for example, an X86 processor. CPU 1418 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and SoC 1404 and / or monitoring the state and health of controller 1436 and / or the in-vehicle infotainment system (the "IVI SoC") 1430 on the chip.
[0158] In at least one embodiment, vehicle 1400 may include a GPU 1420 (e.g., a discrete GPU or dGPU), which may be coupled to the SoC 1404 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, the GPU 1420 may provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and may be used to train and / or update a neural network based at least in part on inputs (e.g., sensor data) from the sensors of the vehicle 1400.
[0159] In at least one embodiment, vehicle 1400 may further include a network interface 1424, which may include, without limitation, a wireless antenna 1426 (e.g., one or more wireless antennas 1426 for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, the network interface 1424 may be used to enable a wireless connection via the Internet with a cloud (e.g., a server and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device). In at least one embodiment, a direct link may be established between vehicle 140 and other vehicles for communicating with other vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. In at least one embodiment, the vehicle-to-vehicle communication link may provide information about vehicles in the vicinity of vehicle 1400 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1400) to vehicle 1400. In at least one embodiment, the aforementioned functionality may be part of a cooperative adaptive cruise control function of vehicle 1400.
[0160] In at least one embodiment, network interface 1424 may include a system-on-chip (SoC) that provides modulation and demodulation functions to enable the controller 1436 to communicate via a wireless network. In at least one embodiment, network interface 1424 may include a radio frequency (RF) front end for up-conversion from baseband to RF and down-conversion from RF to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion can be performed by well-known processes and / or using a superheterodyne process. In at least one embodiment, the RF front end functions may be provided by a separate chip. In at least one embodiment, the network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0161] In at least one embodiment, vehicle 1400 may further include a data store 1428, which may include off-chip (e.g., not on SoC 1404) storage without limitation. In at least one embodiment, data store 1428 may include one or more storage elements including, without limitation, RAM, SRAM, dynamic random access memory ("DRAM"), video random-access memory ("VRAM"), flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0162] In at least one embodiment, vehicle 1400 may further include a GNSS sensor 1458 (e.g., a GPS and / or assisted GPS sensor) to assist with mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1458 may be used, including for example without limitation a GPS that uses a USB connector having a bridge from Ethernet (registered trademark) to serial (e.g., RS-232).
[0163] In at least one embodiment, vehicle 1400 may further include a RADAR sensor 1460. The RADAR sensor 1460 may be used by vehicle 1400 to perform long-range vehicle detection, even in darkness and / or harsh weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. The RADAR sensor 1460 may use the CAN and / or bus 1402 for control (e.g., to transmit data generated by the RADAR sensor 1460) and to access object tracking data, and in some examples, can access Ethernet (registered trademark) to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example without limitation, the RADAR sensor 1460 may be suitable for use in front, rear, and side RADAR. In at least one embodiment, one or more of the RADAR sensors 1460 are pulse Doppler RADAR sensors.
[0164] In at least one embodiment, the RADAR sensor 1460 may include different configurations, such as a narrow field of view for long distances, a wide field of view for short distances, and short distances covering the sides. In at least one embodiment, the long-range RADAR may be used for an adaptive cruise control function. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within a range of 250 m realized by two or more independent scans. In at least one embodiment, the RADAR sensor 1460 may be made easier to distinguish between static and moving objects and may be used by the ADAS system 1438 for emergency braking assistance and forward collision warning. In at least one embodiment, the sensor 1460 included in the long-range RADAR system may include, without limitation, a plurality of (e.g., six or more) fixed RADAR antennas and a monostatic multimode RADAR having high-speed CAN and FlexRay interfaces. In at least one embodiment, when there are six antennas, the four central antennas may generate a concentrated beam pattern designed to record the surroundings of the vehicle 1400 at a higher speed with minimal interference from adjacent lanes. In at least one embodiment, the other two antennas may expand the field of view and enable rapid detection of vehicles entering or exiting the lane of the vehicle 1400.
[0165] In at least one embodiment, the mid-range RADAR system may include, for example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, the short-range RADAR system may include any number of RADAR sensors 1460 designed to be installed at both ends of the rear bumper, without limitation. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor the rear and the blind spots adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in the ADAS system 1438 for blind spot detection and / or lane change assistance.
[0166] In at least one embodiment, vehicle 1400 may further include an ultrasonic sensor 1462. In at least one embodiment, the ultrasonic sensor 1462 may be disposed in front of, behind, and / or on the side of the vehicle 1400 and may be used for parking assistance and / or for generating and updating an occupancy grid. In at least one embodiment, various ultrasonic sensors 1462 may be used, and different ultrasonic sensors 1462 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor 1462 may operate at a functional safety level of ASIL B.
[0167] In at least one embodiment, vehicle 1400 may include a LIDAR sensor 1464. The LIDAR sensor 1464 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor 1464 may be at a functional safety level of ASIL B. In at least one embodiment, vehicle 1400 may include a plurality of LIDAR sensors 1464 (e.g., two, four, six, etc.), and these sensors may use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).
[0168] In at least one embodiment, the LIDAR sensor 1464 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 1464 may, for example, have a advertised range of approximately 100 m, an accuracy of 2 cm to 3 cm, and support a 100 Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors 1464 may be used. In such embodiments, the LIDAR sensor 1464 may be implemented as a small device that can be incorporated in front of, behind, to the sides of, and / or at the corners of the vehicle 1400. In at least one embodiment, the LIDAR sensor 1464 of such embodiments may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, in a range of 200 m. In at least one embodiment, the LIDAR sensor 1464 mounted in the front may be configured to provide a horizontal field of view of 45 degrees to 135 degrees.
[0169] In at least one embodiment, LIDAR technologies such as 3D flash LIDAR may also be used. The 3D flash LIDAR uses a laser flash as a transmission source to irradiate the surroundings of the vehicle 1400 up to approximately 200 m at most. In at least one embodiment, the flash LIDAR unit includes, without limitation, a receptor that records the transit time of the laser pulses and the reflected light at each pixel, which corresponds to the range from the vehicle 1400 to the object. In at least one embodiment, the flash LIDAR enables a very accurate and distortion-free surrounding image to be generated for each laser flash. In at least one embodiment, four flash LIDARs may be introduced, one on each side of the vehicle 1400. In at least one embodiment, the 3D flash LIDAR system includes, without limitation, a LIDAR camera of a semiconductor 3D staring array (such as a non-scanning LIDAR device) without moving parts other than a fan. In at least one embodiment, the flash LIDAR device may use class I (eye-safe) laser pulses of 5 nanoseconds per frame and capture the reflected laser light in the form of a 3D range point cloud and co-registered intensity data.
[0170] In at least one embodiment, the vehicle may further include an IMU sensor 1466. In at least one embodiment, the IMU sensor 1466 may be positioned at the center of the rear axle of the vehicle 1400 in at least one embodiment. In at least one embodiment, the IMU sensor 1466 may include, for example without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other types of sensors. In at least one embodiment for a 6-axis application, the IMU sensor 1466 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment for a 9-axis application, the IMU sensor 1466 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0171] In at least one embodiment, the IMU sensor 1466 may be implemented as a small, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical systems (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 1466 enables the vehicle 1400 to estimate its orientation without the need for input from a magnetic sensor by directly observing changes in velocity and correlating them from the GPS to the IMU sensor 1466. In at least one embodiment, the IMU sensor 1466 and the GNSS sensor 1458 may be combined in a single integrated unit.
[0172] In at least one embodiment, the vehicle 1400 may include a microphone 1496 installed inside and / or around the vehicle 1400. In at least one embodiment, the microphone 1496 may be used, among other things, for the detection and identification of emergency vehicles.
[0173] In at least one embodiment, vehicle 1400 may further include any number of camera types including stereo camera 1468, wide-angle camera 1470, infrared camera 1472, surround camera 1474, long-range camera 1498, mid-range camera 1476, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 1400. In at least one embodiment, the type of cameras used may vary depending on vehicle 1400. In at least one embodiment, any combination of camera types may be used to provide the necessary field of view around vehicle 1400. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, vehicle 1400 may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, or another number of cameras. In at least one embodiment, the cameras may support, by way of non-limiting example, Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet®. In at least one embodiment, each camera is described in more detail above herein with respect to FIGS. 14A and 14B.
[0174] In at least one embodiment, vehicle 1400 may further include vibration sensor 1442. In at least one embodiment, vibration sensor 1442 may measure the vibration of components of vehicle 1400 such as an axle. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 1442 are used, the difference in vibration may be used to determine the amount of friction or slip on the road surface (e.g., if there is a vibration difference between a powered axle and a freely rotating axle).
[0175] In at least one embodiment, vehicle 1400 may include an ADAS system 1438. The ADAS system 1438 may include, without limitation, a SoC in some examples. In at least one embodiment, the ADAS system 1438 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward crash warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions.
[0176] In at least one embodiment, the ACC system may use a RADAR sensor 1460, a LIDAR sensor 1464, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to the vehicle immediately in front of vehicle 1400 and automatically adjusts the speed of vehicle 1400 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 1400 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications such as LC and CW.
[0177] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) to the network interface 1424 and / or the wireless antenna 1426. In at least one embodiment, a direct link may be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link may be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, the concept of V2V communication provides information about the immediately preceding vehicle (e.g., a vehicle in the same lane immediately in front of vehicle 1400), and the concept of I2V communication provides information about traffic even further ahead. In at least one embodiment, the CACC system may include either or both of the I2V and V2V information sources. In at least one embodiment, having information about the vehicle in front of vehicle 1400 can further enhance the reliability of the CACC system, make the traffic flow smoother, and have the potential to reduce traffic jams on the road.
[0178] In at least one embodiment, the FCW system is designed to advise the driver of a hazard, whereby the driver can take corrective measures. In at least one embodiment, the FCW system uses a front camera and / or the RADAR sensor 1460, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as a display, speaker, and / or a vibrating component. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibrations, and / or quick brake pulses.
[0179] In at least one embodiment, the AEB system may detect an imminent frontal collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 1460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the predicted collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-collision braking.
[0180] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibrations of the steering wheel or seat, to advise the driver when vehicle 1400 crosses a lane marking. In at least one embodiment, the LDW system does not activate if the driver indicates an intentional lane departure by activating the turn indicator. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variant of the LDW system. The LKA system provides steering input or brake control to correct vehicle 1400 when vehicle 1400 begins to drift out of a lane.
[0181] In at least one embodiment, the BSW system detects a vehicle in a blind spot of a motor vehicle and warns the driver. In at least one embodiment, the BSW system may provide visual, audible, and / or tactile alerts to indicate that a merge or lane change is not safe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses a turn indicator. In at least one embodiment, the BSW system may use a rear camera and / or a RADAR sensor 1460 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, and these dedicated processors, DSPs, FPGAs, and / or ASICs are electrically coupled to feedback to the driver such as a display, a speaker, and / or a vibration component.
[0182] In at least one embodiment, the RCTW system may provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera when the vehicle 1400 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear RADAR sensors 1460, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver such as a display, a speaker, and / or a vibration component.
[0183] In at least one embodiment, conventional ADAS systems may tend to produce false detection results, which can be annoying and distracting to the driver, but usually are not a major issue. This is because conventional ADAS systems are designed to advise the driver and enable the driver to determine whether a safety-critical condition actually exists and how to appropriately respond to it. In at least one embodiment, if the results conflict, the vehicle 1400 itself determines whether to follow the results from the primary computer (e.g., the first controller 1436) or the results from the secondary computer (e.g., the second controller 1436). For example, in at least one embodiment, the ADAS system 1438 may be a backup and / or secondary computer for resisting perceptual information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may execute various software for redundancy on the hardware components to detect perceptual errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1438 may be provided to the monitoring MCU. In at least one embodiment, when the output from the primary computer conflicts with the output from the secondary computer, the monitoring MCU determines how to reconcile the conflict to ensure safe operation.
[0184] In at least one embodiment, the primary computer may be configured to provide a reliability score indicating the reliability of the selected result of the primary computer to the monitoring MCU. In at least one embodiment, if the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, if the reliability score does not meet the threshold and the primary computer and the secondary computer indicate different results (e.g., conflict), the monitoring MCU may mediate between the computers to determine an appropriate result.
[0185] In at least one embodiment, a neural network trained and configured to determine, at least in part based on outputs from a primary computer and a secondary computer, conditions under which the secondary computer provides a false alarm may be configured to be executed by a monitoring MCU. In at least one embodiment, the neural network of the monitoring MCU may learn when the output of the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metallic object, such as a drain grating or manhole cover, that is not actually a hazard, which triggers an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when bicycles or pedestrians are present and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or a GPU suitable for executing the neural network along with an associated memory. In at least one embodiment, the monitoring MCU may comprise components of the SoC1404 and / or be included as a component thereof.
[0186] In at least one embodiment, the ADAS system 1438 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and the presence of the neural network in the monitoring MCU may improve reliability, safety, and performance. For example, in at least one embodiment, due to various implementations and intentional non-identities, the overall error tolerance of the system is increased, particularly for errors caused by the functions of software (or the software-hardware interface). For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU may have a higher level of confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.
[0187] In at least one embodiment, the output of the ADAS system 1438 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1438 indicates a forward collision warning due to a preceding object, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have a trained, and thus risk-reducing misdetection, neural network as described herein.
[0188] In at least one embodiment, vehicle 1400 may further include an infotainment SoC 1430 (e.g., in-vehicle infotainment system (IVI)). Although the infotainment system 1430 is illustrated and described as an SoC, in at least one embodiment, it may not be an SoC and may include two or more individual components without limitation. In at least one embodiment, the infotainment SoC 1430 may include, without limitation, a combination of hardware and software, and this combination may be used to provide the vehicle 1400 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calls), network connections (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door opening and closing, air filter information, etc.). For example, the infotainment SoC 1430 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connections, a carputer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, a heads-up display ("HUD"), an HMI display 1434, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, information from the ADAS system 1438, autonomous driving information such as vehicle operation plans and trajectories, ambient environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information (e.g., visual and / or auditory) may further be provided to the user of the vehicle using the infotainment SoC 1430.
[0189] In at least one embodiment, the infotainment SoC 1430 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1430 may communicate with other devices, systems, and / or components of the vehicle 1400 via a bus 1402 (e.g., a CAN bus, Ethernet®, etc.). In at least one embodiment, the infotainment SoC 1430 may be coupled to a monitoring MCU, such that when the primary controller 1436 (e.g., the primary and / or backup computer of the vehicle 1400) fails, the GPU of the infotainment system may execute some self-driving functions. In at least one embodiment, the infotainment SoC 1430 may put the vehicle 1400 into a driver-safety stop mode as described herein.
[0190] In at least one embodiment, the vehicle 1400 may further include an instrument cluster 1432 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1432 may include, without limitation, a controller and / or a supercomputer (e.g., an individual controller or supercomputer). In at least one embodiment, the instrument cluster 1432 may include any number and combination of instrument sets, without limitation, a speedometer, a fuel level, a hydraulic pressure, a tachometer, an odometer, a direction indicator, a shift lever position indicator, a seat belt warning light, a parking brake warning light, an engine failure light, an auxiliary restraint system (e.g., an airbag) information, a light control, a safety system control, a navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1430 and the instrument cluster 1432. In at least one embodiment, the instrument cluster 1432 may be included as part of the infotainment SoC 1430, or vice versa.
[0191] FIG. 14D is a diagram of a system 1476 for communicating between a cloud-based server and the autonomous vehicle 1400 of FIG. 14A according to at least one embodiment. In at least one embodiment, the system 1476 may include, without limitation, a server 1478, a network 1490, and any number and type of vehicles including the vehicle 1400. The server 1478 may include, without limitation, a plurality of GPUs 1484(A)-1484(H) (collectively referred to herein as GPU 1484), PCIe switches 1482(A)-1482(H) (collectively referred to herein as PCIe switch 1482), and / or CPUs 1480(A)-1480(B) (collectively referred to herein as CPU 1480). The GPUs 1484, CPUs 1480, and PCIe switches 1482 may be interconnected by high-speed interconnects such as, without limitation, an NVLink interface 1488 and / or a PCIe connection 1486 developed by NVIDIA. In at least one embodiment, the GPUs 1484 are connected to each other via an NVLink and / or an NVSwitch SoC, and the GPU 1484 and the PCIe switch 1482 are connected via a PCIe interconnect. In at least one embodiment, eight GPUs 1484, two CPUs 1480, and four PCIe switches 1482 are illustrated, but this is not limiting. In at least one embodiment, each of the servers 1478 may include, without limitation, any number of GPUs 1484, CPUs 1480, and / or PCIe switches 1482 in any combination. For example, in at least one embodiment, the server 1478 may include eight, sixteen, thirty-two, and / or more GPUs 1484 each.
[0192] In at least one embodiment, server 1478 may receive, via network 1490, image data representing an image indicating an unexpected or changed road condition, such as recently started road construction, from a vehicle. In at least one embodiment, server 1478 may transmit, via network 1490, neural network 1492, updated neural network 1492, and / or map information 1494 including information regarding traffic conditions and road conditions, without limitation, to a vehicle. In at least one embodiment, the update of map information 1494 may include, without limitation, updates to HD map 1422, such as information regarding construction sites, holes, detours, floods, and / or other obstacles. In at least one embodiment, neural network 1492, updated neural network 1492, and / or map information 1494 may be obtained from new training and / or experience represented in data received from any number of vehicles in the environment and / or may be obtained based at least in part on training performed at a data center (e.g., using server 1478 and / or other servers).
[0193] In at least one embodiment, a machine learning model (e.g., a neural network) may be trained using server 1478 based at least in part on training data. In at least one embodiment, the training data may be generated by a vehicle and / or may be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or otherwise pre-processed (e.g., if the associated neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or pre-processed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by a vehicle (e.g., transmitted to the vehicle via network 1490) and / or the machine learning model may be used by server 1478 to remotely monitor the vehicle.
[0194] In at least one embodiment, server 1478 may receive data from a vehicle and apply the data to a state-of-the-art real-time neural network so as to enable real-time intelligent inference. In at least one embodiment, server 1478 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1484, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 1478 may include a deep learning infrastructure that uses a data center powered by a CPU.
[0195] In at least one embodiment, the deep learning infrastructure of server 1478 may be capable of high-speed real-time inference and may use that capability to evaluate and verify the soundness of the processor, software, and / or associated hardware of vehicle 1400. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1400, such as a series of images and / or objects located by vehicle 1400 in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify an object and compare it to the object identified by vehicle 1400, and if the results do not match and the deep learning infrastructure concludes that the AI of vehicle 1400 is malfunctioning, server 1478 may take control from the fail-safe computer of vehicle 1400, notify the occupants, and send a signal to vehicle 1400 instructing it to complete a safe parking maneuver.
[0196] In at least one embodiment, server 1478 may include GPU 1484 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT3). In at least one embodiment, by combining a server powered by a GPU with inference acceleration, real-time response can be enabled. In at least one embodiment, a server powered by a CPU, FPGA, and other processors may be used for inference, such as when performance is not as critical. In at least one embodiment, hardware structure 1315 is used to execute one or more embodiments. Details regarding hardware structure 1315 are provided herein in conjunction with FIG. 13A and / or FIG. 13B.
[0197] Computer system FIG. 15 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof 1500 formed with a processor that may include an execution unit for executing instructions according to at least one embodiment. In at least one embodiment, computer system 1500 may include, without limitation, components such as processor 1502 for using an execution unit that includes logic for executing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 1500 may include a processor such as a PENTIUM® processor family, Xeon™, Itanium® , XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 1500 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0198] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.
[0199] In at least one embodiment, computer system 1500 may include a processor 1502, without limitation, which may include one or more execution units 1508 for performing training and / or inference of a machine learning model according to the techniques described herein, without limitation. In at least one embodiment, system 15 is a single-processor desktop or server system, while in another embodiment, system 15 may be a multi-processor system. In at least one embodiment, processor 1502 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1502 may be coupled to a processor bus 1510, which may transmit digital signals between processor 1502 and other components within computer system 1500.
[0200] In at least one embodiment, the processor 1502 may include, without limitation, a level 1 ("L1") internal cache memory ("cache") 1504. In at least one embodiment, the processor 1502 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to the processor 1502. Other embodiments may also include combinations of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, the register file 1506 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.
[0201] In at least one embodiment, the processor 1502 also includes an execution unit 1508 that includes, without limitation, logic for performing integer and floating point operations. In at least one embodiment, the processor 1502 may also include a microcode ("u-code") read only memory ("ROM") that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 1508 may include logic for handling a packed instruction set 1509. In at least one embodiment, by including the packed instruction set 1509 in the instruction set of the general purpose processor 1502 along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be executed using the packed data of the general purpose processor 1502. In one or more embodiments, by executing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.
[0202] In at least one embodiment, execution unit 1508 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1500 may include memory 1520, without limitation. In at least one embodiment, memory 1520 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 1520 may store instructions 1519 and / or data 1521 represented by data signals that may be executed by processor 1502.
[0203] In at least one embodiment, the system logic chip may be coupled to the processor bus 1510 and the memory 1520. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 1516, and the processor 1502 may communicate with the MCH 1516 via the processor bus 1510. In at least one embodiment, the MCH 1516 may provide a high bandwidth memory path 1518 to the memory 1520 for storing instructions and data and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1516 may direct data signals between the processor 1502, the memory 1520, and other components of the computer system 1500, and may bridge data signals between the processor bus 1510, the memory 1520, and the system I / O 1522. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 1516 may be coupled to the memory 1520 via the high bandwidth memory path 1518, and the graphics / video card 1512 may be coupled to the MCH 1516 via an Accelerated Graphics Port (“AGP”) interconnect 1514.
[0204] In at least one embodiment, computer system 1500 may use a system I / O 1522, which is a proprietary hub interface bus for coupling MCH 1516 to an I / O controller hub (“ICH”) 1530. In at least one embodiment, ICH 1530 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 1520, the chipset, and processor 1502. By way of example, audio controller 1529, firmware hub (“Flash BIOS”) 1528, wireless transceiver 1526, data storage 1524, legacy I / O controller 1523 including user input and keyboard interface, serial expansion ports such as universal serial bus (“USB”), and network controller 1534 may be included, without limitation. In at least one embodiment, data storage 1524 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0205] In at least one embodiment, FIG. 15 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 15 may show an exemplary system-on-chip (“SoC”). In at least one embodiment, the devices shown in FIG. cc may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1500 may be interconnected using a compute express link (CXL) interconnect.
[0206] FIG. 16 is a block diagram showing an electronic device 1600 for utilizing a processor 1610 according to at least one embodiment. In at least one embodiment, the electronic device 1600 may be, without limitation, for example, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0207] In at least one embodiment, the system 1600 may include, without limitation, a processor 1610 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1610 is coupled using a bus or interface such as an I°C bus, a system management bus (“SMBus”), a low pin count (“LPC”) bus, a serial peripheral interface (“SPI”), a high definition audio (“HDA”) bus, a serial advance technology attachment (“SATA”) bus, a universal serial bus (“USB”) (versions 1, 2, 3), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, FIG. 16 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 16 may show an exemplary system on a chip (“SoC”). In at least one embodiment, the devices shown in FIG. 16 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 16 may be interconnected using a Compute Express Link (CXL) interconnect.
[0208] In at least one embodiment, FIG. 16 shows a display 1624, a touch screen 1625, a touch pad 1630, a near field communications unit (NFC) 1645, a sensor hub 1640, a thermal sensor 1646, an express chipset (EC) 1635, a trusted platform module (TPM) 1638, a BIOS / firmware / flash memory (BIOS, FW flash) 1622, a DSP 1660, a drive (SSD or HDD) 1620 such as a solid state disk (SSD) or a hard disk drive (HDD), a wireless local area network unit (WLAN) 1650, a Bluetooth unit 1652, a wireless wide area network unit (WWAN) 1656, a global positioning system (GPS) 1655, a camera (USB3.0 camera) 1654 such as a USB3.0 camera, or a low power double data rate (LPDDR) memory unit (LPDDR3) 1615 implemented, for example, in accordance with the LPDDR3 standard. These components may be implemented in any suitable manner, respectively.
[0209] In at least one embodiment, other components may be communicatively coupled to the processor 1610 via the components described above. In at least one embodiment, an accelerometer 1641, an ambient light sensor ("ALS") 1642, a compass 1643, and a gyroscope 1644 may be communicatively coupled to the sensor hub 1640. In at least one embodiment, a thermal sensor 1639, a fan 1637, a keyboard 1646, and a touch pad 1630 may be communicatively coupled to the EC 1635. In at least one embodiment, a speaker 1663, headphones 1664, and a microphone ("mic") 1665 may be communicatively coupled to the audio unit (audio codec and class D amplifier) 1664, and this audio unit may be communicatively coupled to the DSP 1660. In at least one embodiment, the audio unit 1664 may include, without limitation, for example, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 1657 may be communicatively coupled to the WWAN unit 1656. In at least one embodiment, components such as the WLAN unit 1650 and the Bluetooth unit 1652, as well as the WWAN 1656, may be implemented in a next generation form factor ("NGFF": Next Generation Form Factor).
[0210] FIG. 17 shows a computer system 1700 according to at least one embodiment. In at least one embodiment, the computer system 1700 is configured to implement various processes and methods described throughout this disclosure.
[0211] In at least one embodiment, computer system 1700 includes, without limitation, at least one central processing unit ("CPU") 1702, which is connected to a communication bus 1710 implemented using any suitable protocol, such as, without limitation, PCI: Peripheral Component Interconnect ("peripheral component interconnect"), Peripheral Component Interconnect Express ("PCI-Express": peripheral component interconnect express), AGP: Accelerated Graphics Port ("accelerated graphics port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1700 includes, without limitation, main memory 1704 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1704, which may take the form of random access memory ("RAM": random access memory). In at least one embodiment, a network interface subsystem ("network interface") 1722 provides an interface to other computing devices and networks for receiving data from other systems and transmitting data from computer system 1700 to other systems.
[0212] In at least one embodiment, the computer system 1700 includes, without limitation in at least one embodiment, an input device 1708, a parallel processing system 1712, and a display device 1706, and this display device can be implemented using a conventional cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from an input device 1708 such as a keyboard, a mouse, a touch pad, a microphone, etc. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.
[0213] FIG. 18 shows a computer system 1800 according to at least one embodiment. In at least one embodiment, the computer system 1800 may include, without limitation, a computer 1810 and a USB stick 1820. In at least one embodiment, the computer system 1810 may include, without limitation, any number and type of processors (not shown), as well as memory. In at least one embodiment, the computer 1810 includes, without limitation, servers, cloud instances, laptops, and desktop computers.
[0214] In at least one embodiment, the USB stick 1820 includes, without limitation, a processing unit 1830, a USB interface 1840, and USB interface logic 1850. In at least one embodiment, the processing unit 1830 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1830 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1830 comprises an application specific integrated circuit ("ASIC") optimized to execute any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1830 is a tensor processing unit ("TPC") optimized to execute inference operations of machine learning. In at least one embodiment, the processing core 1830 is a vision processing unit ("VPU") optimized to execute inference operations of machine vision and machine learning.
[0215] In at least one embodiment, the USB interface 1840 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1840 is a USB3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1840 is a USB3.0 Type-A connector. In at least one embodiment, the USB interface logic 1850 may include any amount and type of logic that enables the processing unit 1830 to interface with a device (e.g., computer 1810) via the USB connector 1840.
[0216] FIG. 19A shows an exemplary architecture in which a plurality of GPUs 1910-1913 are communicatively coupled to a plurality of multi-core processors 1905-1906 via high-speed links 1940-1943 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1940-1943 support a communication throughput of 4 GB / sec, 30 GB / sec, 80 GB / sec, or more. Various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0, and NVLink 2.0.
[0217] Further, in one embodiment, two or more of the GPUs 1910-1913 are interconnected via high-speed links 1929-1930, which may be implemented using the same or different protocols / links as those used for the high-speed links 1940-1943. Similarly, two or more of the multi-core processors 1905-1906 may be connected via a high-speed link 1928, which can be a symmetric multi-processor (SMP) bus operating at 20 GB / sec, 30 GB / sec, 120 GB / sec, or more. Alternatively, all communication between the various system components shown in FIG. 19A may be realized using the same protocol / link (e.g., via a common interconnect fabric).
[0218] In one embodiment, each of the multi-core processors 1905-1906 is communicatively coupled to the processor memories 1901-1902 via the memory interconnects 1926-1927 respectively, and each of the GPUs 1910-1913 is communicatively coupled to the GPU memories 1920-1923 via the GPU memory interconnects 1950-1953 respectively. The memory interconnects 1926-1927 and 1950-1953 may utilize the same or different memory access technologies. By way of example, and not limitation, the processor memories 1901-1902 and the GPU memories 1920-1923 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the processor memories 1901-1902 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).
[0219] As described herein, the various processors 1905-1906 and GPUs 1910-1913 may each be physically coupled to a specific memory 1901-1902, 1920-1923, but an integrated memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1901-1902 may each have a 64 GB system memory address space, and each of the GPU memories 1920-1923 may each have a 32 GB system memory address space (in this example, resulting in a total of 256 GB of addressable memory).
[0220] FIG. 19B shows further details of the interconnection between a multi-core processor 1907 and a graphics acceleration module 1946 according to one exemplary embodiment. The graphics acceleration module 1946 may include one or more GPU chips integrated on a line card coupled to the processor 1907 via a high-speed link 1940. Alternatively, the graphics acceleration module 1946 may be integrated in the same package or chip as the processor 1907.
[0221] In at least one embodiment, the illustrated processor 1907 includes a plurality of cores 1960A - 1960D, each core having a translation lookaside buffer 1961A - 1961D and one or more caches 1962A - 1962D. In at least one embodiment, the cores 1960A - 1960D may include various other components (not shown) for executing instructions and processing data. The caches 1962A - 1962D may comprise level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 1956 may be included in the caches 1962A - 1962D and shared by a set of cores 1960A - 1960D. For example, one embodiment of the processor 1907 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. The processor 1907 and the graphics acceleration module 1946 are connected to a system memory 1914, which may include the processor memories 1901 - 1902 of FIG. 19A.
[0222] For the data and instructions stored in the various caches 1962A - 1962D, 1956, and the system memory 1914, coherence is maintained through inter - core communication via the coherence bus 1964. For example, each cache may have associated cache coherence logic / circuitry to communicate via the coherence bus 1964 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via the coherence bus 1964 to monitor cache accesses.
[0223] In one embodiment, a proxy circuit 1925 communicatively couples the graphics acceleration module 1946 to the coherence bus 1964 so that the graphics acceleration module 1946 can participate in the cache coherence protocol as a peer of cores 1960A - 1960D. Specifically, interface 1935 provides a connection to the proxy circuit 1925 via a high - speed link 1940 (e.g., PCIe bus, NVLink, etc.), and interface 1937 connects the graphics acceleration module 1946 to the link 1940.
[0224] In one implementation form, the accelerator integration circuit 1936 provides services for cache management, memory access, content management, and interrupt management instead of the plurality of graphics processing engines 1931, 1932, N of the graphics acceleration module 1946. Each of the graphics processing engines 1931, 1932, N may include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1931, 1932, N may include different types of graphics processing engines such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine within a GPU. In at least one embodiment, the graphics acceleration module 1946 may be a GPU having a plurality of graphics processing engines 1931 to 1932, N, or the graphics processing engines 1931 to 1932, N may be individual GPUs integrated in a common package, line card, or chip.
[0225] In one embodiment, the accelerator integration circuit 1936 includes a memory management unit (MMU) 1939 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1914. The MMU 1939 can also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, the cache 1938 stores commands and data so that the graphics processing engines 1931-1932 can efficiently access N. In one embodiment, the data stored in the cache 1938 and the graphics memories 1933-1934, M are kept coherent with the core caches 1962A-1962D, 1956, and the system memory 1914. As described above, this may be achieved via the proxy circuit 1925 instead of the cache 1938 and the memories 1933-1934, M (e.g., sending updates regarding cache line modifications / accesses in the processor caches 1962A-1962D, 1956 to the cache 1938 and receiving updates from the cache 1938).
[0226] The set of registers 1945 stores context data for the threads executed by the graphics processing engines 1931-1932, N, and the context management circuit 1948 manages thread contexts. For example, the context management circuit 1948 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1948 may store the current register values in a specified area of memory (identified, for example, by a context pointer). Then, when returning to the context, the context management circuit 1948 may restore the register values. In one embodiment, the interrupt management circuit 1947 receives and processes interrupts received from system devices.
[0227] In one implementation, the virtual / effective addresses from the graphics processing engine 1931 are translated by the MMU 1939 into real / physical addresses of the system memory 1914. One embodiment of the accelerator integration circuit 1936 supports a plurality (e.g., 4, 8, 16) of graphics accelerator modules 1946 and / or other accelerator devices. The graphics accelerator module 1946 may be dedicated to a single application executed on the processor 1907 or may be shared among multiple applications. In one embodiment, there is a virtualized graphics execution environment in which the resources of the graphics processing engines 1931-1932, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0228] In at least one embodiment, the accelerator integration circuit 1936 functions as a bridge to the system for the graphics acceleration module 1946 and provides address translation and system memory cache services. Further, the accelerator integration circuit 1936 may provide a virtualization facility for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1931-1932.
[0229] Since the hardware resources of the graphics processing engines 1931-1932, N are explicitly mapped to the physical address space seen by the host processor 1907, any host processor can directly address these resources using the effective address value. In one embodiment, one function of the accelerator integration circuit 1936 is to physically separate the graphics processing engines 1931-1932, N so that they appear as independent units to the system.
[0230] In at least one embodiment, one or more graphics memories 1933-1934, M are each coupled to a respective one of the graphics processing engines 1931-1932, N. The graphics memories 1933-1934, M store instructions and data processed by their respective graphics processing engines 1931-1932, N. The graphics memories 1933-1934, M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.
[0231] In one embodiment, in order to reduce data traffic through link 1940, data stored in graphics memories 1933-1934, M is made to be the data most frequently used by graphics processing engines 1931-1932, N, and preferably data not used (or at least not frequently used) by cores 1960A-1960D. Similarly, the bias mechanism attempts to keep data required by the cores (and thus preferably not required by graphics processing engines 1931-1932, N) in caches 1962A-1962D, 1956 of the cores and in system memory 1914.
[0232] FIG. 19C shows another exemplary embodiment in which the accelerator integration circuit 1936 is integrated within the processor 1907. At least in this embodiment, the graphics processing engines 1931-1932, N communicate directly with the accelerator integration circuit 1936 via the high-speed link 1940 through interfaces 1937 and 1935 (and any form of bus or interface protocol can be utilized in this case as well). The accelerator integration circuit 1936 may perform the same operations as described with respect to FIG. 19B, but potentially may operate at a higher throughput considering its proximity to the coherence bus 1964 and caches 1962A-1962D, 1956. At least one embodiment supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1936 and a programming model controlled by the graphics acceleration module 1946.
[0233] In at least one embodiment, the graphics processing engines 1931-1932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can concentrate other application requirements on the graphics processing engines 1931-1932, N to realize virtualization within a VM / partition.
[0234] In at least one embodiment, the graphics processing engines 1931-1932, N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 1931-1932, N to enable access by each operating system. In a single partition system without a hypervisor, the graphics processing engines 1931-1932, N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1931-1932, N to provide access to each process or application.
[0235] In at least one embodiment, the graphics acceleration module 1946 or individual graphics processing engines 1931-1932, N use a process handle to select process elements. In at least one embodiment, the process elements are stored in the system memory 1914 and can be addressed using the effective address to physical address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1931-1932, N (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.
[0236] FIG. 19D shows an exemplary accelerator integration slice 1990. As used herein, a "slice" comprises a designated portion of the processing resources of accelerator integration circuit 1936. The application effective address space 1982 within system memory 1914 stores a process element 1983. In one embodiment, the process element 1983 is stored in response to a GPU call 1981 from an application 1980 executing on processor 1907. The process element 1983 accommodates the process state of the corresponding application 1980. The work descriptor (WD) 1984 accommodated by the process element 1983 can be a single job requested by the application, or can accommodate a pointer to a queue of jobs. In at least one embodiment, the WD 1984 is a pointer to a job request queue in the application's address space 1982.
[0237] The graphics acceleration module 1946 and / or the individual graphics processing engines 1931-1932, N can be shared by all or a subset of the processes within the system. In at least one embodiment, infrastructure for setting the process state and sending the WD 1984 to the graphics acceleration module 1946 to initiate a job in a virtualized environment may be included.
[0238] In at least one embodiment, a dedicated process programming model is implementation specific. In this model, a single process owns the graphics acceleration module 1946 or an individual graphics processing engine 1931. Since the graphics acceleration module 1946 is owned by a single process, when the graphics acceleration module 1946 is allocated, the hypervisor initializes the accelerator integration circuit 1936 for the owning partition, and the operating system initializes the accelerator integration circuit 1936 for the owning process.
[0239] During operation, the WD fetch unit 1991 within the accelerator integration slice 1990 fetches the following WD 1984, which includes the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1946. As shown, the data from the WD 1984 is stored in the register 1945 and may be used by the MMU 1939, the interrupt management circuit 1947, and / or the context management circuit 1948. For example, one embodiment of the MMU 1939 includes a segment / page walk circuit for accessing the segment / page table 1986 within the OS virtual address space 1985. The interrupt management circuit 1947 may process the interrupt event 1992 received from the graphics acceleration module 1946. When executing a graphics operation, the effective address 1993 generated by the graphics processing engines 1931 - 1932, N is translated to a physical address by the MMU 1939.
[0240] In one embodiment, the same set of registers 1945 is replicated for each of the graphics processing engines 1931 - 1932, N, and / or the graphics acceleration module 1946 and may be initialized by the hypervisor or the operating system. Each of these replicated registers may be included in the accelerator integration slice 1990. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0241] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0242] In one embodiment, each WD1984 is specific to a particular graphics acceleration module 1946 and / or graphics processing engines 1931 - 1932, N. The WD1984 can contain all the information required for the graphics processing engines 1931 - 1932, N to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0243] FIG. 19E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor real address space 1998 in which a process element list 1999 is stored. The hypervisor real address space 1998 is accessible via a hypervisor 1996 that virtualizes the graphics acceleration module engine of the operating system 1995.
[0244] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1946. There are two programming models in which the graphics acceleration module 1946 is shared by multiple processes and partitions: time slice sharing and graphics-directed shared.
[0245] In this model, the system hypervisor 1996 owns the graphics acceleration module 1946 and makes its functions available to all operating systems 1995. In order for the graphics acceleration module 1946 to support the virtualization by the system hypervisor 1996, the graphics acceleration module 1946 may comply with the following: 1) The job requests of the application must be autonomous (i.e., there is no need to maintain the state between jobs), or the graphics acceleration module 1946 must provide a mechanism for saving and restoring the context. 2) The job requests of the application are guaranteed by the graphics acceleration module 1946 to be completed within the specified amount of time, including any translation errors, or the graphics acceleration module 1946 provides a function to preempt the processing of the job. 3) When the graphics acceleration module 1946 operates in the specified shared programming model, fairness must be guaranteed among processes.
[0246] In at least one embodiment, application 1980 needs to make a system call to operating system 1995, along with the type of graphics acceleration module 1946, a work descriptor (WD), a permission mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1946 describes the acceleration function targeted by the system call. In at least one embodiment, the type of graphics acceleration module 1946 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for the graphics acceleration module 1946 and can be in the form of a command of the graphics acceleration module 1946, a virtual address pointer pointing to a user-defined structure, a virtual address pointer pointing to a command queue, or any other data structure for describing the work performed by the graphics acceleration module 1946. In one embodiment, the AMR value is the AMR state for use by the current process. In at least one embodiment, the value passed to the operating system is the same as that of the application setting the AMR. If the implementation of accelerator integration circuit 1936 and graphics acceleration module 1946 does not support the user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to the hypervisor call. Hypervisor 1996 may optionally apply the current authority mask override register (AMOR) value and then put the AMR into process element 1983. In at least one embodiment, the CSRP is one of the registers 1945 that holds the virtual address of an area within the application's address space 1982 for the graphics acceleration module 1946 to save and restore the context state. This pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0247] Upon receiving a system call, the operating system 1995 may verify that the application 1980 is registered and has been granted the right to use the graphics acceleration module 1946. The operating system 1995 then makes a call to the hypervisor 1996 with the information shown in Table 3.
Table 3
[0248] Upon receiving a hypervisor call, the hypervisor 1996 verifies that the operating system 1995 is registered and has been granted the right to use the graphics acceleration module 1946. The hypervisor 1996 then inserts the process element 1983 into the process element link list of the corresponding type of graphics acceleration module 1946. The process element may include the information shown in Table 4.
Table 4
[0249] In at least one embodiment, the hypervisor initializes the registers 1945 of the plurality of accelerator integration slices 1990.
[0250] As shown in FIG. 19F, in at least one embodiment, an integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1901-1902 and GPU memories 1920-1923. In this implementation, operations executed on GPUs 1910-1913 utilize the same virtual / effective memory address space as accessing the processor memories 1901-1902, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1901, a second portion is allocated to a second processor memory 1902, a third portion is allocated to GPU memory 1920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of the processor memories 1901-1902 and GPU memories 1920-1923 such that any processor or GPU can access any physical memory with virtual addresses mapped to physical memory.
[0251] In one embodiment, bias / coherence management circuits 1994A-1994E in one or more of MMUs 1939A-1939E ensure cache coherence between the caches of one or more host processors (e.g., 1905) and the caches of GPUs 1910-1913 and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. Multiple instances of bias / coherence management circuits 1994A-1994E are shown in FIG. 19F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 1905 and / or within the accelerator integration circuit 1936.
[0252] One embodiment enables the memory with GPU 1920 - 1923 to be mapped as part of the system memory and made accessible using the Shared Virtual Memory (SVM) technique, without incurring a performance degradation related to full system cache coherence. In at least one embodiment, the memory with GPU 1920 - 1923 is accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offloading. With this configuration, the host processor 1905 software can set operands and access calculation results without the overhead of conventional I / O DMA data copying. Such conventional copies require driver calls, interrupts, and memory - mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access the memory with GPU 1920 - 1923 without cache coherence overhead can be essential for the execution time of offloaded computations. For example, in the presence of significant streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 1910 - 1913. In at least one embodiment, the efficiency of operand setting, access to results, and GPU calculations may help in determining the effectiveness of GPU offloading.
[0253] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page granularity structure that includes one or two bits per memory page with a GPU (i.e., may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPUs with memory 1920-1923 in a state where a bias cache (e.g., for caching frequently / most recently used entries of the bias table) is or is not in GPU 1910-1913. Alternatively, the entire bias table may be maintained within the GPU.
[0254] In at least one embodiment, an entry in the bias table associated with each access to GPUs with memory 1920-1923 is accessed prior to the actual access to the GPU memory, causing the following operations. First, a local request from GPU 1910-1913 to find its page within the GPU bias is transferred directly to the corresponding GPU memory 1920-1923. A local request from a GPU to find its page in the host bias is transferred to processor 1905 (e.g., via the high-speed link described above). In one embodiment, a request from processor 1905 to find the requested page in the host processor bias completes the request in the same manner as a normal memory read. Alternatively, a request directed to a GPU-biased page may be transferred to GPU 1910-1913. In at least one embodiment, the GPU may then migrate the page to the host processor bias if the current page is not in use. In at least one embodiment, the bias state of a page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply a hardware-based mechanism.
[0255] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the GPU's device driver, and this device driver sends a message to the GPU (or adds a command descriptor to a queue) to change the bias state, and for some transitions, guides the GPU to perform a cache flushing operation on the host. In at least one embodiment, the cache flushing operation is used for the transition from the host processor 1905's bias to the GPU bias, but not for the reverse transition.
[0256] In one embodiment, cache coherence is maintained by the host processor 1905 temporarily rendering GPU-biased pages that cannot be cached. To access these pages, the processor 1905 may request access from the GPU 1910, and the GPU 1910 may either immediately grant access or not grant access. Thus, it is beneficial to make GPU-biased pages requested by the GPU but not by the host processor 1905, or vice versa, in order to reduce communication between the processor 1905 and the GPU 1910.
[0257] A hardware structure 1315 is used to implement one or more embodiments. Details regarding the hardware structure (x) 1315 are provided herein in conjunction with FIGS. 13A and / or 13B.
[0258] FIG. 20 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general-purpose processor cores.
[0259] FIG. 20 is a block diagram showing an exemplary system-on-chip integrated circuit 2000 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 2000 includes one or more application processors 2005 (e.g., CPUs), at least one graphics processor 2010, and may further include an image processor 2015 and / or a video processor 2020, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 2000 includes peripheral devices or bus logic including a USB controller 2025, a UART controller 2030, an SPI / SDIO controller 2035, and an I²S / I²C controller 2040. In at least one embodiment, the integrated circuit 2000 can include a display device 2045 coupled to one or more of a high-definition multimedia interface (HDMI) controller 2050 and a mobile industry processor interface (MIPI) display interface 2055. In at least one embodiment, storage may be provided by a flash memory subsystem 2060 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 2065 to access SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits further include an embedded security engine 2070. In at least one embodiment, some integrated circuits further include an embedded implementation of a physical layer (PHY) library 116.
[0260] Figures 21A - 21B show exemplary integrated circuits and related graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general - purpose processor cores.
[0261] Figures 21A - 21B are block diagrams showing exemplary graphics processors for use within a SoC according to the embodiments described herein. Figure 21A shows an exemplary graphics processor 2110 of a system - on - chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. Figure 21B shows a further exemplary graphics processor 2140 of a system - on - chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 2110 of Figure 21A is a low - power graphics processor core. In at least one embodiment, the graphics processor 2140 of Figure 21B is a high - performance graphics processor core. In at least one embodiment, each of the graphics processors 2110, 2140 can be a variant of the graphics processor 2010 of Figure 20.
[0262] In at least one embodiment, the graphics processor 2110 includes a vertex processor 2105 and one or more fragment processors 2115A - 2115N (e.g., 2115A, 2115B, 2115C, 2115D - 2115N - 1, and 2115N). In at least one embodiment, the graphics processor 2110 can execute different shader programs via separate logic, whereby the vertex processor 2105 is optimized to execute operations for vertex shader programs, while the one or more fragment processors 2115A - 2115N execute fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 2105 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 2115A - 2115N use the primitives and vertex data generated by the vertex processor 2105 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 2115A - 2115N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.
[0263] In at least one embodiment, the graphics processor 2110 further includes one or more memory management units (MMUs) 2120A-2120B, caches 2125A-2125B, and circuit interconnects 2130A-2130B. In at least one embodiment, the one or more MMUs 2120A-2120B provide a virtual-to-physical address mapping for the graphics processor 2110, including the vertex processor 2105 and / or the fragment processors 2115A-2115N, and they may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in the one or more caches 2125A-2125B. In at least one embodiment, the one or more MMUs 2120A-2120B may be synchronized with one or more other MMUs in the system, including one or more MMUs associated with one or more of the application processors 2005, image processors 2015, and / or video processors 2020 of FIG. 20, such that each processor 2005-2020 can participate in a shared or integrated virtual memory system. In at least one embodiment, the one or more circuit interconnects 2130A-2130B enable the graphics processor 2110 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0264] In at least one embodiment, the graphics processor 2140 includes one or more MMUs 2120A-2120B, caches 2125A-2125B, and circuit interconnects 2130A-2130B of the graphics processor 2110 of FIG. 21A. In at least one embodiment, the graphics processor 2140 includes one or more shader cores 2155A-2155N (e.g., 2155A, 2155B, 2155C, 2155D, 2155E, 2155F-2155N-1, and 2155N), which provide an integrated shader core architecture where a single core, or type, or core can execute all types of programmable shader code including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 2140 includes an inter-core task manager 2145 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 2155A-2155N, and a tiling unit 2158 for accelerating tiling operations for tile-based rendering where the rendering operation of a scene is subdivided in the image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches.
[0265] FIGS. 22A-22B show further exemplary graphics processor logic according to embodiments described herein. FIG. 22A shows a graphics core 2200, which in at least one embodiment may be included in the graphics processor 2010 of FIG. 20 and which in at least one embodiment may be the integrated shader cores 2155A-2155N as shown in FIG. 21B. FIG. 22B shows a highly parallel general purpose graphics processing unit 2230 suitable for introduction into a multi-chip module in at least one embodiment.
[0266] In at least one embodiment, the graphics core 2200 includes a shared instruction cache 2202, a texture unit 2218, and a cache / shared memory 2220, which are common to the execution resources within the graphics core 2200. In at least one embodiment, the graphics core 2200 can include a plurality of slices 2201A - 2201N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 2200. The slices 2201A - 2201N can include support logic that includes local instruction caches 2204A - 2204N, thread schedulers 2206A - 2206N, thread dispatchers 2208A - 2208N, and sets of registers 2210A - 2210N. In at least one embodiment, the slices 2201A - 2201N can include a set of additional functional units (AFUs 2212A - 2212N), floating-point units (FPUs 2214A - 2214N), integer arithmetic logic units (ALUs 2216 - 2216N), address calculation units (ACUs 2213A - 2213N), double-precision floating-point units (DPFPUs 2215A - 2215N), and matrix processing units (MPUs 2217A - 2217N).
[0267] In at least one embodiment, FPU2214A to 2214N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and DPFPU2215A to 2215N perform double-precision (64-bit) floating-point operations. In at least one embodiment, ALU2216A to 2216N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured to perform mixed-precision operations. In at least one embodiment, MPU2217A to 2217N can also be configured to perform mixed-precision matrix operations including half-precision floating-point and 8-bit integer operations. In at least one embodiment, MPU2217A to 2217N can perform various matrix operations for accelerating machine learning application frameworks, including enabling support for accelerating general matrix-matrix multiplication (GEMM). In at least one embodiment, AFU2212A to 2212N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric operations (e.g., sine, cosine, etc.).
[0268] FIG. 22B shows a general-purpose processing unit (GPGPU) 2230 that can be configured to execute high-parallel computing operations by an array of graphics processing units in at least one embodiment. In at least one embodiment, GPGPU 2230 can be directly linked to other instances of GPGPU 2230 to generate a plurality of GPU clusters to improve the training speed of a deep neural network. In at least one embodiment, GPGPU 2230 includes a host interface 2232 to enable connection to a host processor. In at least one embodiment, host interface 2232 is a PCI Express interface. In at least one embodiment, host interface 2232 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 2230 receives commands from a host processor and uses a global scheduler 2234 to distribute execution threads associated with these commands to a set of compute clusters 2236A-2236H. In at least one embodiment, compute clusters 2236A-2236H share a cache memory 2238. In at least one embodiment, cache memory 2238 can act as a high-level cache for cache memory within compute clusters 2236A-2236H.
[0269] In at least one embodiment, GPGPU 2230 includes memories 2244A-2244B coupled to compute clusters 2236A-2236H via a set of memory controllers 2242A-2242B. In at least one embodiment, memories 2244A-2244B can include various types of memory devices including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory.
[0270] In at least one embodiment, each of the compute clusters 2236A-2236H includes a set of graphics cores, such as the graphics core 2200 of FIG. 22A, and this set of graphics cores can include multiple types of integer and floating-point logical units that can perform computational operations with various precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each of the compute clusters 2236A-2236H can be configured to perform 16-bit or 32-bit floating-point operations, while another subset of the floating-point units can be configured to perform 64-bit floating-point operations.
[0271] In at least one embodiment, multiple instances of GPGPU2230 can be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 2236A - 2236H for synchronization and data exchange varies across embodiments. In at least one embodiment, multiple instances of GPGPU2230 communicate via host interface 2232. In at least one embodiment, GPGPU2230 includes I / O hub 2239, which couples GPGPU2230 to GPU link 2240 that enables direct connection to other instances of GPGPU2230. In at least one embodiment, GPU link 2240 is coupled to a dedicated GPU - to - GPU bridge that enables communication and synchronization between multiple instances of GPGPU2230. In at least one embodiment, GPU link 2240 is coupled to a high - speed interconnect for sending and receiving data to / from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU2230 are located in separate data processing systems and communicate via a network device accessible via host interface 2232. In at least one embodiment, GPU link 2240 can be configured to enable connection to a host processor in addition to, or instead of, host interface 2232.
[0272] In at least one embodiment, the GPGPU 2230 can be configured to train a neural network. In at least one embodiment, the GPGPU 2230 can be used within an inference platform. In at least one embodiment where the GPGPU 2230 is used for inference, the GPGPU may include fewer compute clusters 2236A - 2236H than when the GPGPU is used for neural network training. In at least one embodiment, the memory technology associated with memories 2244A - 2244B may be different for the inference configuration and the training configuration, and high bandwidth memory technology may be applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 2230 can support inference-specific instructions. For example, in at least one embodiment, the inference configuration can support one or more dot product instructions for 8-bit integers, which may be used during the inference operation of a pre-trained neural network. In at least one embodiment, the inference configuration of the GPGPU 2230 can support the execution of software operations implemented by the software physical layer (PHY) library 116.
[0273] FIG. 23 is a block diagram showing a computing system 2300 according to at least one embodiment. In at least one embodiment, the computing system 2300 includes a processing subsystem 2301 having one or more processors 2302 and a system memory 2304 that communicate via an interconnect path that may include a memory hub 2305. In at least one embodiment, the memory hub 2305 may be a separate component within a chipset component or may be integrated within one or more processors 2302. In at least one embodiment, the memory hub 2305 is coupled to an I / O subsystem 2311 via a communication link 2306. In at least one embodiment, the I / O subsystem 2311 includes an I / O hub 2307 that enables the computing system 2300 to receive input from one or more input devices 2308. In at least one embodiment, the I / O hub 2307 can enable a display controller, which may be included in one or more processors 2302, to provide output to one or more display devices 2310A. In at least one embodiment, one or more display devices 2310A coupled to the I / O hub 2307 can include local, internal, or embedded display devices.
[0274] In at least one embodiment, the processing subsystem 2301 includes one or more parallel processors 2312 coupled to the memory hub 2305 via a bus or other communication link 2313. In at least one embodiment, the communication link 2313 may be one of a number of communication link technologies or protocols based on any number of standards, such as but not limited to PCI Express, or may be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 2312 form a parallel or vector processing system focused on computing that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 2312 form a graphics processing subsystem that can output pixels to one of one or more display devices 2310A coupled via the I / O hub 2307. In at least one embodiment, the one or more parallel processors 2312 may also include a display controller and display interface (not shown) that enables direct connection to one or more display devices 2310B.
[0275] In at least one embodiment, the system storage unit 2314 can be connected to the I / O hub 2307 to provide a storage mechanism for the computing system 2300. In at least one embodiment, an I / O switch 2316 can be used to provide an interface mechanism for enabling communication between the I / O hub 2307 and other components such as a network adapter 2318 and / or a wireless network adapter 2319 that may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 2320. In at least one embodiment, the network adapter 2318 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2319 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0276] In at least one embodiment, the computing system 2300 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 2307. In at least one embodiment, the communication paths interconnecting the various components of FIG. 23 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces and / or protocols such as NV-Link high-speed interconnect or interconnect protocol.
[0277] In at least one embodiment, one or more parallel processors 2312 incorporate circuitry optimized for graphics and video processing, such as a video output circuit, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2312 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 2300 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2312, memory hub 2305, processor 2302, and I / O hub 2307 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 2300 may be integrated in a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 2300 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.
[0278] Processor FIG. 24A shows a parallel processor 2400 according to at least one embodiment. In at least one embodiment, the various components of parallel processor 2400 may be implemented using one or more integrated circuit devices such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 2400 is a variant of one or more parallel processors 2312 shown in FIG. 23 according to an exemplary embodiment.
[0279] In at least one embodiment, the parallel processor 2400 includes a parallel processing unit 2402. In at least one embodiment, the parallel processing unit 2402 includes an I / O unit 2404 that enables communication with other devices including other instances of the parallel processing unit 2402. In at least one embodiment, the I / O unit 2404 may be directly connected to other devices. In at least one embodiment, the I / O unit 2404 is connected to other devices via the use of a hub or switch interface such as the memory hub 2305. In at least one embodiment, the connection between the memory hub 2305 and the I / O unit 2404 forms a communication link 2313. In at least one embodiment, the I / O unit 2404 is connected to a host interface 2406 and a memory crossbar 2416, where the host interface 2406 receives commands targeted for execution of processing operations and the memory crossbar 2416 receives commands targeted for execution of memory operations.
[0280] In at least one embodiment, when the host interface 2406 receives a command buffer via the I / O unit 2404, the host interface 2406 can direct the work operations for executing these commands towards the front end 2408. In at least one embodiment, the front end 2408 is coupled to a scheduler 2410, and this scheduler is configured to distribute commands or other work items to the processing cluster array 2412. In at least one embodiment, the scheduler 2410 ensures that the processing cluster array 2412 is properly configured and in an effective state before tasks are distributed to the processing cluster array 2412 of the processing cluster array 2412. In at least one embodiment, the scheduler 2410 is implemented via firmware logic running on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 2410 can be configured to execute complex scheduling and work distribution operations at both a coarse granularity and a fine granularity, enabling rapid preemption of threads and context switching of the threads running on the processing array 2412. In at least one embodiment, the host software can prove the scheduling workload on the processing array 2412 via one of a plurality of graphics processing doorbells. In at least one embodiment, then, the workload can be automatically distributed across the entire processing array 2412 by the scheduler 2410 logic within the microcontroller including the scheduler 2410.
[0281] In at least one embodiment, the processing cluster array 2412 can include up to "N" processing clusters (e.g., cluster 2414A, cluster 2414B - cluster 2414N). In at least one embodiment, each of the clusters 2414A - 2414N of the processing cluster array 2412 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 2410 can use various scheduling and / or workload distribution algorithms to distribute work to the clusters 2414A - 2414N of the processing cluster array 2412, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 2410, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 2412. In at least one embodiment, different clusters 2414A - 2414N of the processing cluster array 2412 can be allocated to process different types of programs or execute different types of calculations.
[0282] In at least one embodiment, the processing cluster array 2412 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2412 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 2412 can include logic for performing processing tasks including filtering of video and / or audio data, execution of modeling operations including physical operations, and execution of data conversion.
[0283] In at least one embodiment, the processing cluster array 2412 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2412 can include texture sampling logic for performing texture operations, as well as additional logic, including but not limited to, mosaic logic and other vertex processing logic, to support the execution of such graphics processing operations. In at least one embodiment, the processing cluster array 2412 can be configured to execute graphics processing related shader programs such as, but not limited to, vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2402 can transfer data from the system memory through the I / O unit 2404 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2422) during processing and then written back to the system memory.
[0284] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2402, the scheduler 2410 can be configured to divide the processing workload into tasks of approximately equal size so as to more effectively distribute the graphics processing operations among the plurality of clusters 2414A - 2414N of the processing cluster array 2412. In at least one embodiment, a portion of the processing cluster array 2412 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2414A - 2414N can be stored in a buffer so that the intermediate data can be transmitted among the clusters 2414A - 2414N for further processing.
[0285] In at least one embodiment, the processing cluster array 2412 can receive processing tasks to be executed via a scheduler 2410, and the scheduler 2410 receives commands defining the processing tasks from a front end 2408. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands (e.g., which program to execute) defining how the data is to be processed. In at least one embodiment, the scheduler 2410 may be configured to fetch the index corresponding to the task, or may receive the index from the front end 2408. In at least one embodiment, the front end 2408 can be configured to ensure that the processing cluster array 2412 is configured in an active state before the workload specified by an incoming command buffer (e.g., batch buffer, push buffer, etc.) is started.
[0286] In at least one embodiment, each of one or more instances of the parallel processing unit 2402 can be coupled to a parallel processor - memory 2422. In at least one embodiment, the parallel processor - memory 2422 can be accessed via a memory crossbar 2416, and the memory crossbar 2416 can receive memory requests from the processing cluster array 2412 as well as the I / O unit 2404. In at least one embodiment, the memory crossbar 2416 can access the parallel processor - memory 2422 via a memory interface 2418. In at least one embodiment, the memory interface 2418 can include a plurality of partition units (e.g., partition unit 2420A, partition unit 2420B - partition unit 2420N), and each of these units can be coupled to a portion (e.g., a memory unit) of the parallel processor - memory 2422. In at least one embodiment, the number of partition units 2420A - 2420N is configured to be equal to the number of memory units, such that the first partition unit 2420A has a corresponding first memory unit 2424A, the second partition unit 2420B has a corresponding memory unit 2424B, and the Nth partition unit 2420N has a corresponding Nth memory unit 2424N. In at least one embodiment, the number of partition units 2420A - 2420N may not be equal to the number of memory devices.
[0287] In at least one embodiment, the memory units 2424A - 2424N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 2424A - 2424N may also include, but are not limited to, 3D stacked memory including high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 2422, a render target such as a frame buffer or texture map can be stored across the memory units 2424A - 2424N so that the partition units 2420A - 2420N can write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 2422 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.
[0288] In at least one embodiment, any one of clusters 2414A - 2414N of the processing cluster array 2412 can process data that is to be written to any one of memory units 2424A - 2424N within the parallel processor memory 2422. In at least one embodiment, the memory crossbar 2416 can be configured to transfer the output of each of clusters 2414A - 2414N to any partition unit 2420A - 2420N that can perform further processing operations on the output, or to another cluster 2414A - 2414N. In at least one embodiment, each of clusters 2414A - 2414N can communicate with the memory interface 2418 through the memory crossbar 2416 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2416 has a connection to the memory interface 2418 for communicating with the I / O unit 2404, as well as a connection to a local instance of the parallel processor memory 2422, enabling processing units within different processing clusters 2414A - 2414N to communicate with the system memory or other memory not local to the parallel processing unit 2402. In at least one embodiment, the memory crossbar 2416 can use virtual channels to separate traffic streams between clusters 2414A - 2414N and partition units 2420A - 2420N.
[0289] In at least one embodiment, multiple instances of the parallel processing unit 2402 may be provided on a single add-in card or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2402 may be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2402 can include a higher precision floating point unit compared to other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2402 or the parallel processor 2400 can be implemented in a variety of configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0290] FIG. 24B is a block diagram of a partition unit 2420 according to at least one embodiment. In at least one embodiment, the partition unit 2420 is an instance of one of the partition units 2420A - 2420N of FIG. 24A. In at least one embodiment, the partition unit 2420 includes an L2 cache 2421, a frame buffer interface 2425, and a ROP: raster operations unit 2426. The L2 cache 2421 is a read / write cache configured to perform load and store operations received from the memory crossbar 2416 and the ROP 2426. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 2421 to the frame buffer interface 2425 to be processed. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 2425 to be processed. In at least one embodiment, the frame buffer interface 2425 interfaces with one of the memory units of the parallel processor memory, such as the memory units 2424A - 2424N (e.g., within the parallel processor memory 2422 of FIG. 24).
[0291] In at least one embodiment, ROP2426 is a processing unit that performs raster operations such as stencil, z-test, and blending. In at least one embodiment, ROP2426 then outputs processed graphics data stored in the graphics memory. In at least one embodiment, ROP2426 includes compression logic for compressing depth or color data written to the memory and decompressing depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. In at least one embodiment, the type of compression performed by ROP2426 can be changed based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0292] In at least one embodiment, ROP2426 is included within each processing cluster (e.g., clusters 2414A - 2414N of FIG. 24) rather than within the partitioning unit 2420. In at least one embodiment, read and write requests for pixel data rather than pixel fragment data are transmitted via the memory crossbar 2416. In at least one embodiment, the processed graphics data may be displayed on a display device such as one of the one or more display devices 2310 of FIG. 23, routed to be further processed by the processor 2302, or routed to be further processed by one of the processing entities within the parallel processor 2400 of FIG. 24A.
[0293] FIG. 24C is a block diagram of a processing cluster 2414 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 2414A-2414N of FIG. 24. In at least one embodiment, the processing cluster 2414 may be configured to execute a number of threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, a single instruction multiple data (SIMD) instruction issue technique is used to support parallel execution of a number of threads without providing a plurality of independent instruction units. In at least one embodiment, a single instruction multiple thread (SIMT) technique is used to support parallel execution of a number of threads that are overall synchronized, using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0294] In at least one embodiment, the operation of processing cluster 2414 can be controlled via pipeline manager 2432 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, pipeline manager 2432 receives instructions from scheduler 2410 of FIG. 24 and manages the execution of these instructions via graphics multiprocessor 2434 and / or texture unit 2436. In at least one embodiment, graphics multiprocessor 2434 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within processing cluster 2414. In at least one embodiment, one or more instances of graphics multiprocessor 2434 can be included within processing cluster 2414. In at least one embodiment, graphics multiprocessor 2434 can process data, and data crossbar 2440 may be used to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, pipeline manager 2432 can facilitate the distribution of processed data by specifying the destination of the processed data that will be distributed through data crossbar 2440.
[0295] In at least one embodiment, each graphics multiprocessor 2434 within processing cluster 2414 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction has completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. In at least one embodiment, different operations can be executed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0296] In at least one embodiment, the instructions sent to processing cluster 2414 configure threads. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, a thread group executes a program on different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 2434. In at least one embodiment, a thread group may include fewer threads than the number of processing engines within graphics multiprocessor 2434. In at least one embodiment, if a thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during cycles in which the thread group is being processed. In at least one embodiment, a thread group may also include more threads than the number of processing engines within graphics multiprocessor 2434. In at least one embodiment, if a thread group includes more threads than the number of processing engines within graphics multiprocessor 2434, processing can be executed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 2434.
[0297] In at least one embodiment, the graphics multi-processor 2434 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multi-processor 2434 can forego the internal cache and use the cache memory (e.g., L1 cache 2448) within the processing cluster 2414. In at least one embodiment, each graphics multi-processor 2434 can also access the L2 cache within a partition unit (e.g., partition units 2420A - 2420N of FIG. 24), and these caches can be shared among all processing clusters 2414 and may be used to transfer data between threads. In at least one embodiment, the graphics multi-processor 2434 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 2402 may be used as global memory. In at least one embodiment, the processing cluster 2414 includes multiple instances of the graphics multi-processor 2434 that can share common instructions and data, and these may be stored in the L1 cache 2448.
[0298] In at least one embodiment, each processing cluster 2414 may include an MMU 2445 (Memory Management Unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2445 may be within the memory interface 2418 of FIG. 24. In at least one embodiment, the MMU 2445 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles (tiling is described in detail) and optionally cache line indices. In at least one embodiment, the MMU 2445 may include a translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 2434 or L1 cache, or within the processing cluster 2414. In at least one embodiment, the physical address is processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0299] In at least one embodiment, each graphics multiprocessor 2434 is coupled to a texture unit 2436 such that the processing cluster 2414 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multiprocessor 2434 and, if necessary, fetched from the L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 2434 outputs the processed tasks to the data crossbar 2440 to provide the processed tasks to another processing cluster 2414 for further processing, or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 2416. In at least one embodiment, the pre-ROP 2442 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 2434 and direct the data to the ROP unit, which may be located within a partitioning unit (e.g., partitioning units 2420A - 2420N of FIG. 24) as described herein. In at least one embodiment, the pre-ROP 2442 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.
[0300] Figure 24D shows a graphics multi-processor 2434 according to at least one embodiment. In at least one embodiment, the graphics multi-processor 2434 is coupled to a pipeline manager 2432 of a processing cluster 2414. In at least one embodiment, the graphics multi-processor 2434 has an execution pipeline including, but not limited to, an instruction cache 2452, an instruction unit 2454, an address mapping unit 2456, a register file 2458, one or more general-purpose graphics processing unit (GPGPU) cores 2462, and one or more load / store units 2466. The GPGPU cores 2462 and the load / store units 2466 are coupled to a cache memory 2472 and a shared memory 2470 via a memory and cache interconnect 2468.
[0301] In at least one embodiment, the instruction cache 2452 receives a stream of instructions to be executed from the pipeline manager 2432. In at least one embodiment, the instructions are cached in the instruction cache 2452 and dispatched for execution by the instruction unit 2454. In at least one embodiment, the instruction unit 2454 can dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within the GPGPU core 2462. In at least one embodiment, instructions can access any of a local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, the address mapping unit 2456 can be used to translate an address in the unified address space to an individual memory address accessible by the load / store unit 2466.
[0302] In at least one embodiment, register file 2458 provides a set of registers to the functional units of graphics multiprocessor 2434. In at least one embodiment, register file 2458 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU cores 2462, load / store unit 2466) of graphics multiprocessor 2434. In at least one embodiment, register file 2458 is divided among the respective functional units such that each functional unit is allocated a dedicated portion of register file 2458. In one embodiment, register file 2458 is divided among different warps being executed by graphics multiprocessor 2434.
[0303] In at least one embodiment, each GPGPU core 2462 can include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of graphics multiprocessor 2434. The GPGPU cores 2462 may have the same architecture or different architectures. In at least one embodiment, a first portion of GPGPU core 2462 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU can implement the IEEE 754-2008 standard for floating point operations or enable variable precision floating point operations. In at least one embodiment, graphics multiprocessor 2434 can further include one or more fixed function units or special function units for performing specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores can also include fixed or special function logic.
[0304] In at least one embodiment, the GPGPU core 2462 includes SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 2462 can physically execute SIMD4, SIMD8, and SIMD16 instructions and can logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core may be generated at compile time by a shader compiler or may be automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logical unit.
[0305] In at least one embodiment, the memory and cache interconnect 2468 is an interconnect network that connects each functional unit of the graphics multiprocessor 2434 to the register file 2458 and the shared memory 2470. In at least one embodiment, the memory and cache interconnect 2468 is a crossbar interconnect that enables the load / store unit 2466 to implement load and store operations between the shared memory 2470 and the register file 2458. In at least one embodiment, the register file 2458 can operate at the same frequency as the GPGPU core 2462, and thus the data transfer between the GPGPU core 2462 and the register file 2458 is very low latency. In at least one embodiment, the shared memory 2470 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 2434. In at least one embodiment, the cache memory 2472 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2436. In at least one embodiment, the shared memory 2470 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 2462 can programmatically store data in the shared memory in addition to automatically cached data stored in the cache memory 2472.
[0306] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into the same package or die as the core and communicatively coupled to the core via an internal (i.e., inside the package or die) processor bus / interconnect. In at least one embodiment, regardless of the method of connecting the GPU, the processor core may distribute work to the GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions. In at least one embodiment, the GPU uses dedicated circuitry / logic to efficiently process software functions implemented by a software physical layer (PHY) library 116.
[0307] FIG. 25 shows a multi-GPU computing system 2500 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2500 can include a processor 2502 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2506A-D via a host interface switch 2504. In at least one embodiment, the host interface switch 2504 is a PCI Express switch device that couples the processor 2502 to a PCI Express bus, via which the processor 2502 can communicate with the GPGPUs 2506A-D. The GPGPUs 2506A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2516. In at least one embodiment, the GPU-to-GPU link 2516 is connected to each of the GPGPUs 2506A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 2516 enables direct communication between each of the GPGPUs 2506A-D without requiring communication via the host interface bus 2504 to which the processor 2502 is connected. In at least one embodiment, when there is GPU-to-GPU traffic destined for the P2P GPU link 2516, the host interface bus 2504 is kept available for access to system memory or for communication with other instances of the multi-GPU computing system 2500, for example, via one or more network devices. In at least one embodiment, the GPGPUs 2506A-D are connected to the processor 2502 via the host interface switch 2504, and in at least one embodiment, the processor 2502 includes direct support for the P2P GPU link 2516 and can be directly connected to the GPGPUs 2506A-D.
[0308] FIG. 26 is a block diagram of a graphics processor 2600 according to at least one embodiment. In at least one embodiment, the graphics processor 2600 includes a ring interconnect 2602, a pipeline front end 2604, a media engine 2637, and graphics cores 2680A - 2680N. In at least one embodiment, the ring interconnect 2602 couples the graphics processor 2600 to other graphics processors or other processing units including one or more general - purpose processor cores. In at least one embodiment, the graphics processor 2600 is one of a number of processors integrated within a multi - core processing system.
[0309] In at least one embodiment, the graphics processor 2600 receives a batch of commands via the ring interconnect 2602. In at least one embodiment, incoming commands are interpreted by the command streamer 2603 of the pipeline front end 2604. In at least one embodiment, the graphics processor 2600 includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 2680A - 2680N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2603 supplies the commands to the geometry pipeline 2636. In at least one embodiment, for at least some media processing commands, the command streamer 2603 supplies the commands to the video front end 2634, and the video front end 2634 is coupled to the media engine 2637. In at least one embodiment, the media engine 2637 includes a Video Quality Engine (VQE) 2630 for post - processing of video and images, and a multi - format encode / decode (MFX) 2633 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2636 and the media engine 2637 each generate execution threads for the thread execution resources provided by at least one graphics core 2680A.
[0310] In at least one embodiment, the graphics processor 2600 includes a scalable thread execution resource characterized by modular cores 2680A-2680N (which may also be referred to as core slices), each modular core having a plurality of sub-cores 2650A-2650N, 2660A-2660N (which may also be referred to as core sub-slices). In at least one embodiment, the graphics processor 2600 can have any number of graphics cores 2680A-2680N. In at least one embodiment, the graphics processor 2600 includes a graphics core 2680A having at least a first sub-core 2650A and a second sub-core 2660A. In at least one embodiment, the graphics processor 2600 is a low-power processor having a single sub-core (e.g., 2650A). In at least one embodiment, the graphics processor 2600 includes a plurality of graphics cores 2680A-2680N, each including a first set of sub-cores 2650A-2650N and a second set of sub-cores 2660A-2660N. In at least one embodiment, each sub-core of the first set of sub-cores 2650A-2650N includes at least an execution unit 2652A-2652N and a first set of media / texture samplers 2654A-2654N. In at least one embodiment, each sub-core of the second set of sub-cores 2660A-2660N includes at least an execution unit 2662A-2662N and a second set of samplers 2664A-2664N. In at least one embodiment, each sub-core 2650A-2650N, 2660A-2660N shares a set of shared resources 2670A-2670N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0311] FIG. 27 is a block diagram showing the micro-architecture of a processor 2700 that may include a logic circuit for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2700 may execute instructions including x86 instructions, AMR instructions, special instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, the processor 2710 may include registers for storing packed data, such as 64-bit wide MMX (trademark) registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers available in both integer and floating-point formats may operate on packed data elements with single instruction multiple data (``SIMD'') and streaming SIMD extensions (``SSE'') instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as ``SSEx'') technologies may hold operands of such packed data. In at least one embodiment, the processor 2710 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0312] In at least one embodiment, the processor 2700 includes an in-order front end ("front end") 2701 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2701 may include several units. In at least one embodiment, an instruction prefetcher 2726 fetches instructions from memory and supplies the instructions to an instruction decoder 2728, and the instruction decoder decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2728 decodes the received instruction into one or more operations called "microinstructions" or "micro-operations" that the machine can execute (also called "micro-ops" or "uops"). In at least one embodiment, the instruction decoder 2728 parses the instruction into an opcode and corresponding data, as well as a control field, such that these are used by the microarchitecture and the operations according to at least one embodiment may be executed. In at least one embodiment, a trace cache 2730 may assemble the decoded uops into a program-order sequence or trace in a uop queue 2734 so that they can be executed. In at least one embodiment, when the trace cache 2730 encounters a complex instruction, a microcode ROM 2732 provides the uops necessary for the completion of the operation.
[0313] In at least one embodiment, there are instructions that can be converted into a single micro-op, and there are also instructions that require several micro-ops to complete the entire operation. In at least one embodiment, if five or more micro-ops are required to complete an instruction, the instruction decoder 2728 may access the microcode ROM 2732 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2728. In at least one embodiment, if a large number of micro-ops are required to complete an operation, the instruction may be stored in the microcode ROM 2732. In at least one embodiment, the trace cache 2730 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry point programmable logic array ("PLA") to complete one or more instructions from the microcode ROM 2732 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2732 finishes sequencing the micro-ops for the instruction, the front end 2701 of the machine may resume fetching micro-ops from the trace cache 2730.
[0314] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 2703 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth the flow of instructions and change their order, optimizing performance when instructions are scheduled to flow down the pipeline and be executed. The out-of-order execution engine 2703 includes, without limitation, an allocator / register renamer 2740, a memory uop queue 2742, an integer / floating point uop queue 2744, a memory scheduler 2746, a fast scheduler 2702, a slow / general-purpose floating point scheduler ("slow / general-purpose FP scheduler") 2704, and a simple floating point scheduler ("simple FP scheduler") 2706. In at least one embodiment, the fast scheduler 2702, the slow / general-purpose floating point scheduler 2704, and the simple floating point scheduler 2706 are also collectively referred to herein as "uop schedulers 2702, 2704, 2706". In at least one embodiment, the allocator / register renamer 2740 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2740 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2740 also distributes the entry of each uop to one of two uop queues in front of the memory scheduler 2746 and the uop schedulers 2702, 2704, 2706, namely, the memory uop queue 2742 for memory operations and the integer / floating point uop queue 2744 for non-memory operations. In at least one embodiment, the uop schedulers 2702, 2704, 2706 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2702 of at least one embodiment may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2704 and the simple floating-point scheduler 2706 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2702, 2704, 2706 arbitrate dispatch ports to schedule uops for execution.
[0315] In at least one embodiment, the execution block b11 includes, without limitation, an integer register file / bypass network 2708, a floating-point register file / bypass network (referred to herein as the "FP register file / bypass network") 2710, address generation units (referred to herein as "AGUs") 2712 and 2714, high-speed arithmetic logic units (referred to herein as "high-speed ALUs") 2716 and 2718, low-speed arithmetic logic units (referred to herein as "low-speed ALUs") 2720, floating-point ALUs (referred to herein as "FPs") 2722, and floating-point move units (referred to herein as "FP moves") 2724. In at least one embodiment, the integer register file / bypass network 2708 and the floating-point register file / bypass network 2710 are also referred to herein as the "register files 2708, 2710". In at least one embodiment, the AGUs 2712 and 2714, the high-speed ALUs 2716 and 2718, the low-speed ALUs 2720, the floating-point ALUs 2722, and the floating-point move units 2724 are also referred to herein as the "execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724". In at least one embodiment, the execution block b11 may include any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination without limitation.
[0316] In at least one embodiment, register files 2708, 2710 may be disposed between uop schedulers 2702, 2704, 2706 and execution units 2712, 2714, 2716, 2718, 2720, 2722, and 2724. In at least one embodiment, integer register file / bypass network 2708 executes integer operations. In at least one embodiment, floating-point register file / bypass network 2710 executes floating-point operations. In at least one embodiment, each of register files 2708, 2710 may include, without limitation, a bypass network that may bypass or transfer recently completed results not yet written to the register file to new dependent uops. In at least one embodiment, register files 2708, 2710 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2708 may include, without limitation, two separate register files, namely, one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating-point instructions typically have operands with a width of 64 to 128 bits, floating-point register file / bypass network 2710 may include, without limitation, 128-bit wide entries.
[0317] In at least one embodiment, execution units 2712, 2714, 2716, 2718, 2720, 2722, 2724 may execute instructions. In at least one embodiment, register files 2708, 2710 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2700 may include any number and combination of execution units 2712, 2714, 2716, 2718, 2720, 2722, 2724 without limitation. In at least one embodiment, floating-point ALU 2722 and floating-point shift unit 2724 may execute floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, floating-point ALU 2722 includes, without limitation, a 64-bit floating-point divider and may execute division, square root, and other micro-ops. In at least one embodiment, instructions containing floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2716, 2718. In at least one embodiment, fast ALUs 2716, 2718 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, since low-speed ALU 2720 may include, without limitation, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing, most complex integer operations proceed to low-speed ALU 2720. In at least one embodiment, memory load / store operations may be performed by AGUs 2712, 2714. In at least one embodiment, fast ALU 2716, fast ALU 2718, and low-speed ALU 2720 may execute integer operations with 64-bit data operands. In at least one embodiment, fast ALU 2716, fast ALU 2718, and low-speed ALU 2720 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, floating-point ALU 2722 and floating-point shift unit 2724 may be implemented to support a wide range of operands having various bit widths.In at least one embodiment, the floating point ALU 2722 and the floating point shift unit 2724 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0318] In at least one embodiment, the uop schedulers 2702, 2704, 2706 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, since uops may be scheduled and executed speculatively in the processor 2700, the processor 2700 may also include logic to handle memory misses. In at least one embodiment, when a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed through a scheduler with temporarily incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, dependent operations may need to be replayed, and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0319] In at least one embodiment, the term "register" may refer to a storage location of an on-board processor that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be one that can be used from outside the processor (from the perspective of a programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0320] FIG. 28 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2800 includes one or more processors 2802 and one or more graphics processors 2808, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 2802 or processor cores 2807. In at least one embodiment, system 2800 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile device, a portable device, or an embedded device.
[0321] In at least one embodiment, system 2800 may include, or be incorporated in, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 2800 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 2800 may also include, be coupled to, or be integrated within wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2800 is a television or set-top box device having one or more processors 2802 and a graphical interface generated by one or more graphics processors 2808.
[0322] In at least one embodiment, each of the one or more processors 2802 includes one or more processor cores 2807 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2807 is configured to process a particular instruction set 2809. In at least one embodiment, instruction set 2809 may facilitate computing via a complex instruction set computing (CISC), a reduced instruction set computing (RISC), or a very long instruction word (VLIW). In at least one embodiment, the processor cores 2807 may each process different instruction sets 2809, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, the processor cores 2807 may also include other processing devices such as a digital signal processor (DSP).
[0323] In at least one embodiment, processor 2802 includes cache memory 2804. In at least one embodiment, processor 2802 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared among various components of processor 2802. In at least one embodiment, processor 2802 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and this cache may be shared among processor cores 2807 using known cache coherence techniques. In at least one embodiment, further, register file 2806 is included in processor 2802, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and...
Claims
1. A processor comprising one or more circuits that cause one or more Fifth Generation New Radio (5G-NR) operations to be performed in parallel based at least in part on one or more computing resources on which the one or more operations are performed.
2. 2. The processor of claim 1, wherein causing 5G-NR operations to be performed in parallel includes grouping the one or more operations based at least in part on one or more attributes that cause operations of each group to be performed using the one or more computing resources.
3. 3. The processor of claim 2, wherein the one or more attributes indicate one or more 5G-NR cells.
4. 2. The processor of claim 1, wherein causing 5G-NR operations to be performed in parallel includes receiving one or more parameters indicating the one or more computing resources on which the one or more operations are performed.
5. 2. The processor of claim 1, wherein causing 5G-NR operations to be performed in parallel includes configuring the one or more operations to be performed in parallel based at least on one or more parameters indicative of the one or more computing resources.
6. 2. The processor of claim 1 , wherein the one or more computing resources include one or more kernels that perform the one or more operations, each kernel of the one or more kernels performing one or more groups of the one or more computing operations based at least in part on parameters indicative of one or more attributes of the one or more computing operations.
7. 2. The processor of claim 1 , wherein the one or more circuits cause a software library to receive one or more parameters indicating the one or more computing resources on which the one or more operations are to be executed, and to group the one or more operations to be executed in parallel using the one or more computing resources.
8. 2. The processor of claim 1, wherein the one or more operations include one or more physical layer (PHY) operations from one or more devices associated with one or more cells of a 5G-NR network.
9. The processor of claim 1 , wherein the one or more circuits further cause the 5G-NR operations to be performed in parallel by one or more parallel processing units.
10. 1. A method comprising performing one or more Fifth Generation New Radio (5G-NR) operations in parallel based at least in part on one or more computing resources on which the one or more 5G-NR operations are performed.
11. 11. The method of claim 10, further comprising grouping the one or more operations by a 5G-NR physical layer (PHY) library, the 5G-NR PHY library grouping the one or more operations based at least in part on one or more attributes such that each group of operations is performed using the one or more computing resources, wherein the 5G-NR PHY library receives the one or more attributes as a result of one or more function calls to an application programming interface.
12. 11. The method of claim 10, wherein the one or more computing resources include one or more software kernels that perform the one or more operations using one or more parallel processing units.
13. 11. The method of claim 10, wherein a 5G-NR physical layer (PHY) library receives one or more parameters configuring each of the one or more operations as a result of one or more function calls to the 5G-NR PHY library, and stores each of the one or more parameters based at least in part on whether each of the one or more parameters is updated when the one or more operations are performed.
14. 11. The method of claim 10, wherein a 5G-NR physical layer (PHY) library determines which of the one or more computing resources are used to perform the one or more operations in parallel based at least in part on one or more attributes of the one or more operations, the one or more attributes indicating at least a 5G-NR cell.
15. The method of claim 10, wherein the one or more operations correspond to one or more 5G-NR cells, and a 5G-NR physical layer (PHY) library selects the one or more computing resources on which the one or more operations are to be performed based at least in part on the one or more 5G-NR cells.
16. 11. The method of claim 10, further comprising: causing a 5G-NR physical layer (PHY) library to receive one or more parameters at least indicative of the one or more computing resources and configure the one or more operations to be performed by the one or more computing resources based at least in part on the one or more parameters.
17. The method of claim 10 , wherein the one or more computing resources comprise at least a parallel processing unit of a 5G-NR baseband device that performs the one or more computing operations.
18. A system comprising one or more processors that cause one or more Fifth Generation New Radio (5G-NR) operations to be performed in parallel based at least in part on one or more computing resources on which the one or more operations are performed.
19. 20. The system of claim 18, wherein the one or more computing resources comprise at least one parallel processing unit, and the one or more operations are performed in parallel by one or more kernels executed by the at least one parallel processing unit, and the one or more kernels are selected by the software library based at least in part on one or more parameters received by a software library.
20. The system of claim 19, wherein the one or more parameters indicate at least one attribute for each of the one or more operations, the at least one attribute indicating one or more 5G-NR cells that generate information processed by the one or more operations.
21. 20. The system of claim 18, comprising instructions that, when executed by the one or more processors, implement the software library to batch the one or more operations into groups according to one or more parameters received as a result of one or more function calls to a software library, wherein operations of each group are executed in parallel using the one or more computing resources.
22. 20. The system of claim 18, wherein the one or more processors cause the one or more operations to be executed in parallel during one or more execution slots, the one or more execution slots comprising periods during which the one or more computing resources are available to execute the one or more operations.
23. 20. The system of claim 18, further comprising a software library, the software library comprising instructions that, when executed, cause the software library to receive one or more parameters indicating one or more configurations of the one or more operations, and group the one or more operations to be executed in parallel using the one or more computing resources, the software library grouping the one or more operations based at least in part on the one or more configurations, the one or more configurations indicating the one or more computing resources available to perform the one or more operations.
24. 20. The system of claim 18, wherein the one or more computing resources comprise one or more parallel processing units that perform the first group of one or more operations and the second group of one or more operations in parallel.
25. When executed by one or more processors, the one or more processors are caused to at least: A machine-readable medium having stored thereon a set of instructions that causes one or more Fifth Generation New Radio (5G-NR) operations to be performed in parallel based at least in part on one or more computing resources on which the one or more operations are performed.
26. 26. The machine-readable medium of claim 25, further comprising instructions that, when executed by the one or more processors, implement a 5G-NR physical layer (PHY) library that causes the one or more processors to group the one or more operations into one or more groups, each of the one or more groups being executed by one or more software kernels determined by the software library based at least in part on the one or more computing resources.
27. 26. The machine-readable medium of claim 25, further comprising instructions that, when executed by the one or more processors, implement a 5G-NR physical layer (PHY) library that causes the one or more processors to receive one or more parameters that configure the one or more operations to be executed in parallel, the one or more parameters including information indicative of the one or more computing resources on which the one or more operations are to be executed.
28. 26. The machine-readable medium of claim 25, further comprising instructions that, when executed by the one or more processors, implement the 5G-NR physical layer (PHY) library to cause the one or more processors to group the one or more operations based at least in part on one or more attributes of the one or more operations indicated by one or more parameters provided to a 5G-NR PHY library, the one or more attributes being usable by the 5G-NR PHY library to select the one or more computing resources on which the one or more operations are to be performed.
29. 26. The machine-readable medium of claim 25, wherein the one or more computing resources comprise at least one parallel processing unit, the at least one parallel processing unit comprising one or more execution units that perform one or more groups of the one or more operations in parallel.
30. 26. The machine-readable medium of claim 25, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to group the one or more operations according to one or more parameters received as a result of one or more function calls to an interface provided by a software library, and for each group, perform the one or more operations using one or more kernels, the one or more kernels implementing the software library executed in parallel using the one or more computing resources.
31. 26. The machine-readable medium of claim 25, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to group the one or more operations according to at least one attribute of the one or more operations and perform each group of the one or more operations in parallel using the one or more computing resources, wherein the at least one attribute is indicative of a 5G-NR cell.
32. 26. The machine-readable medium of claim 25, wherein the one or more computing resources comprise at least one parallel processing unit, the at least one parallel processing unit operable to perform the one or more operations in parallel.
Citation Information
Patent Citations
Distributed batch normalization using estimates and rollback
US20200160123A1
Parallel de-rate-matching and layer demapping for physical uplink shared channel
WO2021080784A1