Parallel de-rate matching and layer demapping for physical uplink shared channel
Parallel processing techniques for de-rate matching and layer demapping in 5G NR systems improve the efficiency of wireless communication signal processing by reducing computational demands and processing time.
Patent Information
- Application Number
- JP2025047993
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-10-22
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-30
AI Technical Summary
The processing of wireless communication signals and data for decoding requires significant computing resources and time, necessitating improvements in efficiency.
Implementing parallel processors or computing systems for de-rate matching and layer demapping of wireless communication information, utilizing techniques such as layer demapping, descrambling, and de-rate matching in a parallel manner, particularly in the context of 5G New Radio (NR) physical uplink shared channels.
This approach reduces computational demands and processing time, enhancing the efficiency of wireless communication signal processing.
Smart Images

Figure 2025111440000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the priority of U.S. Patent Application No. 16 / 660,536, filed on October 22, 2019, entitled "PARALLEL DE-RATE-MATCHING AND LAYER DEMAPPING FOR PHYSICAL UPLINK SHARED CHANNEL", and incorporates the entire contents thereof by reference herein for all purposes.
[0002] At least one embodiment relates to processing resources used to process wireless communication information for decoding. For example, at least one embodiment relates to parallel processors or computing systems used for de-rate matching and layer demapping of wireless communication information by various novel techniques described herein.
Background Art
[0003] The processing of wireless communication signals and data for decoding may use significant computing resources and time. Solutions for processing wireless communication signals and data can be improved.
Brief Description of the Drawings
[0004]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13A
Figure 13B
Figure 13C
Figure 13D
Figure 13E
Figure 13F
Figure 14
Figure 15A
Figure 15B
Figure 16A
Figure 16B
Figure 17
Figure 18A
Figure 18B
Figure 18C
Figure 18D
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27A
Figure 27B
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Best Mode for Carrying Out the Invention
[0005] FIG. 1 is a block diagram showing a 5th generation (5G) New Radio (NR) signal processing environment 100 including a rate matching 102 and a scramble demodulation 104 according to at least one embodiment. In at least one embodiment, the scramble demodulation 104 is referred to as descrambling. In at least one embodiment, at least one of the rate matching 102 and the scramble demodulation 104 is implemented in a parallel manner, such as those described with respect to at least one of FIGS. 3-6. In at least one embodiment, at least one of the rate matching 102 and the scramble demodulation 104 is implemented by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processors or components thereof described and / or illustrated herein. In at least one embodiment, at least a part of the 5G NR signal processing environment 100 is included in a virtual radio access network (vRAN). In at least one embodiment, the 5G NR signal processing environment 100 includes a 5G vRAN stack 106 having a lower physical (PHY) layer 108, an upper PHY layer 110, a media access control (MAC) layer 112, a radio link control (RLC) layer 114, and a packet data convergence protocol (PDCP) layer 116. In at least one embodiment, the lower PHY layer 108 and the upper PHY layer 110 are not separately referenced but are referred to as the PHY layer. In at least one embodiment, the 5G vRAN stack 106 communicates with at least one user equipment (UE) 118 shown as UE1 - UEn via a radio frequency (RF) layer 120 and a wireless channel 122. In at least one embodiment, the 5G vRAN stack 106 communicates with a 5G packet core 124 using internet protocol (IP) packets.
[0006] In at least one embodiment, the lower PHY layer 108 and the upper PHY layer 110 include signal processing components 126 shown in an enlarged block diagram between an analog-to-digital converter (ADC) / digital-to-analog converter (DAC) 128 and the MAC layer 112. In at least one embodiment, the uplink path 130 includes orthogonal frequency division multiplexing (OFDM) demodulation 132, receiver (Rx) beamforming 134, channel estimation 136, channel equalization 138, descrambling demodulation 104, derate matching 102, low density parity check (LDPC) decoding 140, and transport block cyclic redundancy check (CRC) 142. In at least one embodiment, the downlink path 144 includes CRC segmentation 146, LDPC encoding 148, rate matching 150, scrambling modulation 152, precoding 154, transmission (Tx) beamforming 156, and OFDM modulation 158. In at least one embodiment, the uplink PHY layer operates as a virtual network function (VNF). In at least one embodiment, the VNF operating on the uplink PHY layer operates in a cluster computing environment. In at least one embodiment, the uplink PHY layer processes data related to a multiple-input multiple-output (MIMO) layer. In at least one embodiment, the uplink path 130 includes resource demapping.
[0007] FIG. 2 is a block diagram illustrating a 5G NR physical uplink shared channel (PUSCH) processing pipeline 200 according to at least one embodiment. In at least one embodiment, the 5G NR PUSCH processing pipeline 200 includes layer demapping 202, as well as descrambling and derate matching 204. In at least one embodiment, the descrambling and derate matching 204 corresponds to the descrambling demodulation 104 and derate matching 102 of FIG. 1. In at least one embodiment, at least one of the layer demapping 202, as well as the descrambling and derate matching 204, is implemented in a parallel manner, such as those described with respect to at least one of FIGS. 3-6. In at least one embodiment, at least one aspect of the 5G NR PUSCH processing pipeline 200 corresponds to at least one aspect of the uplink path 130 of FIG. 1. In at least one embodiment, at least one of the layer demapping 202, as well as the descrambling and derate matching 204, is implemented by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some of the other processors or components thereof described and / or illustrated herein.
[0008] In at least one embodiment, the RF layer 206 receives signals from at least one antenna 208. In at least one embodiment, the RF layer 206 receives signals from a plurality of antennas 208. In at least one embodiment, the antenna 208 provides reception of 5G NR MIMO signals. In at least one embodiment, the antenna 208 is a 5G NR antenna. In at least one embodiment, the PUSCH processing pipeline 200 includes performing fast fourier transform (FFT) and cyclic prefix (CP) removal in block 210. In at least one embodiment, the PUSCH processing pipeline 200 includes Rx beamforming in the midband in block 212. In at least one embodiment, the PUSCH processing pipeline 200 includes demodulation reference signal (DMRS) resource element (RE) demapping in block 214. In at least one embodiment, the PUSCH processing pipeline 200 includes orthogonal cover code (OCC) removal in block 216. In at least one embodiment, the PUSCH processing pipeline 200 includes interpolation based at least in part on the output of the OCC removal 216 in block 218 and pre-computing a DMR interpolation filter in block 220. In at least one embodiment, the PUSCH processing pipeline 200 includes equalizer filter calculation in block 222. In at least one embodiment, the PUSCH processing pipeline 200 includes PUSCH RE demapping in block 224 and equalization based at least in part on the outputs of the PUSCH RE demapping 224 and the equalizer filter calculation 222 in block 226. In at least one embodiment, the channel estimation 136 of FIG. 1 includes at least one element of the PUSCH processing pipeline 200, such as DMRS RE demapping 214, OCC removal 216, interpolation 218, and pre-computation of the DMRS interpolation filter 220.In at least one embodiment, the channel equalization 138 of FIG. 1 includes at least one element of the PUSCH processing pipeline 200, such as PUSCH RE demapping 224 and equalization 226.
[0009] In at least one embodiment, the PUSCH processing pipeline 200 includes soft demapping 228. In at least one embodiment, the PUSCH processing pipeline 200 performs layer demapping 202 based at least in part on the output of the soft demapping 228. In at least one embodiment, the PUSCH processing pipeline 200 includes LDPC decoding 230 based at least in part on the output of descrambling and rate matching 204. In at least one embodiment, the PUSCH processing pipeline 200 includes channel bonding (CB), carrier aggregation, and cyclic redundancy check (CRC) at block 232. In at least one embodiment, the PUSCH processing pipeline 200 provides the output of block 232 to the upper layer 234.
[0010] FIG. 3 is a block diagram showing layer demapping and deratematching 300 according to at least one embodiment. In at least one embodiment, layer demapping and deratematching 300 is implemented by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processors or components thereof described and / or illustrated herein. In at least one embodiment, 5G NR signal data is stored in an input memory array layout 304 according to a layer-time-frequency resource grid 302. In at least one embodiment, the 5G NR signal data of the input memory array layout 304 corresponds to signal information received from a plurality of antennas at a point in time after soft demapping, such as soft demapping 228 in the PUSCH processing pipeline, such as antenna 208 of FIG. 2. In at least one embodiment, the 5G NR signal data of the layer-time-frequency resource grid 302 is referred to as information received from a plurality of 5G NR antennas. In at least one embodiment, the 5G signal data of the layer-time-frequency resource grid 302 is referred to as information transmitted using a plurality of 5G NR signals. In at least one embodiment, the layer of the layer-time-frequency resource grid 302 refers to a MIMO layer. In at least one embodiment, layer demapping refers to demapping data of the MIMO layer.
[0011] In at least one embodiment, layer demapping and derate matching 300 takes the equalizer soft demapping output and processes it for further handling by the LDPC decoder. In at least one embodiment, the data of the layer-time-frequency resource grid 302 and the input memory array layout 304 is stored in transport blocks shown as TB 1 to TB N. In at least one embodiment, the data of the input memory array layout 304 includes data corresponding to log-likelihood ratios (LLRs). In at least one embodiment, the data of the input memory array layout 304 includes floating-point values representing the LLRs. In at least one embodiment, the data of the input memory array layout 304 includes integer values representing quantization of the LLRs. In at least one embodiment, the data of the input memory array layout 304 includes soft bits. In at least one embodiment, the data of the input memory array layout 304 is a vector of LLRs.
[0012] In at least one embodiment, layer demapping, such as layer demapping 202 of FIG. 2, includes extracting transport block 306 from data within input memory array layout 304 stored according to layer-time-frequency resource grid 302 in transport block extraction 308. In at least one embodiment, the transport block is segmented into a plurality of code blocks. In at least one embodiment, layer demapping includes, at code block extraction 312, extracting code block 310 from transport block 306, and for clarity, a portion of the extracted code blocks of TB 1 is shown. In at least one embodiment, layer demapping includes performing transport block extraction and code block extraction in a combined form such that code block 310 is extracted from the data of input memory array layout 304. In at least one embodiment, the code block includes data representing quadrature amplitude modulation (QAM) symbols. In at least one embodiment, the code block includes data representing some other type of symbol, such as frequency quadrature amplitude modulation (FQAM) symbols. In at least one embodiment, each code block includes one QAM symbol.
[0013] In at least one embodiment, layer demapping and derate matching 300 includes descrambling code block 310 to generate a descrambled code block 314. For clarity, in descrambling 316, one descrambled code block CBi is shown. In at least one embodiment, descrambling 316 corresponds to at least one of the descrambling of descrambling demodulation 104 of FIG. 1 or descrambling and derate matching 204 of FIG. 2. In at least one embodiment, layer demapping and derate matching 300 includes block interleaving the descrambled code block 314 to generate an interleaved code block 318. For clarity, in interleaving 320, one interleaved code block CBi is shown. In at least one embodiment, interleaving 320 uses a block interleaver that operates on each code block.
[0014] In at least one embodiment, layer demapping and derate matching 300 includes rate extension and filler bit insertion for the deinterleaved code block 318 to generate a derate-matched code block 322. For clarity, in the rate extension 324, one extended code block Exp CBi is shown. In at least one embodiment, the rate extension 324 corresponds to at least one of the derate matchings of the derate matching 102 in FIG. 1 or the descrambling and derate matching 204 in FIG. 2. In at least one embodiment, at least one of the derate-matched code blocks 322 includes a filler bit 326. In at least one embodiment, at least one of the derate-matched code blocks 322, although not shown for clarity, includes padding such as zero-padding. In at least one embodiment, the rate extension 324 includes extending the rate-matched (punctured) codeword to a full-length codeword by writing the LLRs to the corresponding positions, padding zeros for the punctured bits, and adding a predetermined value to the filler bit positions.
[0015] In at least one embodiment, layer demapping and derate matching 300 includes soft combining the derate-matched code block 322 with the corresponding code block 328 previously stored in the hybrid automatic repeat request (HARQ) buffer to generate a combined code block 330 in the HARQ buffer. For clarity, in soft combining 332, one code block 328 before combining and one combined code block 330 after combining in the HARQ buffer are shown. In at least one embodiment, soft combining 332 is performed as part of the derate matching 102 of FIG. 1 or as part of the derate matching of the descrambling and derate matching 204 of FIG. 2, so that the combined code block 330 can be used for LDPC decoding. In at least one embodiment, the previously stored corresponding code block 328 includes filler bits 334. In at least one embodiment, the filler bits 334 are in a different position than the filler bits 326. In at least one embodiment, both the filler bits 334 and the filler bits 326 are present within the combined code block 330 as shown. In at least one embodiment, soft combining 332 includes adding the value of the data in the derate-matched code block 322 to the value of the data in the corresponding previously stored code block 328. In at least one embodiment, the value in the HARQ buffer is repositioned by the a += operation during soft combining 332. In at least one embodiment, soft combining 332 includes combining the LLR with the contents of the HARQ buffer, which may include the LLR received in a previous HARQ transmission. In at least one embodiment, layer demapping and derate matching 300 uses a gather / scatter technique, using a gather operation to read from the input memory array layout 304 and a scatter operation to write to the output buffer in soft combining 332.
[0016] In at least one embodiment, at least one parallel processor performs layer demapping and derate matching 300 in parallel. In at least one embodiment, at least some aspects of layer demapping and derate matching 300 are performed by software executed on a graphics processing unit (GPU). In at least one embodiment, a plurality of thread blocks each having a plurality of threads perform layer demapping and derate matching 300 in parallel. In at least one embodiment, a thread block is referred to as a group of threads. In at least one embodiment, each code block is handled by a different thread block. In at least one embodiment, if a code block includes elements such as LLRs that are more than the maximum number of available threads in the initially assigned thread block, at least one additional thread block is assigned to handle the elements that exceed the maximum number of available threads in the initially assigned thread block. In at least one embodiment, if a code block includes elements such as LLRs that are more than the maximum number of available threads in the initially assigned thread block, the threads of the initially assigned thread block loop through the elements that exceed the maximum number of available threads, so that a subset more than one of the code block elements is sequentially and in parallel handled by the threads of the initially assigned thread block.
[0017] In at least one embodiment, for each code block, the corresponding thread block performs transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate extension 324, and soft combine 332. In at least one embodiment, for the thread block designated as i, the thread reads in[j] and writes out[k]+=s[m]×in[j]. Here, k = mapping_function(j). <parameters>) where out[k] is the HARQ buffer or a given HARQ process, s[j] is in the form of the set {+1, -1}, <parameters>includes an input index, a thread block index (mapped to a code block index and a transport block index), a transport block size, the number of multiple input multiple output (MIMO) layers, a modulation index, a coding rate, a code-based graph, and a redundancy version. In at least one embodiment, <parameters>And / or a subset of additional parameters is used. In at least one embodiment, k is used to demap j. In at least one embodiment, mapping_function is referred to as a demapping function. In at least one embodiment, mapping_function is referred to as a layer demapping function. In at least one embodiment, [j] refers to an array of values corresponding to the LLRs. In at least one embodiment, [j] corresponds to data in at least one of the layer-time-frequency resource grid 302, the input memory array layout 402 of FIG. 4, and the input buffer 502 of FIG. 5. In at least one embodiment, out[k] corresponds to at least one of the HARQ buffer described with respect to the previously stored corresponding code block 328 and the combined code block 330, the output memory array layout 404 of FIG. 4, and the output buffer 506 of FIG. 5.
[0018] In at least one embodiment, the thread block index includes a first thread block index corresponding to the x-dimension of the two-dimensional thread block array and a second thread block index corresponding to the y-dimension of the two-dimensional thread block array. In at least one embodiment, the thread index that identifies a particular thread within the thread block is <parameters>and is used as a parameter in at least one of <scr_parameters>. In at least one embodiment, <parameters>At least one of them includes at least one aspect described in the technical specification (TS) of the Third Generation Partnership Project (3GPP (registered trademark)), such as TS 38.212, TS 38.211, and / or TS 38.214 Release 15, version 15.6.0, or any other version and / or release.
[0019] In at least one embodiment, descrambling 316 is performed by changing the sign of in[j] based on the scrambling sequence s[m], where m = scr_mapping_function(j, <scr_parameters>). In at least one embodiment, s[m] is a pseudo-random sequence based at least in part on scr_mapping_function. In at least one embodiment, scr_mapping_function is referred to as a descrambling function. In at least one embodiment, changing the sign of in[j] includes multiplying in[j] by a negative sign in response to s[m] indicating that the sign of in[j] should be changed. In at least one embodiment, keeping the sign of in[j] the same value includes multiplying in[j] by a positive sign in response to s[m] indicating that the sign of in[j] should not be changed. In at least one embodiment, <scr_parameters> is described with respect to mapping_function <parameters>has the same parameters. In at least one embodiment, <scr_parameters> excludes the redundant version and <parameters>by including all parameters of, etc., <parameters>is a subset of. In some embodiments, <scr_parameters> is, <parameters>It includes at least one parameter not included in. In at least one embodiment, the soft combine 332 is performed when in[j] is aggregated to out[k]. In at least one embodiment, out[k] is initialized to zero.
[0020] In at least one embodiment, at least one circuit of at least one processor, system, and / or other device described herein decodes information received from a plurality of 5G NR antennas, such as antenna 208, in parallel by corresponding multiple processor pipelines. In at least one embodiment, the information received from the plurality of 5G NR antennas has been processed by soft demapping, such as soft demapping 228, and the decoded information includes at least one aspect of layer demapping and derate matching 300, such as layer demapping, descrambling, and derate matching. In at least one embodiment, the corresponding multiple processor pipelines refer to using at least one thread block per code block to perform layer demapping and derate matching 300. In at least one embodiment, instructions stored in a machine-readable medium, when executed, cause a parallel processor to transmit information using a plurality of 5G NR signals, such as data in a layer-time-frequency resource grid 302, and decode by the parallel processor by scheduling a plurality of thread groups respectively corresponding to at least one of the plurality of 5G NR signals in the parallel processor. In at least one embodiment, transmitting and decoding information using a plurality of 5G NR signals includes at least one aspect of layer demapping and derate matching 300, such as layer demapping, descrambling, and derate matching. In at least one embodiment, the plurality of thread groups respectively corresponding to at least one of the plurality of 5G NR signals refer to using at least one thread block per code block to perform layer demapping and derate matching 300.
[0021] FIG. 4 is a block diagram showing layer demapping and derate matching using a single read / write operation 400 according to at least one embodiment. In at least one embodiment, layer demapping and derate matching using a single read / write operation 400 is performed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or components thereof described and / or illustrated herein. In at least one embodiment, a thread of a thread block reads data elements from an input memory array layout 402, performs layer demapping and derate matching operations, and writes to an output memory array layout 404. In at least one embodiment, the input memory array layout 402 corresponds to the input memory array layout 304 of FIG. 3. In at least one embodiment, the output memory array layout 404 corresponds to a HARQ buffer layout, such as where the combined code block 330 of FIG. 3 is written. In at least one embodiment, the input memory array layout 402 is referred to as a first mapping configuration and the output memory array layout 404 is referred to as a second mapping configuration.
[0022] In at least one embodiment,
Number
[0023] In at least one embodiment, the threads of thread block 406 and the threads of thread block 408 perform layer demapping and derate matching, such as that described with respect to layer demapping and derate matching 300 of FIG. 3. In at least one embodiment, at least one thread of thread block 406 and / or at least one thread of thread block 408 perform actions not involving reading from input memory array layout 402, such as writing filler bits or padding zeros to output memory array layout 404. In at least one embodiment, at least one of thread block 406 and thread block 408 is started using a kernel startup function. In at least one embodiment, the kernel startup function is as contemplated with respect to layer demapping and derate matching 300 of FIG. 3 <parameters>By passing at least one parameter from <scr_parameters>, thread blocks 406 and 408 are launched. In at least one embodiment, the kernel launch function is used to index at least one of the thread blocks and threads, <parameters>Calculate at least one parameter of <scr_parameters>. In at least one embodiment, a single read and write operation from global memory is performed for each element (LLR floating point value). In at least one embodiment, all of transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate extension 324, and soft combine 332 are performed using a single read write operation.
[0024] In at least one embodiment, a subset of threads inserts filler bits as part of rate extension 324. In at least one embodiment, the number F of filler bits is a parameter provided to the kernel by a kernel startup function. In at least one embodiment, the number of filler bits is calculated using K', based at least in part on the difference between the number K of information bits corresponding to a selected LDPC lifting size Zc and the number K' of bits in a code block. In at least one embodiment, the number of filler bits is calculated by a central processing unit (CPU) and passed to a parallel processor such as a GPU using a kernel startup function. In at least one embodiment, a parallel processor such as a GPU calculates the number of filler bits. In at least one embodiment, the CPU calculates a pointer to the position where each code block starts in the input memory array layout 402 and passes the pointer as a startIndex parameter to a parallel processor such as a GPU using a kernel startup function. In at least one embodiment, a parallel processor such as a GPU calculates startIndex. In at least one embodiment, the CPU calculates a code block size E and passes E to a parallel processor such as a GPU using a kernel startup function. In at least one embodiment, a parallel processor such as a GPU calculates E.
[0025] Figure 5 is a block diagram showing layer demapping and derate matching using shared memory operation 500 according to at least one embodiment. In at least one embodiment, layer demapping and derate matching using shared memory operation 500 is performed by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or its components described and / or illustrated herein. In at least one embodiment, data elements of input buffer 502 are read in parallel by a plurality of threads into shared memory buffer 504 in shared memory. In at least one embodiment, input buffer 502 has an input memory array layout corresponding to at least one of input memory array layout 402 or input memory array layout 304. In at least one embodiment, input buffer 502 is in global memory. In at least one embodiment, input buffer 502 is in local memory. In at least one embodiment, the first thread, indicated as threadIdx.x = 0 (also referred to as thread 0), reads the first data element from the first position of input buffer 502 to the first position of shared memory buffer 504. In at least one embodiment, the second thread, indicated as threadIdx.x = 1 (also referred to as thread 1), reads the second data element from the second position of input buffer 502 to the second position of shared memory buffer 504. In at least one embodiment, thread 0 and thread 1 read consecutive positions from global memory. In at least one embodiment, a thread block including thread 0 and thread 1 reads N_tbsz consecutive elements of input buffer 502. In at least one embodiment, N refers to the number of thread blocks, and tbsz refers to the thread block size. In at least one embodiment, N_tbsz refers to the number of elements in a total of N thread blocks. In at least one embodiment, N_tbsz refers to the number of elements in a particular thread block. In at least one embodiment, the threads are synchronized after reading the data elements into shared memory buffer 504.In at least one embodiment, the threads are synchronized using the a_syncthreads() operation, but it should be understood that in other embodiments, other thread synchronization techniques may be used.
[0026] In at least one embodiment, after the data elements are read into the shared memory buffer 504, the threads of the thread block perform layer demapping and deratematching of the data elements of the code block. In at least one embodiment, the threads of the thread block use the data elements read into the shared memory buffer 504 to perform at least one of transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate extension 324, and soft combining 332. In at least one embodiment, threadIdx.x = 0 reads from the first position of the shared memory buffer 504, performs at least one of transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate extension 324, and soft combining 332, and writes the first position of the first transport block TB 1 to the output buffer 506. In at least one embodiment, the output buffer 506 is in global memory. In at least one embodiment, the output buffer 506 is a HARQ buffer. In at least one embodiment, the output buffer 506 corresponds to the output memory array layout 404. In at least one embodiment, threadIdx.x = 1 reads from the position of the shared memory buffer 504, performs at least one of transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate extension 324, and soft combining 332, and writes the second position of TB 1 to the output buffer 506. In at least one embodiment, threadIdx.x = 1 is at the position 0 + of the shared memory buffer 504
Number
[0027] In at least one embodiment, threads of a thread block, such as Thread 0 and Thread 1, perform transport block extraction 308 and code block extraction 312 when reading from input buffer 502 to shared memory buffer 504. In at least one embodiment, threads of a thread block, such as Thread 0 and Thread 1, perform descrambling 316, deinterleaving 320, and rate expansion 324 on data elements in shared memory buffer 504. In at least one embodiment, threads of a thread block, such as Thread 0 and Thread 1, perform soft combine 332 when writing to output buffer 506.
[0028] In at least one embodiment, a block of elements (LLR floating-point values) is read into shared memory buffer 504 in shared memory using a combined global memory read operation. In at least one embodiment, the size of the block of elements is equal to the size of the thread block. In at least one embodiment, each thread executes transport block extraction 308, code block extraction 312, descrambling 316, deinterleaving 320, rate expansion 324, and soft combine 332 and writes to global memory, such as to output buffer 506, using a combined global memory write operation. In at least one embodiment, the combined global memory read and write operations are used in combination with shared memory to provide faster processing times. In at least one embodiment, a thread writes continuously to shared memory, such as shared memory buffer 504, and reads discontinuously from shared memory. In at least one embodiment, a thread writes discontinuously to shared memory, such as shared memory buffer 504, and reads continuously from shared memory.
[0029] FIG. 6 shows a flowchart of a layer demapping and deratematching technique 600 according to at least one embodiment. In at least one embodiment, the technique 600 is implemented by at least one circuit, at least one system, at least one processor, at least one graphics processing unit, at least one parallel processor, and / or at least some other processor or components thereof described and / or illustrated herein. In at least one embodiment, a plurality of threads of at least one thread block perform at least one aspect of the technique 600 in parallel. In at least one embodiment, the technique 600 includes, at block 602, receiving information from a plurality of 5G NR antennas such as antenna 208. In at least one embodiment, at block 604, the technique 600 includes extracting from the received information a transport block and a code block. In at least one embodiment, extracting from the received information a transport block and a code block corresponds to transport block extraction 308 and code block extraction 312 of FIG. 3.
[0030] In at least one embodiment, at block 606, technique 600 includes layer demapping the extracted code blocks. In at least one embodiment, the layer demapping of block 606 corresponds to the layer demapping 202 of FIG. 2. In at least one embodiment, at block 608, technique 600 includes descrambling the code blocks. In at least one embodiment, descrambling the code blocks at block 608 corresponds to at least one of the descrambling 316 of FIGS. 1-3, the descrambling of the descrambling and rate matching 204, and the descrambling of the scramble demodulation 104. In at least one embodiment, at block 610, technique 600 includes deinterleaving the code blocks. In at least one embodiment, deinterleaving the code blocks at block 610 corresponds to the deinterleaving 320 of FIG. 3. In at least one embodiment, at block 612, technique 600 includes rate matching the code blocks. In at least one embodiment, rate matching the code blocks at block 612 corresponds to at least one of the rate matching 102 of FIGS. 1-3, the rate matching of the descrambling and rate matching 204, and the rate expansion 324. In at least one embodiment, at block 614, technique 600 includes soft combining the layer demapped, descrambled, deinterleaved, and rate matched code blocks using the content of the HARQ buffer. In at least one embodiment, the soft combining of block 614 corresponds to the soft combining 332 of FIG. 3. In at least one embodiment, at block 616, technique 600 includes performing other actions.
[0031] Data center FIG. 7 shows an example data center 700 in which at least one embodiment may be used. In at least one embodiment, data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0032] In at least one embodiment, as shown in FIG. 7, the data center infrastructure layer 710 may include a resource orchestrator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1) to 716(N), where “N” represents any positive integer. In at least one embodiment, the node C.R. 716(1) to 716(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., semiconductor drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R. 716(1) to 716(N) may be servers having one or more of the computing resources described above.
[0033] In at least one embodiment, the grouped computing resources 714 may include a separate group of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). The separate group of node C.R.s within the grouped computing resources 714 may be configured or allocated to support one or more workloads, and may include grouped compute resources, network resources, memory resources, or storage resources. In at least one embodiment, some node C.R.s including a CPU or processor may be grouped within one or more racks to provide compute resources that support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0034] In at least one embodiment, the resource orchestrator 712 may configure or otherwise control one or more node C.R.s 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the resource orchestrator 712 may include a software design infrastructure ("SDI") management entity for the data center 700. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.
[0035] In at least one embodiment, as shown in FIG. 7, the framework layer 720 includes a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, the framework layer 720 may include a framework that supports software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 732 or the application 742 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be an open-source software web application framework, such as Apache Spark (trademark) (hereinafter "Spark"), which may be free and may use the distributed file system 738 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be capable of configuring different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 to support large-scale data processing. In at least one embodiment, the resource manager 736 may be capable of managing clustered or grouped computing resources mapped or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 714 in the data center infrastructure layer 710.In at least one embodiment, the resource manager 736 may manage these mapped or allocated computing resources in cooperation with the resource orchestrator 712.
[0036] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the node C.R. 716(1) - 716(N), the grouped computing resources 714, and / or the distribution file system 738 of the framework layer 720. In at least one embodiment, the one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0037] In at least one embodiment, the application 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node C.R. 716(1) - 716(N), the grouped computing resources 714, and / or the distribution file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning applications including machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0038] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource orchestrator 712 may implement any number and type of self-corrective measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective measures may prevent the data center operator of the data center 700 from determining a configuration that may be defective and may eliminate parts of the data center that are not fully utilized and / or have low performance.
[0039] In at least one embodiment, the data center 700 may include tools, services, software, or other resources that train one or more machine learning models according to one or more embodiments described herein or that predict or infer information using one or more machine learning models. For example, in at least one embodiment, the machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 700. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 700 by using weight parameters calculated by one or more of the training techniques described herein.
[0040] In at least one embodiment, the data center may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service that enables a user to perform training or inference of information, such as image recognition, voice recognition, or other artificial intelligence services.
[0041] In at least one embodiment, at least one component illustrated or described with respect to FIG. 7 is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one of the grouped computing resources 714 and node C.R. 716 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one of the grouped computing resources 714 and node C.R. 716 is used for layer demapping, descrambling, and de - rate matching of 5G NR PUSCH data that is soft - demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0042] Autonomous vehicle FIG. 8A shows an example of an autonomous vehicle 800 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 800 (or, referred to herein as "vehicle 800") may be a passenger vehicle such as, without limitation, a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 800 may be a trailer truck of a semi - tractor used for cargo transportation. In at least one embodiment, vehicle 800 may be an aircraft, robotic vehicle, or another type of vehicle.
[0043] The autonomous vehicle may be described from the perspective of the automation level defined by the National Highway Traffic Safety Administration (NHTSA), a department of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE)'s "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806 issued on June 15, 2018, Standard No. J3016-201609 issued on September 30, 2016, and the old and new versions of this standard). In one or more embodiments, the vehicle 800 may be capable of corresponding to the functionality according to one or more of automation levels 1 to 5 of the autonomous driving level. For example, in at least one embodiment, the vehicle 800 may be capable of corresponding to conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5) according to the embodiment.
[0044] In at least one embodiment, the vehicle 800 may include components such as, without limitation, a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, the vehicle 800 may include a propulsion system 850 such as, without limitation, an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, the propulsion system 850 may be connected to the drive train of the vehicle 800, and the drive train may include, without limitation, a transmission for enabling the propulsion of the vehicle 800. In at least one embodiment, the propulsion system 850 may be controlled in response to receiving a signal from the throttle / accelerator 852.
[0045] In at least one embodiment, a steering system 854, which may include a steering wheel without limitation, is used to steer a vehicle 800 (e.g., along a desired path or route) when the propulsion system 850 is operating (e.g., when the vehicle is moving). In at least one embodiment, the steering system 854 may receive a signal from a steering actuator 856. The steering wheel may be optional with respect to fully automated (level 5) functionality. In at least one embodiment, a brake sensor system 846 may be used to operate the vehicle brakes in response to receiving a signal from a brake actuator 848 and / or a brake sensor.
[0046] In at least one embodiment, controller 836 may include one or more system-on-chips (“SoCs”) (not shown in FIG. 8A) and / or graphics processing units (“GPUs”) without limitation and provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 800. For example, in at least one embodiment, controller 836 may transmit signals to operate the vehicle brakes via brake actuator 848, signals to operate steering system 854 via steering actuator 856, and signals to operate propulsion system 850 via throttle / accelerator 852. Controller 836 may include one or more integrated (e.g., monolithic) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in operating vehicle 800. In at least one embodiment, controller 836 may include a first controller 836 for autonomous driving functions, a second controller 836 for functional safety functions, a third controller 836 for artificial intelligence functions (e.g., computer vision), a fourth controller 836 for infotainment functions, a fifth controller 836 for redundancy in emergencies, and / or other controllers. In at least one embodiment, a single controller 836 may handle two or more of the above functionalities, two or more controllers 836 may handle a single functionality, and / or any combination thereof may be possible.
[0047] In at least one embodiment, the controller 836 provides signals for controlling one or more components and / or systems of the vehicle 800 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may be received from, for example and without limitation, a global navigation satellite system (GNSS) sensor 858 (e.g., a global positioning system sensor), a RADAR sensor 860, an ultrasonic sensor 862, a LIDAR sensor 864, an inertial measurement unit (IMU) sensor 866 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 896, a stereo camera 868, a wide-angle camera 870 (e.g., a fish-eye camera), an infrared camera 872, a surround camera 874 (e.g., a 360-degree camera), a long-range camera (not shown in FIG. 8A), a mid-range camera (not shown in FIG. 8A), a speed sensor 844 (e.g., for measuring the speed of the vehicle 800), a vibration sensor 842, a steering sensor 840, a brake sensor (e.g., as part of the brake sensor system 846), and / or other types of sensors.
[0048] In at least one embodiment, one or more of the controllers 836 receive an input (e.g., represented by input data) from the instrument cluster 832 of the vehicle 800 and provide an output (e.g., represented by output data, display data, etc.) via the human-machine interface (「HMI」) display 834, an audible annunciator, a speaker, and / or via other components of the vehicle 800. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high-definition map (not shown in FIG. 8A), location data (e.g., the location of the vehicle 800 on a map, etc.), direction, the location of other vehicles (e.g., an occupancy grid), information about objects and the state of the objects recognized by the controller 836. For example, in at least one embodiment, the HMI display 834 may display information about the presence of one or more objects (e.g., road signs, warning signs, signal changes, etc.) and / or information about driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B in 3.22 km (2 miles), etc.).
[0049] In at least one embodiment, vehicle 800 further includes a network interface 824, which may use a wireless antenna 826 and / or a modem for communicating via one or more networks. For example, in at least one embodiment, network interface 824 may be capable of communicating via Long-Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA (R)”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile communication (“GSM”), IMT-CDMA multi-carrier (“CDMA2000”), and the like. Also, in at least one embodiment, wireless antenna 826 may use local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, and / or low power wide-area networks (“LPWAN”) such as LoRaWAN, SigFox to enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.).
[0050] In at least one embodiment, the signal received from antenna 208 of FIG. 2 is from vehicle 800 and may be processed as described with respect to at least one of FIGS. 1-6 to provide vehicle 800 with information for its autonomous operation, such as weather data, navigation data, road condition data, etc., and / or may be used to provide a remote operator with the ability to remotely control vehicle 800.
[0051] FIG. 8B shows an example of the camera position and field of view for the autonomous vehicle 800 of FIG. 8A according to at least one embodiment. In at least one embodiment, the camera and its respective field of view are one example of an embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the vehicle 800.
[0052] In at least one embodiment, the camera type of the camera may include, but is not limited to, a digital camera that may be adapted to be used with the components and / or systems of the vehicle 800. The camera may operate at automotive safety integrity level (ASIL) B and / or another ASIL. In at least one embodiment, the camera type may be capable of corresponding to any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a color filter array of red, clear, clear, clear (RCCC: red clear clear clear), a color filter array of red, clear, clear, blue (RCCB: red clear clear blue), a color filter array of red, blue, green, clear (RBGC: red blue green clear), a color filter array of Foveon X3, a color filter array of a Bayer sensor (RGGB), a color filter array of a monochrome sensor, and / or another type of color filter array. In at least one embodiment, a clear pixel camera, such as a camera having an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.
[0053] In at least one embodiment, one or more cameras may be used to implement advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0054] In at least one embodiment, one or more of the cameras may be attached to a mounting assembly such as a custom-designed (three-dimensional (3D) printed) assembly to eliminate stray light and reflections from inside the vehicle (e.g., reflections reflected from the dashboard to the windshield) that may interfere with the camera's image data capture performance. Referring to the door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed such that the camera mounting plate conforms to the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. In at least one embodiment, in the case of a side view camera, the camera may also be integrated within four pillars located at each corner of the vehicle.
[0055] In at least one embodiment, a camera (e.g., a front camera) having a field of view that includes a portion of the environment in front of vehicle 800 is used for the surrounding view to facilitate identification of the front path and obstacles, and may be used with one or more of controller 836 and / or the control SoC to assist in providing information essential for the generation of an occupancy grid and / or the determination of a preferred vehicle path. In at least one embodiment, many of the same ADAS functions as LIDAR, including without limitation emergency braking, pedestrian detection, and collision avoidance, may be implemented using a front camera. In at least one embodiment, the front camera may also be used for ADAS functions and systems, including without limitation other functions such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0056] In at least one embodiment, various cameras may be used in a front configuration, including, for example, a platform of a monocular camera including a CMOS ("complementary metal oxide semiconductor") color imaging device. In at least one embodiment, a wide-angle camera 870 may be used to perceive objects (e.g., pedestrians, cross traffic, or bicycles) entering the view from the surroundings. Although only one wide-angle camera 870 is shown in FIG. 8B, in other embodiments, any number (including zero) of wide-angle cameras 870 may be present on vehicle 800. In at least one embodiment, any number of long-range cameras 898 (e.g., a pair of stereo cameras for a long-range view) may be used for depth-based object detection, especially for objects for which a neural network has not yet been trained thereon. In at least one embodiment, the long-range cameras 898 may also be used for object detection and classification, and basic object tracking.
[0057] In at least one embodiment, any number of stereo cameras 868 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 868 may include an integrated control unit with a scalable processing unit, which control unit may provide a programmable logic (“FPGA”) and a multi-core microprocessor having an integrated controller area network (“CAN”) or Ethernet® interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 800, including distance estimation for all points within the image. In at least one embodiment, one or more of the stereo cameras 868 may include, without limitation, a compact stereo vision sensor, which sensor may measure the distance from the vehicle 800 to a target object and use the generated information (e.g., metadata) to activate functions such as autonomous emergency braking and lane departure warning, and may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip. In at least one embodiment, in addition to, or instead of, those described herein, other types of stereo cameras 868 may be used.
[0058] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view that includes a portion of the environment to the side of vehicle 800 may be used for the surrounding view to provide information for creating and updating an occupancy grid and for generating a side collision warning. For example, in at least one embodiment, surround cameras 874 (e.g., four surround cameras 874 as shown in FIG. 8B) may be positioned on vehicle 800. The surround cameras 874 may include, without limitation, any number and combination of wide-angle cameras 870, fish-eye cameras, 360-degree cameras, and / or others. For example, in at least one embodiment, four fish-eye cameras may be positioned in front of, behind, and to the sides of vehicle 800. In at least one embodiment, vehicle 800 may use three surround cameras 874 (e.g., left, right, and rear), and one or more other cameras (e.g., a front camera) may be utilized as a fourth surround camera.
[0059] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment behind vehicle 800 may be used for parking assistance, the surrounding view, and rear collision warning, and an occupancy grid may be created and updated. In at least one embodiment, a variety of cameras may be used, including but not limited to cameras suitable as front cameras as described herein (e.g., long-range camera 898, and / or mid-range camera 876, stereo camera 868), infrared camera 872, etc.
[0060] In at least one embodiment, the signal received from antenna 208 in FIG. 2 is from vehicle 800 and may be processed as described with respect to at least one of FIGS. 1-6 to provide vehicle 800 with information for its autonomous operation, such as weather data, navigation data, road condition data, etc., and / or may be used to provide a remote operator with the ability to remotely control vehicle 800.
[0061] FIG. 8C is a block diagram showing a system architecture example related to the autonomous vehicle 800 of FIG. 8A according to at least one embodiment. In at least one embodiment, the components, features, and systems of the vehicle 800 in FIG. 8C are each shown as being connected via a bus 802. In at least one embodiment, the bus 802 may include, without limitation, a CAN data interface (or, as referred to herein, a "CAN bus"). In at least one embodiment, CAN may be a network inside the vehicle 800 that is used to assist in controlling various features and functions of the vehicle 800, such as brake actuation, acceleration, brake control, steering, windshield wipers, etc. In at least one embodiment, the bus 802 may be configured to have dozens or even hundreds of nodes, each having its own unique identifier (e.g., CAN ID). In at least one embodiment, the bus 802 may be read to find steering wheel angle, ground speed, engine revolutions per minute ("RPM"), button position, and / or other vehicle state indicators. In at least one embodiment, the bus 802 may be a CAN bus compliant with ASIL B.
[0062] In at least one embodiment, in addition to or instead of CAN, FlexRay and / or Ethernet® may be used. In at least one embodiment, any number of buses 802 may be present, including, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 802 may be used to perform different functions and / or to provide redundancy. For example, a first bus 802 may be used for a collision avoidance function and a second bus 802 may be used for actuation control. In at least one embodiment, each bus 802 may communicate with any of the components of vehicle 800, and two or more buses 802 may communicate with the same component. In at least one embodiment, each of any number of system-on-chips (“SoC”) 804, controllers 836, and / or each computer within the vehicle may be accessible to the same input data (e.g., an input from a sensor of vehicle 800) and may be connected to a common bus such as a CAN bus.
[0063] In at least one embodiment, vehicle 800 may include one or more controllers 836, such as those described herein with respect to FIG. 8A. Controller 836 may be used for various functions. In at least one embodiment, controller 836 may be coupled to any of the various other components and systems of vehicle 800 and may be used for control of vehicle 800, the artificial intelligence of vehicle 800, and / or the infotainment of vehicle 800.
[0064] In at least one embodiment, vehicle 800 may include any number of SoCs 804. Each SoC 804 may include, without limitation, a central processing unit (“CPU”) 806, a graphics processing unit (“GPU”) 808, a processor 810, a cache 812, an accelerator 814, a data store 816, and / or other components and features not shown. In at least one embodiment, the SoC 804 may be used to control the vehicle 800 in various platforms and systems. For example, in at least one embodiment, the SoC 804 may be incorporated into a system (e.g., the system of vehicle 800) having a high-definition (“HD”) map 822 that may obtain map refreshes and / or updates via network interface 824 from one or more servers (not shown in FIG. 8C).
[0065] In at least one embodiment, the CPU 806 may include a CPU cluster, or a CPU complex (or “CCPLEX” as referred to herein). In at least one embodiment, the CPU 806 may include multiple cores and / or a level 2 (“L2”) cache. For example, in at least one embodiment, the CPU 806 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 806 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2MB L2 cache). In at least one embodiment, the CPU 806 (e.g., CCPLEX) may be configured to support simultaneous cluster operation that enables any combination of clusters of the CPU 806 to be activated at any given time.
[0066] In at least one embodiment, one or more of the CPUs 806 may implement a power management function, which may include, without limitation, one or more of the following features: individual hardware blocks can be automatically clock-gated during idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to the execution of a wait for interrupt ("WFI") / wait for event ("WFE") instruction; each core can be power-gated independently; when all cores are clock-gated or power-gated, each core cluster can be clock-gated independently; and / or when all cores are power-gated, each core cluster can be power-gated independently. In at least one embodiment, the CPU 806 may further implement an extended algorithm for managing power states, where the allowed power states and the expected wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. In at least one embodiment, the processing cores may support, in software, a simple sequence for entering a power state with the work offloaded to microcode.
[0067] In at least one embodiment, GPU 808 may include an integrated GPU (or, as referred to herein, an "iGPU"). In at least one embodiment, GPU 808 may be programmable and may be efficient for parallel workloads. In at least one embodiment, GPU 808 may use an extended tensor instruction set. In one embodiment, GPU 808 may include one or more streaming microprocessors, each streaming microprocessor may include a level 1 ("L1") cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache having a storage capacity of 512 KB). In at least one embodiment, GPU 808 may include at least eight streaming microprocessors. In at least one embodiment, GPU 808 may use a compute application programming interface (API). In at least one embodiment, GPU 808 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0068] In at least one embodiment, one or more of the GPUs 808 may be power optimized to provide optimal performance in automotive and embedded use cases. For example, in one embodiment, the GPU 808 can be fabricated on a fin field effect transistor ( "FinFET"). In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores partitioned into multiple blocks. For example, without limitation, 64 PF32 cores and 32 PF64 cores can be partitioned into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR cores for deep learning matrix operations, a level zero ( "L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor may include independent parallel data paths for integers and floating points to provide efficient execution of workloads by a combination of computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0069] In at least one embodiment, one or more of the GPUs 808 include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem, and in some examples, may provide a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as five synchronous random access memories of the graphics double data rate type (GDDR5), may be used.
[0070] In at least one embodiment, the GPU 808 may include integrated memory technology. In at least one embodiment, address translation service (ATS) support may be used to enable the GPU 808 to directly access the page table of the CPU 806. In at least one embodiment, when the GPU 808 memory management unit (MMU) encounters a miss, an address translation request may be sent to the CPU 806. In at least one embodiment, in response, the CPU 806 may search its page table for the virtual-to-physical address mapping and send the translation back to the GPU 808. In at least one embodiment, the integrated memory technology may enable a single integrated virtual address space for both the memory of the CPU 806 and the GPU 808, thereby simplifying the programming of the GPU 808 and the porting of applications to the GPU 808.
[0071] In at least one embodiment, the GPU 808 may include any number of access counters that can record the frequency of access of the GPU 808 to the memory of other processors. In at least one embodiment, the access counter may assist in ensuring that memory pages are moved to the physical memory of the processor that most frequently accesses the page, thereby improving the efficiency of the memory range shared among the processors.
[0072] In at least one embodiment, one or more of the SoCs 804 may include any number of caches 812, including those described herein. For example, in at least one embodiment, the cache 812 can include a level 3 ("L3") cache that is available to both the CPU 806 and the GPU 808 (e.g., connected to both the CPU 806 and the GPU 808). In at least one embodiment, the cache 812 may include a write-back cache that can record the state of the line, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4MB or more, depending on the embodiment, although smaller cache sizes may be used.
[0073] In at least one embodiment, one or more of the SoCs 804 may include one or more accelerators 814 (e.g., hardware accelerators, software accelerators, or combinations thereof). In at least one embodiment, the SoC 804 may include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memories. In at least one embodiment, a large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU 808 and offload some of the tasks of the GPU 808 (e.g., free up more cycles of the GPU 808 to perform other tasks). In at least one embodiment, the accelerator 814 can be used for workloads (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.) that are stable enough to accept acceleration. In at least one embodiment, the CNN may include region-based, i.e., region convolutional neural networks (“RCNNs”), and (e.g., used for object detection) fast RCNNs, or other types of CNNs.
[0074] In at least one embodiment, the accelerator 814 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include, without limitation, one or more Tensor processing units (TPUs), which may be further configured to provide up to 10 trillion operations per second for deep learning applications and inferences. In at least one embodiment, the TPU may be an accelerator configured and optimized to perform image processing functions (e.g., for CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating point operations, as well as for inferences. In at least one embodiment, the design of the DLA can improve performance per millimeter over a typical general-purpose GPU and typically far exceeds the performance of a CPU. In at least one embodiment, the TPU may perform several functions, including, for example, a single instance of a convolutional function that supports INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions. In at least one embodiment, the DLA may execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including, without limitation, CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection, identification, and detection using data from a microphone 896, CNNs for face recognition and vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events.
[0075] In at least one embodiment, the DLA may implement any function of the GPU 808. For example, by using an inference accelerator, the designer may target either the DLA or the GPU 808 for any function. For example, in at least one embodiment, the designer may concentrate the processing of CNN and floating-point operations on the DLA and leave other functions to the GPU 808 and / or other accelerators 814.
[0076] In at least one embodiment, the accelerator 814 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS) 838, autonomous driving, augmented reality (AR) applications, and / or virtual reality (VR) applications. The PVA may maintain a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors, without limitation.
[0077] In at least one embodiment, the RISC core may interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or others. In at least one embodiment, each RISC core may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0078] In at least one embodiment, the DMA may enable the components of the PVA to access system memory independently of the CPU 806. In at least one embodiment, the DMA may support any number of features used to optimize the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0079] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or vector memory (e.g., "VMEM"). In at least one embodiment, the VPU core may include a digital signal processor such as, for example, a single instruction multiple data ("SIMD"), very long instruction word ("VLIW") digital signal processor. In at least one embodiment, the combination of SIMD and VLIW may improve throughput and speed.
[0080] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to execute independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on successive images or on portions of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of the PVAs. In at least one embodiment, the PVA may include additional error correcting code ( "ECC") memory to enhance the overall security of the system.
[0081] In at least one embodiment, the accelerator 814 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory (SRAM) to provide high-bandwidth, low-latency SRAM to the accelerator 814. In at least one embodiment, the on-chip memory may include at least 4MB of SRAM, which may consist of, for example and without limitation, eight field-configurable memory blocks, which may be accessible from either the PVA or the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone that provides fast access to the memory for the PVA and DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using the APB).
[0082] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and DLA provide a ready signal and a valid signal before transmitting any control signal / address / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signal / address / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, the interface may comply with the standards of the International Organization for Standardization (ISO) 26262 or the International Electrotechnical Commission (IEC) 61508, although other standards and protocols may be used.
[0083] In at least one embodiment, one or more of the SoCs 804 may include a hardware accelerator for real-time ray tracing. In at least one embodiment, the hardware accelerator for real-time ray tracing is used to quickly and efficiently determine the position and extent of an object (e.g., within a world model) for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of a SONAR system, for simulation of general waveform propagation, for comparison with LIDAR data for localization and / or other functions, and / or for real-time visualization simulation for other uses.
[0084] In at least one embodiment, the accelerator 814 (e.g., a hardware accelerator cluster) has various uses for autonomous driving. In at least one embodiment, the PVA may be a programmable vision accelerator that can be used in the main processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well-suited to algorithm domains that require predictable processing with low power and low latency. In other words, the PVA functions well even with a small data set in semi-dense or dense regular calculations that require a predictable runtime with low latency and low power. In at least one embodiment, in an autonomous vehicle such as the vehicle 800, the PVA is designed to execute conventional computer vision algorithms because they are effective for object detection and integer arithmetic.
[0085] For example, according to at least one embodiment of the technology, PVA is used to implement computer stereo vision. In at least one embodiment, although an algorithm based on semi-global matching may be used in some examples, this is not intended to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may implement a computer stereo vision function for inputs from two monocular cameras.
[0086] In at least one embodiment, PVA may be used to implement dense optical flow. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and for example, by processing raw time-of-flight data, processed time-of-flight data is provided.
[0087] In at least one embodiment, for example without limitation, a DLA may be used to execute any type of network for enhancing control and driving safety, including a neural network that outputs a measure of reliability for each object detection. In at least one embodiment, reliability may be represented or interpreted as the probability of each detection compared to other detections or as providing its relative "weight". In at least one embodiment, reliability enables the system to make further decisions regarding which detections should be considered positive detections rather than false detections. For example, in at least one embodiment, the system may set a threshold for reliability and consider only detections that exceed the threshold as positive detections. In one embodiment where an automatic emergency braking ("AEB") system is used, a false detection would cause the vehicle to automatically apply emergency brakes, which is clearly undesirable. In at least one embodiment, very highly reliable detections may be considered as triggers for AEB. In at least one embodiment, the DLA may execute the neural network to regress a confidence value. In at least one embodiment, the neural network may take as its input at least some subset of parameters such as, among others, the dimensions of the bounding box, ground estimation obtained (e.g., from another subsystem), the output from the IMU sensor 866 correlated with the orientation of the vehicle 800, distance, and the 3D location estimation of the object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 864 or the RADAR sensor 860).
[0088] In at least one embodiment, one or more of the SoCs 804 may include a data store 816 (e.g., memory). In at least one embodiment, the data store 816 may be on-chip memory of the SoC 804, and this memory may store neural networks executed on the GPU 808 and / or DLA. In at least one embodiment, the capacity of the data store 816 may be sufficient to store multiple instances of the neural network for redundancy and security. In at least one embodiment, the data store 812 may include an L2 or L3 cache.
[0089] In at least one embodiment, one or more of the SoCs 804 may include any number of processors 810 (e.g., embedded processors). The processor 810 may include a boot and power management processor, which may be a dedicated processor and subsystem that handles boot power as well as management functions and related security execution. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 804 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance in transitioning the system to a low power state, management of the thermal and temperature sensors of the SoC 804, and / or management of the power state of the SoC 804. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 804 may use the ring oscillator to detect the temperature of the CPU 806, GPU 808, and / or accelerator 814. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine, put the SoC 804 in a low power state, and / or put the vehicle 800 in a driver-safety stop mode (e.g., safely stop the vehicle 800).
[0090] In at least one embodiment, the processor 810 may further include a set of embedded processors that can serve as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables complete hardware support for multi-channel audio via a multi-interface and a wide variety of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0091] In at least one embodiment, the processor 810 may further include an always-on processor engine that provides the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0092] In at least one embodiment, the processor 810 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for handling safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, support peripherals (such as timers and interrupt controllers, etc.), and / or routing logic. In the safety mode, in at least one embodiment, two or more cores may operate in a lockstep mode and function as a single core having comparison logic for detecting any differences between these operations. In at least one embodiment, the processor 810 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, the processor 810 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0093] In at least one embodiment, the processor 810 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image in the window of the playback device. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 870, the surround camera 874, and / or the in-cabin monitoring camera / sensor. In at least one embodiment, the in-cabin monitoring camera / sensor is preferably monitored by a neural network running on another instance of the SoC 804 that is configured to identify events in the cabin and respond thereto as appropriate. In at least one embodiment, the in-cabin system may perform, without limitation, lip reading to activate cellular service, make a phone call, write an email, change the destination of the vehicle, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable otherwise.
[0094] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in the video, the noise reduction appropriately weights the spatial information to reduce the weight of the information provided by adjacent frames. In at least one embodiment, when an image or a portion of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from the previous image to reduce the noise of the current image.
[0095] In at least one embodiment, the video image synthesizer may also be configured to perform stereo parallelization on the input stereo lens frames. In at least one embodiment, the video image synthesizer may further be used to synthesize a user interface when the desktop of the operating system is in use, and the GPU 808 does not need to continuously render a new surface. In at least one embodiment, the video image synthesizer may be used to offload the GPU 808 to improve performance and responsiveness when the power of the GPU 808 is turned on and it is performing active 3D rendering.
[0096] In at least one embodiment, one or more of the SoCs 804 may further include a camera serial interface of a mobile industry processor interface ("MIPI") for receiving inputs from video and cameras, a high-speed interface, and / or a video input block that may be used for the input functions of cameras and related pixels. In at least one embodiment, one or more of the SoCs 804 may further include an input / output controller, which may be controlled by software and may be used to receive I / O signals that are not bound to a specific role.
[0097] In at least one embodiment, one or more of the SoCs 804 may further include a wide range of peripheral device interfaces that enable communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. The SoC 804 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet®), sensors (e.g., LIDAR sensor 864, RADAR sensor 860, etc. that may be connected via Ethernet®), data from bus 802 (e.g., speed, steering wheel position, etc. of vehicle 800), data from GNSS sensor 858 (e.g., connected via Ethernet® or CAN bus), etc. In at least one embodiment, one or more of the SoCs 804 may further include a dedicated high-performance large-capacity storage controller, which may include its own DMA engine and may be used to free the CPU 806 from routine data management tasks.
[0098] In at least one embodiment, the SoC 804 may be an end-to-end platform with a flexible architecture spanning automation levels 3 - 5, thereby providing a comprehensive functional safety architecture that leverages computer vision and ADAS techniques for diversity and redundancy and efficiently utilizes them, and providing a platform for a flexible and reliable driving software stack, along with deep learning tools. In at least one embodiment, the SoC 804 may be faster, more reliable, and more energy - and space - efficient than conventional systems. For example, in at least one embodiment, the accelerator 814, when combined with the CPU 806, GPU 808, and data store 816, may provide a fast and efficient platform for level 3 - 5 autonomous vehicles.
[0099] In at least one embodiment, the computer vision algorithm may be executed on a CPU, and this algorithm may be configured using a high-level programming language such as the C programming language to execute various processing algorithms over various visual data. However, in at least one embodiment, the CPU often cannot meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs cannot execute in real time complex object detection algorithms used in ADAS applications within a vehicle and in realistic level 3-5 autonomous vehicles.
[0100] The embodiments described herein enable multiple neural networks to be implemented simultaneously and / or sequentially and combine the results to enable level 3-5 autonomous driving functions. For example, in at least one embodiment, the CNN running on the DLA or an individual GPU (e.g., GPU 820) may include text and word recognition and enable a supercomputer to read and understand traffic signs including signs that the neural network has not been specifically trained for. In at least one embodiment, the DLA may further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass that semantic understanding to a path planning module running on the CPU complex.
[0101] In at least one embodiment, for level 3, 4, or 5 operation, multiple neural networks may be executed simultaneously. For example, in at least one embodiment, a warning sign that reads "Caution: Frozen when flashing" in conjunction with the electro-optical may be interpreted separately or collectively by several neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), the text "Frozen when flashing" may be interpreted by a second introduced neural network, and when a flashing light is detected, this neural network notifies the vehicle's (preferably running on the CPU complex) route planning software that a frozen state exists. In at least one embodiment, the flashing light may be identified by operating a third introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may be executed simultaneously, such as within the DLA and / or on the GPU 808.
[0102] In at least one embodiment, a CNN for face recognition and vehicle owner identification may use data from a camera sensor to identify the presence of an approved driver and / or owner of the vehicle 800. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle, turn on the lights when the owner approaches the driver's door, and disable the vehicle in security mode when the owner leaves the vehicle. In this way, the SoC 804 provides security against theft and / or carjacking.
[0103] In at least one embodiment, the CNN for detecting and identifying emergency vehicles may detect and identify the sirens of emergency vehicles using data from the microphone 896. In at least one embodiment, the SoC 804 uses the CNN to classify environmental and urban sounds, as well as to classify visual data. In at least one embodiment, the CNN executed on the DLA is trained to identify the relative speed at which an emergency vehicle is approaching (e.g., by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area where the vehicle is operating, as identified by the GNSS sensor 858. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and in the case of the United States, the CNN attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine is used to reduce the speed of the vehicle, move it to the side of the road, stop the vehicle, and / or use the ultrasonic sensor 862 in combination to idle the vehicle until the emergency vehicle has passed.
[0104] In at least one embodiment, the vehicle 800 may include a CPU 818 (e.g., an individual CPU or dCPU), which may be coupled to the SoC 804 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU 818 may include, for example, an X86 processor. The CPU 818 may be used to perform any of a variety of functions, including, for example, mediating potentially inconsistent results between the ADAS sensors and the SoC 804, and / or monitoring the state and health of the controller 836 and / or the in-vehicle infotainment system (the "infotainment SoC") 830 on the chip.
[0105] In at least one embodiment, vehicle 800 may include a GPU 820 (e.g., a discrete GPU, i.e., a dGPU), which may be coupled to the SoC 804 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, the GPU 820 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from the sensors of vehicle 800.
[0106] In at least one embodiment, vehicle 800 may further include a network interface 824, which may include, without limitation, a wireless antenna 826 (e.g., a cellular antenna, a Bluetooth antenna, or one or more wireless antennas 826 for different communication protocols). In at least one embodiment, the network interface 824 may be used to wirelessly connect to the cloud (e.g., servers and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device) through the Internet. In at least one embodiment, a direct link may be established between vehicle 800 and other vehicles for communicating with them, and / or an indirect link (e.g., across a network and through the Internet) may be established. In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide vehicle 800 with information about vehicles in the vicinity of vehicle 800 (e.g., vehicles in front of, to the side of, and / or behind vehicle 800). In at least one embodiment, the foregoing functions may be part of the cooperative adaptive cruise control function of vehicle 800.
[0107] In at least one embodiment, network interface 824 may include a system-on-chip (SoC) that provides modulation and demodulation functions to enable the controller 836 to communicate via a wireless network. In at least one embodiment, network interface 824 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion can be performed by well-known processes and / or using a superheterodyne process. In at least one embodiment, the radio frequency front end functions may be provided by a separate chip. In at least one embodiment, the network interface may include wireless capabilities for communicating via LTE, WCDMA (registered trademark), UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0108] In at least one embodiment, vehicle 800 may further include a data store 828, which may include off-chip (e.g., not on SoC 804) storage without limitation. In at least one embodiment, data store 828 may include one or more storage elements including, without limitation, RAM, SRAM, dynamic random access memory ("DRAM"), video random-access memory ("VRAM"), flash, hard disk, and / or other components and / or devices that may store at least one bit of data.
[0109] In at least one embodiment, vehicle 800 may further include a GNSS sensor 858 (e.g., a GPS and / or assisted GPS sensor) that aids in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 858 may be used, including, for example and without limitation, a GPS that uses a USB connector with a bridge from Ethernet (registered trademark) to serial (e.g., RS-232).
[0110] In at least one embodiment, vehicle 800 may further include a RADAR sensor 860. The RADAR sensor 860 may be used by vehicle 800 to detect vehicles at long range even in darkness and / or adverse weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. The RADAR sensor 860 may use the CAN and / or bus 802 for control (e.g., to transmit data generated by the RADAR sensor 860) and to access object tracking data, and in some examples access Ethernet (registered trademark) to access raw data. In at least one embodiment, a variety of types of RADAR sensors may be used. For example and without limitation, the RADAR sensor 860 may be suitable for use in front, rear, and side RADAR. In at least one embodiment, one or more of the RADAR sensors 860 are pulse Doppler RADAR sensors.
[0111] In at least one embodiment, the RADAR sensor 860 may include different configurations, such as a narrow field of view for long distances, a wide field of view for short distances, and a short distance that covers the sides. In at least one embodiment, the long-range RADAR may be used for an adaptive cruise control function. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within a range of 250 m, realized by two or more independent scans. In at least one embodiment, the RADAR sensor 860 may be made easier to distinguish between static and moving objects and may be used by the ADAS system 838 for emergency braking assistance and forward collision warning. The sensors 860 included in the long-range RADAR system may include, without limitation, a plurality (for example, six or more) of fixed RADAR antennas, as well as a monostatic multi-mode RADAR having high-speed CAN and FlexRay interfaces. In at least one embodiment, when there are six antennas, the four central antennas may generate a focused beam pattern designed to record the surroundings of the vehicle 800 at a higher speed with minimal interference from traffic in adjacent lanes. In at least one embodiment, the other two antennas may be capable of expanding the field of view to quickly detect vehicles entering or exiting the lane of the vehicle 800.
[0112] In at least one embodiment, the mid-range RADAR system may include, by way of example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, the short-range RADAR system may include any number of RADAR sensors 860 designed to be installed at both ends of the rear bumper, without limitation. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may create two beams that constantly monitor the rear and the blind spots adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used by the ADAS system 838 for blind spot detection and / or lane change assistance.
[0113] In at least one embodiment, vehicle 800 may further include an ultrasonic sensor 862. The ultrasonic sensor 862 may be positioned in front of, behind, and / or to the side of the vehicle 800 and may be used for parking assistance and / or for creating and updating an occupancy grid. In at least one embodiment, a variety of ultrasonic sensors 862 may be used, and different ultrasonic sensors 862 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor 862 may operate at a functional safety level of ASIL B.
[0114] In at least one embodiment, vehicle 800 may include a LIDAR sensor 864. The LIDAR sensor 864 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor 864 may be at a functional safety level of ASIL B. In at least one embodiment, vehicle 800 may include a plurality of LIDAR sensors 864 (e.g., two, four, six, etc.), and these sensors may use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).
[0115] In at least one embodiment, the LIDAR sensor 864 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 864 may, for example, have a claimed range of about 100 m, an accuracy of 2 cm to 3 cm, and support a 100 Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors 864 may be used. In such an embodiment, the LIDAR sensor 864 may be incorporated in the front, rear, sides, and / or corners of the vehicle 800 and may be implemented as a small device. In at least one embodiment, the LIDAR sensor 864 of such an embodiment may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, in a range of 200 m. In at least one embodiment, the LIDAR sensor 864 mounted in the front may be configured to provide a horizontal field of view of 45 degrees to 135 degrees.
[0116] In at least one embodiment, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmission source to irradiate the area around vehicle 800 up to approximately 200 m at most. In at least one embodiment, the flash LIDAR unit includes, without limitation, a receptor that records the transit time of the laser pulses and the reflected light at each pixel, which corresponds to the range from vehicle 800 to an object. In at least one embodiment, flash LIDAR makes it possible to generate a very accurate and distortion-free surrounding image for each laser flash. In at least one embodiment, four flash LIDAR sensors may be introduced, one on each side of vehicle 800. In at least one embodiment, the 3D flash LIDAR system includes, without limitation, a LIDAR camera of a semiconductor 3D staring array (such as a non-scanning LIDAR device) without moving parts other than a fan. In at least one embodiment, the flash LIDAR device may use Class I (eye-safe) laser pulses of 5 nanoseconds per frame and capture the reflected laser light in the form of 3D range point clouds and co-registered intensity data.
[0117] In at least one embodiment, the vehicle may further include an IMU sensor 866. In at least one embodiment, the IMU sensor 866 may be disposed at the center of the rear axle of vehicle 800 in at least one embodiment. In at least one embodiment, the IMU sensor 866 may include, without limitation, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other types of sensors. In at least one embodiment for a six-axis application, the IMU sensor 866 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment for a nine-axis application, the IMU sensor 866 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0118] In at least one embodiment, the IMU sensor 866 may be implemented as a small high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical systems (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, with the IMU sensor 866, the vehicle 800 may be able to estimate the orientation without the need for input from a magnetic sensor by directly observing the speed change and correlating it from GPS to the IMU sensor 866. In at least one embodiment, the IMU sensor 866 and the GNSS sensor 858 may be combined into a single integrated unit.
[0119] In at least one embodiment, the vehicle 800 may include a microphone 896 disposed within and / or around the vehicle 800. In at least one embodiment, the microphone 896 may be used, among other things, for the detection and identification of emergency vehicles.
[0120] In at least one embodiment, vehicle 800 may further include any number of camera types, including stereo camera 868, wide-angle camera 870, infrared camera 872, surround camera 874, long-range camera 898, mid-range camera 876, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 800. In at least one embodiment, the type of camera used may vary depending on vehicle 800. In at least one embodiment, any combination of camera types may be used to provide the required field of view around vehicle 800. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, vehicle 800 may include six cameras, seven cameras, ten cameras, twelve cameras, or another number of cameras. The cameras may support, by way of non-limiting example, Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet®. In at least one embodiment, the cameras are each described in further detail herein with respect to FIGS. 8A and 8B.
[0121] In at least one embodiment, vehicle 800 may further include vibration sensor 842. Vibration sensor 842 may measure the vibration of components of vehicle 800, such as an axle. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 842 are used, the difference in vibration may be used to determine the amount of friction or slip on the road surface (e.g., if there is a vibration difference between a power-driven axle and a freely rotating axle).
[0122] In at least one embodiment, vehicle 800 may include an ADAS system 838. The ADAS system 838 may include, without limitation, a SoC in some examples. In at least one embodiment, the ADAS system 838 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward crash warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions.
[0123] In at least one embodiment, the ACC system may use a RADAR sensor 860, a LIDAR sensor 864, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to the vehicle immediately in front of vehicle 800 and automatically adjusts the speed of vehicle 800 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 800 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications such as LC and CW.
[0124] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a wireless link or indirectly through a network connection (e.g., through the Internet) via network interface 824 and / or wireless antenna 826. In at least one embodiment, a direct link may be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link may be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, the concept of V2V communication provides information about the immediately preceding vehicle (e.g., a vehicle in the same lane immediately in front of vehicle 800), and the concept of I2V communication provides information about traffic further ahead. In at least one embodiment, the CACC system may include either or both of the I2V and V2V information sources. Given information about the vehicle ahead of vehicle 800, in at least one embodiment, the CACC system may further enhance reliability, make traffic flow smoother, and have the potential to reduce traffic jams on the road.
[0125] In at least one embodiment, the FCW system is designed to advise the driver of a hazard, whereby the driver may take corrective action. In at least one embodiment, the FCW system uses a front camera and / or RADAR sensor 860, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibration, and / or quick brake pulses.
[0126] In at least one embodiment, the AEB system may detect an imminent frontal collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the predicted collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-crash braking.
[0127] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibrations of the steering wheel or seat, to advise the driver when vehicle 800 crosses a lane marker. In at least one embodiment, the LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variant of the LDW system. The LKA system provides steering input or brake control to correct vehicle 800 if vehicle 800 begins to drift out of a lane.
[0128] In at least one embodiment, the BSW system detects a vehicle in a blind spot of a motor vehicle and warns the driver. In at least one embodiment, the BSW system may provide visual, audible, and / or tactile alerts to indicate that a merge or lane change is not safe. In at least one embodiment, the BSW system may provide an additional warning when the driver uses a turn indicator. In at least one embodiment, the BSW system may use a rear camera and / or a RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, and these dedicated processors, DSPs, FPGAs, and / or ASICs are electrically coupled to feedback to the driver, such as a display, a speaker, and / or a vibration component.
[0129] In at least one embodiment, the RCTW system may provide visual, audible, and / or tactile notifications when an object is detected outside the range of a rear camera when the vehicle 800 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear RADAR sensors 860, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, such as a display, a speaker, and / or a vibration component.
[0130] In at least one embodiment, conventional ADAS systems may sometimes produce false detection results, which can be annoying and distracting to the driver, but usually are not a big deal. This is because conventional ADAS systems are designed to allow the driver to determine whether there is truly a safety-critical situation and how to appropriately respond to it. In at least one embodiment, if the results are inconsistent, the vehicle 800 itself determines whether to follow the results from the primary computer (e.g., the first controller 836) or the results from the secondary computer (e.g., the second controller 836). For example, in at least one embodiment, the ADAS system 838 may be a backup and / or secondary computer for resisting perception information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may execute various software for redundancy on the hardware components to detect perception errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 838 may be provided to the monitoring MCU. In at least one embodiment, when the output from the primary computer conflicts with the output from the secondary computer, the monitoring MCU determines how to reconcile the conflict to ensure safe operation.
[0131] In at least one embodiment, the primary computer may be configured to provide a reliability score indicating the reliability of the selected result of the primary computer to the monitoring MCU. In at least one embodiment, if the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, if the reliability score does not meet the threshold and the primary computer and the secondary computer show different results (e.g., conflict), the monitoring MCU may mediate between the computers to determine an appropriate result.
[0132] In at least one embodiment, a neural network trained and configured to determine, at least in part based on outputs from the primary computer and the secondary computer, conditions under which the secondary computer provides a false alarm may be configured to be executed by the monitoring MCU. In at least one embodiment, the neural network of the monitoring MCU may learn when the output of the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metal object, such as a drain grate or manhole cover, that is not actually a hazard but triggers an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when bicycles or pedestrians are present and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or GPU suitable for executing the neural network along with an associated memory. In at least one embodiment, the monitoring MCU may comprise components of the SoC 804 and / or be included as a component thereof.
[0133] In at least one embodiment, the ADAS system 838 may include a secondary computer that implements ADAS functions using conventional rules of computer vision. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and the presence of the neural network in the monitoring MCU may improve reliability, safety, and performance. For example, in at least one embodiment, due to various implementations and intentional non-identities, the overall error tolerance of the system is increased, particularly with respect to errors caused by the functions of software (or the software-hardware interface). For example, in at least one embodiment, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU may have a higher level of confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.
[0134] In at least one embodiment, the output of the ADAS system 838 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 838 indicates a forward collision warning due to an object immediately ahead, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have a trained, and thus unique, neural network that reduces the risk of false detection, as described herein.
[0135] In at least one embodiment, vehicle 800 may further include an infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI)). Although the infotainment system 830 is illustrated and described as an SoC, in at least one embodiment, it may not be an SoC and may include two or more individual components without limitation. In at least one embodiment, the infotainment SoC 830 may include, without limitation, a combination of hardware and software, and this combination may be used to provide the vehicle 800 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connection (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, wireless data system, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door opening and closing, air filter information, etc.). For example, the infotainment SoC 830 may include a radio, a disk player, a navigation system, a video player, USB and Bluetooth connections, a car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, a heads-up display ("HUD"), an HMI display 834, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 830 may further be used to provide the vehicle user with information such as information from the ADAS system 838, autonomous driving information such as vehicle operation plans, trajectories, etc., ambient environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information (e.g., visual and / or auditory).
[0136] In at least one embodiment, the infotainment SoC 830 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 830 may communicate with other devices, systems, and / or components of the vehicle 800 via a bus 802 (e.g., a CAN bus, Ethernet®, etc.). In at least one embodiment, the infotainment SoC 830 may be coupled to a monitoring MCU such that, when the primary controller 836 (e.g., the primary and / or backup computer of the vehicle 800) fails, the GPU of the infotainment system may perform some self-driving functions. In at least one embodiment, the infotainment SoC 830 may put the vehicle 800 into a driver-safe stop mode as described herein.
[0137] In at least one embodiment, the vehicle 800 may further include an instrument cluster 832 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 832 may include, without limitation, a controller and / or a supercomputer (e.g., an individual controller or supercomputer). In at least one embodiment, the instrument cluster 832 may include any number and combination of instrument sets, without limitation, a speedometer, a fuel level, a hydraulic pressure, a tachometer, an odometer, a direction indicator, a shift lever position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, an auxiliary restraint system (e.g., an airbag) information, a light control, a safety system control, a navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 830 and the instrument cluster 832. In at least one embodiment, the instrument cluster 832 may be included as part of the infotainment SoC 830, or vice versa.
[0138] In at least one embodiment, the signal received from antenna 208 of FIG. 2 is from vehicle 800 and is processed as described with respect to at least one of FIGS. 1-6 to provide vehicle 800 with information for its autonomous operation, such as weather data, navigation data, road condition data, etc., and / or may be used to provide a remote operator with the ability to remotely control vehicle 800.
[0139] FIG. 8D is a diagram of a system 876 for communication between a cloud-based server and autonomous vehicle 800 of FIG. 8A according to at least one embodiment. In at least one embodiment, system 876 may include, without limitation, server 878, network 890, and any number and type of vehicles including vehicle 800. Server 878 may include, without limitation, multiple GPUs 884(A)-884(H) (collectively referred to herein as GPU 884), PCIe switches 882(A)-882(H) (collectively referred to herein as PCIe switch 882), and / or CPUs 880(A)-880(B) (collectively referred to herein as CPU 880). The GPUs 884, CPUs 880, and PCIe switches 882 may be interconnected by high-speed interconnects such as, without limitation, NVLink interface 888 developed by NVIDIA and / or PCIe connection 886. In at least one embodiment, the GPUs 884 are connected to each other via NVLink and / or NVSwitchSoC, and the GPUs 884 and PCIe switches 882 are connected via PCIe interconnect. In at least one embodiment, eight GPUs 884, two CPUs 880, and four PCIe switches 882 are illustrated, but this is not limiting. In at least one embodiment, server 878 may each include any number of GPUs 884, CPUs 880, and / or PCIe switches 882 in any combination. For example, in at least one embodiment, server 878 may each include 8, 16, 32, and / or more than 32 GPUs 884.
[0140] In at least one embodiment, server 878 may receive, via network 890, image data representing an image indicative of an unexpected or changed road condition, such as a recently started road construction, from a vehicle. In at least one embodiment, server 878 may transmit, via network 890, neural network 892, updated neural network 892, and / or map information 894 including information regarding traffic conditions and road conditions, without limitation, to a vehicle. In at least one embodiment, the update of map information 894 may include, without limitation, updates to HD map 822, such as information regarding construction sites, holes, detours, floods, and / or other obstacles. In at least one embodiment, neural network 892, updated neural network 892, and / or map information 894 may be obtained from new training and / or experiences represented in data received from any number of vehicles within the environment, and / or may be obtained, at least in part, based on training performed at a data center (e.g., using server 878 and / or other servers).
[0141] In at least one embodiment, a machine learning model (e.g., a neural network) may be trained using server 878, at least in part, based on training data. The training data may be generated by a vehicle and / or may be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or otherwise pre-processed (e.g., if the associated neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or pre-processed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by a vehicle (e.g., transmitted to the vehicle via network 890), and / or the machine learning model may be used by server 878 to remotely monitor the vehicle.
[0142] In at least one embodiment, server 878 may receive data from a vehicle and apply the data to a state-of-the-art real-time neural network to enable real-time intelligent inference. In at least one embodiment, server 878 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 884, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 878 may include a deep learning infrastructure that uses a CPU-powered data center.
[0143] In at least one embodiment, the deep learning infrastructure of server 878 may be capable of high-speed real-time inference and may use that capability to evaluate and verify the health of the processor, software, and / or associated hardware of vehicle 800. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 800, such as a series of images and / or objects located by vehicle 800 in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify an object and compare it to the object identified by vehicle 800. If the results do not match and the deep learning infrastructure concludes that the AI of vehicle 800 is malfunctioning, server 878 may take control from vehicle 800's fail-safe computer, notify the occupants, and send a signal to vehicle 800 instructing it to complete a safe shutdown procedure.
[0144] In at least one embodiment, server 878 may include GPU 884 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT3). In at least one embodiment, by combining a server powered by a GPU with inference acceleration, real-time response can be enabled. In at least one embodiment, a server powered by a CPU, FPGA, and other processors may be used for inference, such as when performance is not as critical.
[0145] Computer system FIG. 9 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination 900 thereof, formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 900 may include components such as processor 902 that uses an execution unit that includes logic for implementing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein without limitation. In at least one embodiment, computer system 900 may include a processor such as a PENTIUM® processor family, Xeon™, Itanium®, XScale™, and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 900 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may also be used.
[0146] Embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and laptop PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that may execute one or more instructions according to at least one embodiment.
[0147] In at least one embodiment, computer system 900 may include, without limitation, a processor 902, which may include, without limitation, one or more execution units 908 that perform training and / or inference of a machine learning model according to the techniques described herein. In at least one embodiment, system 900 is a single-processor desktop or server system, but in another embodiment, system 900 may be a multi-processor system. In at least one embodiment, processor 902 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that implements a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 902 may be coupled to a processor bus 910, which may transmit data signals between processor 902 and other components within computer system 900.
[0148] In at least one embodiment, the processor 902 may include, without limitation, a level 1 ("L1") internal cache memory ("cache") 904. In at least one embodiment, the processor 902 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to the processor 902. Other embodiments may also include combinations of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, the register file 906 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.
[0149] In at least one embodiment, an execution unit 908 that includes, without limitation, logic for performing integer and floating point operations is also within the processor 902. The processor 902 may also include a microcode ("u-code") read only memory ("ROM") that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 908 may include logic for handling a packed instruction set 909. In at least one embodiment, by including the packed instruction set 909 in the instruction set of the general purpose processor 902 along with the associated circuitry for executing the instructions, operations used by many multimedia applications may be performed using the packed data of the general purpose processor 902. In one or more embodiments, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.
[0150] In at least one embodiment, execution unit 908 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 900 may include memory 920 without limitation. In at least one embodiment, memory 920 may be implemented as a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or other memory device. Memory 920 may store instructions 919 and / or data 921 represented by data signals that may be executed by processor 902.
[0151] In at least one embodiment, a system logic chip may be coupled to processor bus 910 and memory 920. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (MCH) 916, and processor 902 may communicate with MCH 916 via processor bus 910. In at least one embodiment, MCH 916 may provide a high-bandwidth memory path 918 to memory 920 for storing instructions and data, as well as for storing graphics commands, data, and textures. In at least one embodiment, MCH 916 may direct data signals between processor 902, memory 920, and other components of computer system 900, and may bridge data signals between processor bus 910, memory 920, and system I / O 922. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 916 may be coupled to memory 920 via high-bandwidth memory path 918, and graphics / video card 912 may be coupled to MCH 916 through an Accelerated Graphics Port (AGP) interconnect 914.
[0152] In at least one embodiment, computer system 900 may use a system I / O 922, which is a proprietary hub interface bus that couples MCH 916 to an I / O controller hub ("ICH") 930. In at least one embodiment, ICH 930 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 920, the chipset, and processor 902. By way of example, and without limitation, it may include an audio controller 929, a firmware hub ("flash BIOS") 928, a wireless transceiver 926, data storage 924, a legacy I / O controller 923 that includes a user input and keyboard interface, a serial expansion port such as a Universal Serial Bus ("USB"), and a network controller 934. Data storage 924 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0153] In at least one embodiment, FIG. 9 shows a system that includes interconnected hardware devices or "chips," while in other embodiments, FIG. 9 may show an exemplary system-on-chip ("SoC"). In at least one embodiment, the devices shown in FIG. 9 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 900 are interconnected using a Compute Express Link (CXL) interconnect.
[0154] In at least one embodiment, at least one component illustrated or described with respect to FIG. 9 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one of processor 902 and graphics card 912 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one of processor 902 and graphics card 912 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, processor 902 executes a kernel launch function that passes parameters to at least one kernel on graphics card 912 that layer demaps, descrambles, and derate matches data from 5G NR antennas in parallel.
[0155] FIG. 10 is a block diagram illustrating an electronic device 1000 for utilizing a processor 1010 according to at least one embodiment. In at least one embodiment, the electronic device 1000 may be, for example and without limitation, a notebook, tower server, rack server, blade server, laptop, desktop, tablet, mobile device, telephone, embedded computer, or any other suitable electronic device.
[0156] In at least one embodiment, system 1000 may include a processor 1010 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices, without limitation. In at least one embodiment, the processor 1010 is coupled using a bus or interface such as an I°C bus, a system management bus ("SMBus"), a low pin count ("LPC") bus, a serial peripheral interface ("SPI"), a high definition audio ("HDA") bus, a serial advance technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, FIG. 10 shows a system including interconnected hardware devices or "chips," although in other embodiments, FIG. 10 may show an exemplary system on a chip ("SoC"). In at least one embodiment, the devices shown in FIG. 10 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 10 are interconnected using a compute express link ("CXL") interconnect.
[0157] In at least one embodiment, FIG. 10 shows a display 1024, a touch screen 1025, a touch pad 1030, a near field communications unit (NFC) 1045, a sensor hub 1040, a thermal sensor 1046, an express chipset (EC) 1035, a trusted platform module (TPM) 1038, a BIOS / firmware / flash memory (BIOS, FW flash) 1022, a DSP 1060, a drive (SSD or HDD) 1020 such as a solid state disk (SSD) or a hard disk drive (HDD), a wireless local area network unit (WLAN) 1050, a Bluetooth unit 1052, a wireless wide area network unit (WWAN) 1056, a global positioning system (GPS) 1055, a camera (USB3.0 camera) 1054 such as a USB3.0 camera, or a low power double data rate (LPDDR) memory unit (LPDDR3) 1015 implemented, for example, in accordance with the LPDDR3 standard. These components may be implemented in any suitable manner, respectively.
[0158] In at least one embodiment, through the components described above, other components may be communicatively coupled to the processor 1010. In at least one embodiment, the accelerometer 1041, the ambient light sensor (ALS) 1042, the compass 1043, and the gyroscope 1044 may be communicatively coupled to the sensor hub 1040. In at least one embodiment, the thermal sensor 1039, the fan 1037, the keyboard 1046, and the touch pad 1030 may be communicatively coupled to the EC 1035. In at least one embodiment, the speaker 1063, the headphones 1064, and the microphone (mic) 1065 may be communicatively coupled to the audio unit (audio codec and class D amplifier) 1064, and this audio unit may be communicatively coupled to the DSP 1060. In at least one embodiment, the audio unit 1064 may include, for example and without limitation, an audio coder / decoder (codec) and a class D amplifier. In at least one embodiment, the SIM card (SIM) 1057 may be communicatively coupled to the WWAN unit 1056. In at least one embodiment, components such as the WLAN unit 1050 and the Bluetooth unit 1052, as well as the WWAN unit 1056, may be implemented in the next generation form factor (NGFF).
[0159] In at least one embodiment, at least one component illustrated or described with respect to FIG. 10 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, processor 1010 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines of processor 1010. In at least one embodiment, processor 1010 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft-demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0160] FIG. 11 shows a computer system 1100 according to at least one embodiment. In at least one embodiment, computer system 1100 is configured to implement the various processes and methods described throughout this disclosure.
[0161] In at least one embodiment, computer system 1100 includes, without limitation, at least one central processing unit ("CPU") 1102, which is connected to a communication bus 1110 implemented using any suitable protocol, such as, without limitation, PCI: Peripheral Component Interconnect ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express": peripheral component interconnect express), AGP: Accelerated Graphics Port ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1100 includes, without limitation, main memory 1104 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1104, which may take the form of random access memory ("RAM": random access memory). In at least one embodiment, a network interface subsystem ("network interface") 1122 provides an interface with other computing devices and networks for receiving data from other systems and transmitting data from computer system 1100 to other systems.
[0162] In at least one embodiment, computer system 1100 includes, without limitation in at least one embodiment, input device 1108, parallel processing system 1112, and display device 1106, and this display device can be implemented using a conventional cathode ray tube ("CRT"), liquid crystal display ("LCD"), light emitting diode ("LED"), plasma display, or other suitable display technology. In at least one embodiment, user input is received from input device 1108 such as a keyboard, mouse, touch pad, microphone, etc. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.
[0163] In at least one embodiment, a computer program in the form of machine-readable and executable code or a computer control logic algorithm is stored in main memory 1104 and / or secondary storage. When executed by one or more processors, the computer program enables the system 1100 to perform various functions according to at least one embodiment. Memory 1104, storage, and / or any other storage are possible examples of computer-readable media. In at least one embodiment, secondary storage may refer to any suitable storage device or system such as a hard disk drive and / or a removable storage drive, which represent floppy (registered trademark) disk drives, magnetic tape drives, compact disk drives, digital versatile disk (「DVD」) drives, recording devices, universal serial bus (「USB」) flash memories, etc. In at least one embodiment, the architectures and / or functions of the various above-described drawings are implemented in the context of CPU 1102, a parallel processing system 1112, an integrated circuit capable of realizing at least a part of the functions of both CPU 1102 and the parallel processing system 1112, a chipset (for example, a group of integrated circuits that function as units for performing related functions and are designed to be sold), and any suitable combination of integrated circuits.
[0164] In at least one embodiment, the architectures and / or functions of the various foregoing drawings are implemented in the context of a general purpose computer system, a circuit board system, a game console system dedicated for entertainment purposes, and an application specific system, etc. In at least one embodiment, computer system 1100 may take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless portable device), a personal digital assistant (“PDA”), a digital camera, a vehicle, a head-mounted display, a portable electronic device, a mobile phone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.
[0165] In at least one embodiment, parallel processing system 1112 includes, without limitation, a plurality of parallel processing units (“PPU”) 1114, and associated memory 1116. In at least one embodiment, PPU 1114 is connected to a host processor or other peripheral device via interconnect 1118 and switch 1120 or multiplexer. In at least one embodiment, parallel processing system 1112 distributes computational tasks across PPU 1114, which can be made parallelizable, for example, as part of the distribution of computational tasks across thread blocks of a plurality of graphics processing units (“GPU”). In at least one embodiment, the memory is shared across some or all of PPU 1114 and is accessible (e.g., for read and / or write access), but such shared memory may have a performance disadvantage over the use of local memory and registers resident in PPU 1114. In at least one embodiment, the operation of PPU 1114 is synchronized by using commands such as _syncthreads(), where all threads within a block (e.g., operating across multiple PPU 1114) reach a certain execution point of the code before proceeding.
[0166] In at least one embodiment, at least one component illustrated or described with respect to FIG. 11 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one of parallel processing system 1112 and CPU 1102 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one PPU 1114 of parallel processing system 1112 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, CPU 1102 performs a kernel startup function that passes parameters to at least one kernel on PPU 1114 that layer demaps, descrambles, and derate matches data from a 5G NR antenna in parallel.
[0167] FIG. 12 shows a computer system 1200 according to at least one embodiment. In at least one embodiment, computer system 1200 includes, without limitation, computer 1210 and USB stick 1220. In at least one embodiment, computer 1210 may include, without limitation, any number and type of processors (not shown), as well as memory (not shown). In at least one embodiment, computer 1210 includes, without limitation, servers, cloud instances, laptops, and desktop computers.
[0168] In at least one embodiment, the USB stick 1220 includes, without limitation, a processing unit 1230, a USB interface 1240, and USB interface logic 1250. In at least one embodiment, the processing unit 1230 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1230 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1230 comprises an application-specific integrated circuit (an "ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1230 is a tensor processing unit (a "TPC") optimized to perform inference operations of machine learning. In at least one embodiment, the processing core 1230 is a vision processing unit (a "VPU") optimized to perform inference operations of machine vision and machine learning.
[0169] In at least one embodiment, the USB interface 1240 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1240 is a USB3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1240 is a USB3.0 Type-A connector. In at least one embodiment, the USB interface logic 1250 may include any amount and type of logic that enables the processing unit 1230 to interface with a device (e.g., computer 1210) via the USB connector 1240.
[0170] In at least one embodiment, at least one component illustrated or described with respect to FIG. 12 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, computer 1210 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, computer 1210 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0171] FIG. 13A shows an exemplary architecture in which a plurality of GPUs 1310-1313 are communicatively coupled to a plurality of multi-core processors 1305-1306 through high-speed links 1340-1343 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, high-speed links 1340-1343 support a communication throughput of 4 GB / second, 30 GB / second, 80 GB / second, or more. Various interconnect protocols may be used, including but not limited to PCIe4.0 or 5.0, and NVLink2.0.
[0172] In addition, in one embodiment, two or more of GPUs 1310-1313 are interconnected through high-speed links 1329-1330, which may be implemented using the same or different protocols / links as those used for high-speed links 1340-1343. Similarly, two or more of multi-core processors 1305-1306 may be connected via high-speed link 1328, which may be a symmetric multi-processor (SMP) bus operating at 20 GB / second, 30 GB / second, 120 GB / second, or more. Alternatively, all communication between the various system components shown in FIG. 13A may be realized using the same protocol / link (e.g., through a common interconnect fabric).
[0173] In one embodiment, each of the multi-core processors 1305-1306 is communicatively coupled to the processor memories 1301-1302 via memory interconnects 1326-1327 respectively, and each of the GPUs 1310-1313 is communicatively coupled to the GPU memories 1320-1323 through GPU memory interconnects 1350-1353 respectively. The memory interconnects 1326-1327 and 1350-1353 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memories 1301-1302 and the GPU memories 1320-1323 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of the processor memories 1301-1302 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).
[0174] As described herein, the various processors 1305-1306 and GPUs 1310-1313 may each be physically coupled to specific memories 1301-1302, 1320-1323, but an integrated memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1301-1302 may have a 64 GB system memory address space, and each of the GPU memories 1320-1323 may have a 32 GB system memory address space (in this example, resulting in a total of 256 GB of addressable memory).
[0175] FIG. 13B shows further details of the interconnection between a multi-core processor 1307 and a graphics acceleration module 1346 according to one exemplary embodiment. The graphics acceleration module 1346 may include one or more GPU chips integrated on a line card coupled to the processor 1307 via a high-speed link 1340. Alternatively, the graphics acceleration module 1346 may be integrated in the same package or chip as the processor 1307.
[0176] In at least one embodiment, the illustrated processor 1307 includes a plurality of cores 1360A - 1360D, each core having a translation lookaside buffer 1361A - 1361D and one or more caches 1362A - 1362D. In at least one embodiment, the cores 1360A - 1360D may include various other components (not shown) for executing instructions and processing data. The caches 1362A - 1362D may comprise level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 1356 may be included in the caches 1362A - 1362D and shared by a set of the cores 1360A - 1360D. For example, one embodiment of the processor 1307 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. The processor 1307 and the graphics acceleration module 1346 are connected to a system memory 1314, which may include the processor memories 1301 - 1302 of FIG. 13A.
[0177] For the data and instructions stored in various caches 1362A - 1362D, 1356, and system memory 1314, coherence is maintained through inter - core communication via coherence bus 1364. For example, each cache may have associated cache coherence logic / circuitry to communicate via coherence bus 1364 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via coherence bus 1364 to monitor cache accesses.
[0178] In one embodiment, proxy circuit 1325 communicatively couples graphics acceleration module 1346 to coherence bus 1364, enabling graphics acceleration module 1346 to participate in the cache coherence protocol as a peer of cores 1360A - 1360D. In particular, interface 1335 provides a connection to proxy circuit 1325 through high - speed link 1340 (e.g., PCIe bus, NVLink, etc.), and interface 1337 connects graphics acceleration module 1346 to link 1340.
[0179] In one implementation, the accelerator integration circuit 1336 provides services for cache management, memory access, context management, and interrupt management instead of the multiple graphics processing engines 1331, 1332, N of the graphics acceleration module 1346. Each of the graphics processing engines 1331, 1332, N may include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1331, 1332, N may include different types of graphics processing engines, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine, within the GPU. In at least one embodiment, the graphics acceleration module 1346 may be a GPU having multiple graphics processing engines 1331 to 1332, N, or the graphics processing engines 1331 to 1332, N may be individual GPUs integrated in a common package, line card, or chip.
[0180] In one embodiment, the accelerator integration circuit 1336 includes a memory management unit (MMU) 1339 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1314. The MMU 1339 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 1338 stores commands and data for efficient access by the graphics processing engines 1331-1332, N. In one embodiment, the data stored in the cache 1338 and the graphics memories 1333-1334, M are kept coherent with the core caches 1362A-1362D, 1356, and the system memory 1314. As described above, this may be achieved via the proxy circuit 1325 instead of the cache 1338 and the memories 1333-1334, M (e.g., sending updates regarding cache line modifications / accesses in the processor caches 1362A-1362D, 1356 to the cache 1338 and receiving updates from the cache 1338).
[0181] The set of registers 1345 stores context data for the threads executed by the graphics processing engines 1331 - 1332, N, and the context management circuit 1348 manages the thread contexts. For example, the context management circuit 1348 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1348 may store the current register values in a specified area of memory (identified, for example, by a context pointer). Then, the register values may be restored when returning to the context. In one embodiment, the interrupt management circuit 1347 receives and processes interrupts received from system devices.
[0182] In one implementation, the virtual / effective addresses from the graphics processing engine 1331 are translated to real / physical addresses of the system memory 1314 by the MMU 1339. One example of the accelerator integration circuit 1336 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1346, and / or other accelerator devices. The graphics accelerator module 1346 may be dedicated to a single application executed on the processor 1307, or may be shared among multiple applications. In one embodiment, there is a virtualized graphics execution environment where the resources of the graphics processing engines 1331 - 1332, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and the priorities associated with the VMs and / or applications.
[0183] In at least one embodiment, the accelerator integration circuit 1336 functions as a bridge to the system for the graphics acceleration module 1346 and provides address translation and cache services for system memory. Additionally, the accelerator integration circuit 1336 may provide a virtualization facility for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1331-1332.
[0184] Since the hardware resources of the graphics processing engines 1331-1332, N are explicitly mapped to the physical address space seen by the host processor 1307, any host processor can directly address these resources using the effective address value. In one embodiment, one function of the accelerator integration circuit 1336 is to physically separate the graphics processing engines 1331-1332, N so that they appear as independent units to the system.
[0185] In at least one embodiment, one or more graphics memories 1333-1334, M are each coupled to a respective one of the graphics processing engines 1331-1332, N. The graphics memories 1333-1334, M store the instructions and data processed by the respective graphics processing engines 1331-1332, N. The graphics memories 1333-1334, M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.
[0186] In one embodiment, to reduce data traffic through link 1340, a bias technique is used to ensure that the data stored in graphics memories 1333-1334, M is data that will be most frequently used by graphics processing engines 1331-1332, N and preferably not used (or at least not frequently used) by cores 1360A-1360D. Similarly, the bias mechanism attempts to keep data that the cores need (and thus preferably that graphics processing engines 1331-1332, N do not need) within caches 1362A-1362D, 1356 of the cores and system memory 1314.
[0187] FIG. 13C shows another exemplary embodiment in which the accelerator integration circuit 1336 is integrated within the processor 1307. In this embodiment, graphics processing engines 1331-1332, N communicate directly with the accelerator integration circuit 1336 through high-speed link 1340 via interface 1337 and interface 1335 (again, any form of bus or interface protocol may be utilized). The accelerator integration circuit 1336 may perform the same operations as described with respect to FIG. 13B, but potentially operate at a higher throughput given its proximity to coherence bus 1364 and caches 1362A-1362D, 1356. One embodiment supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1336 and a programming model controlled by the graphics acceleration module 1346.
[0188] In at least one embodiment, the graphics processing engines 1331-1332, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can concentrate other application requirements on the graphics processing engines 1331-1332, N to provide virtualization within a VM / partition.
[0189] In at least one embodiment, the graphics processing engines 1331-1332, N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 1331-1332, N to enable access by each operating system. In a single partition system without a hypervisor, the graphics processing engines 1331-1332, N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1331-1332, N to provide access to each process or application.
[0190] In at least one embodiment, the graphics acceleration module 1346 or individual graphics processing engines 1331-1332, N select process elements using a process handle. In one embodiment, the process elements are stored in the system memory 1314 and can be addressed using the translation technique from the effective address to the physical address described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1331-1332, N (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.
[0191] Figure 13D shows an exemplary accelerator integration slice 1390. As used herein, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuit 1336. The application effective address space 1382 within the system memory 1314 stores process elements 1383. In one embodiment, the process elements 1383 are stored in response to GPU calls 1381 from an application 1380 executing on the processor 1307. The process elements 1383 encompass the process state for the corresponding application 1380. A work descriptor (WD) 1384 included in the process element 1383 can be a single job requested by the application or may include a pointer to a queue of jobs. In at least one embodiment, the WD 1384 is a pointer to a job request queue in the application's address space 1382.
[0192] The graphics acceleration module 1346 and / or the individual graphics processing engines 1331 - 1332, N can be shared by all or a subset of the processes within the system. In at least one embodiment, infrastructure may be included to set the process state and send the WD 1384 to the graphics acceleration module 1346 to initiate a job in a virtualized environment.
[0193] In at least one embodiment, a dedicated process programming model is implementation - specific. In this model, a single process owns the graphics acceleration module 1346 or an individual graphics processing engine 1331. Since the graphics acceleration module 1346 is owned by a single process, when the graphics acceleration module 1346 is assigned, the hypervisor initializes the accelerator integration circuit 1336 for the owning partition, and the operating system initializes the accelerator integration circuit 1336 for the owning process.
[0194] During operation, the WD fetch unit 1391 within the accelerator integration slice 1390 fetches the next WD 1384, which includes instructions for work to be performed by one or more graphics processing engines of the graphics acceleration module 1346. As shown, the data from the WD 1384 is stored in the register 1345 and may be used by the MMU 1339, the interrupt management circuit 1347, and / or the context management circuit 1348. For example, one embodiment of the MMU 1339 includes a segment / page walk circuit for accessing the segment / page table 1386 within the OS virtual address space 1385. The interrupt management circuit 1347 may process the interrupt event 1392 received from the graphics acceleration module 1346. When performing a graphics operation, the effective address 1393 generated by the graphics processing engines 1331 - 1332, N is translated to a physical address by the MMU 1339.
[0195] In one embodiment, the same set of registers 1345 is replicated for each of the graphics processing engines 1331 - 1332, N, and / or the graphics acceleration module 1346 and may be initialized by the hypervisor or the operating system. These replicated registers may each be included in the accelerator integration slice 1390. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0196] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0197] In one embodiment, each WD 1384 is specific to a particular graphics acceleration module 1346 and / or graphics processing engines 1331 - 1332, N. The WD 1384 can either contain all the information required for the graphics processing engines 1331 - 1332, N to perform their work, or be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0198] Figure 13E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor physical address space 1398 in which a process element list 1399 is stored. The hypervisor physical address space 1398 is accessible via a hypervisor 1396 that virtualizes the graphics acceleration module engine of the operating system 1395.
[0199] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1346. There are two programming models, time - slice sharing and graphics - directed shared, in which the graphics acceleration module 1346 is shared by multiple processes and partitions.
[0200] In this model, the system hypervisor 1396 owns the graphics acceleration module 1346 and makes its functions available to all operating systems 1395. To support the virtualization by the system hypervisor 1396, the graphics acceleration module 1346 may comply with the following: 1) The job requests of the application must be autonomous (i.e., there is no need to maintain the state between jobs), or the graphics acceleration module 1346 must provide a mechanism for saving and restoring the context. 2) The job requests of the application are guaranteed by the graphics acceleration module 1346 to complete within a specified amount of time, including any translation errors, or the graphics acceleration module 1346 provides a function to preempt the job processing. 3) When the graphics acceleration module 1346 operates in a specified shared programming model, fairness must be guaranteed among processes.
[0201] In at least one embodiment, application 1380 needs to make a system call to operating system 1395 with the type of graphics acceleration module 1346, a work descriptor (WD), a permission mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1346 describes the acceleration function targeted by the system call. In at least one embodiment, the type of graphics acceleration module 1346 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1346 and can be in the form of a command for graphics acceleration module 1346, a virtual address pointer pointing to a user-defined structure, a virtual address pointer pointing to a command queue, or any other data structure for describing the work performed by graphics acceleration module 1346. In one embodiment, the AMR value is the AMR state for use by the current process. In at least one embodiment, the value passed to the operating system is the same as the application that sets the AMR. If the implementation of accelerator integration circuit 1336 and graphics acceleration module 1346 does not support the user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to the hypervisor call. Hypervisor 1396 may optionally apply the current permission mask override register (AMOR) value and then put the AMR into process element 1383. In at least one embodiment, the CSRP is one of registers 1345 that includes the virtual address of an area within the application's address space 1382 for graphics acceleration module 1346 to save and restore the context state. This pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0202] When receiving a system call, the operating system 1395 may verify that the application 1380 is registered and has been granted the right to use the graphics acceleration module 1346. Then, the operating system 1395 makes a call to the hypervisor 1396 with the information shown in Table 3. [Table 3]
[0203] When receiving a hypervisor call, the hypervisor 1396 verifies that the operating system 1395 is registered and has been granted the right to use the graphics acceleration module 1346. Then, the hypervisor 1396 inserts the process element 1383 into the process element link list of the corresponding type of graphics acceleration module 1346. The process element may include the information shown in Table 4. [Table 4]
[0204] In at least one embodiment, the hypervisor initializes the registers 1345 of the plurality of accelerator integration slices 1390.
[0205] As shown in Fig. 13F, in at least one embodiment, an integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1301-1302 and GPU memories 1320-1323. In this implementation, operations executed on GPUs 1310-1313 utilize the same virtual / effective memory address space as accessing processor memories 1301-1302, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1301, a second portion is allocated to a second processor memory 1302, and a third portion is allocated to GPU memory 1320, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of processor memories 1301-1302 and GPU memories 1320-1323 such that any processor or GPU can access any physical memory with virtual addresses mapped to physical memory.
[0206] In one embodiment, bias / coherence management circuits 1394A-1394E in one or more of MMUs 1339A-1339E ensure cache coherence between caches of one or more host processors (e.g., 1305) and caches of GPUs 1310-1313 and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. Multiple instances of bias / coherence management circuits 1394A-1394E are shown in Fig. 13F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 1305 and / or within accelerator integration circuit 1336.
[0207] One embodiment enables mapping the GPUs' memories 1320 - 1323 as part of the system memory and making them accessible using shared virtual memory (SVM) technology, without incurring performance degradation related to full system cache coherence. In at least one embodiment, the GPUs' memories 1320 - 1323 are accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offloading. This configuration enables the host processor 1305 software to set operands and access computation results without the overhead of conventional I / O DMA data copies. Such conventional copies involve driver calls, interrupts, and memory mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPUs' memories 1320 - 1323 without cache coherence overhead may be essential to the execution time of offloaded computations. For example, with significant streaming write memory traffic, cache coherence overhead may significantly reduce the effective write bandwidth seen by the GPUs 1310 - 1313. In at least one embodiment, the efficiency of operand setting, access to results, and GPU computation may help in determining the effectiveness of GPU offloading.
[0208] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, which may be a page-granularity structure that includes one or two bits per memory page with a GPU (i.e., it may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPUs with memory 1320 - 1323, with or without a bias cache in the GPUs 1310 - 1313 (e.g., to cache frequently used / recently used entries of the bias table). Alternatively, the entire bias table may be maintained within the GPU.
[0209] In at least one embodiment, the entries of the bias table associated with each access to the GPUs with memory 1320 - 1323 are accessed prior to the actual access to the GPU memory, resulting in the following operations. First, local requests from the GPUs 1310 - 1313 to find their pages within the GPU bias are transferred directly to the corresponding GPUs with memory 1320 - 1323. Local requests from the GPUs to find their pages in the host bias are transferred to the processor 1305 (e.g., through the high-speed link described above). In one embodiment, requests from the processor 1305 to find the requested page in the host processor bias complete the request in the same manner as a normal memory read. Alternatively, requests directed to a GPU-biased page may be transferred to the GPUs 1310 - 1313. In at least one embodiment, the GPU may then migrate the page to the host processor bias if the page is not currently in use. In at least one embodiment, the bias state of a page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply a hardware-based mechanism.
[0210] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the GPU's device driver, and this device driver sends a message to the GPU (or adds a command descriptor to a queue) to change the bias state, and for some transitions, directs the GPU to perform a cache flushing operation at the host. In at least one embodiment, the cache flushing operation is used for the transition from the host processor 1305's bias to the GPU bias, but not for the opposite transition.
[0211] In one embodiment, cache coherence is maintained by the host processor 1305 temporarily rendering GPU-biased pages that cannot be cached. To access these pages, the processor 1305 may request access from the GPU 1310, and the GPU 1310 may either immediately grant access or not. Thus, to reduce communication between the processor 1305 and the GPU 1310, it is beneficial to make GPU-biased pages that are requested by the GPU but not by the host processor 1305, or vice versa.
[0212] In at least one embodiment, at least one component illustrated or described with respect to FIGS. 13A - 13F is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one GPU and / or multi - core processor illustrated or described with respect to FIGS. 13A - 13F is used to decode in parallel information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one GPU, such as 1310, 1311, 1312, and / or 1313, is used for layer demapping, descrambling, and de - rate matching of 5G NR PUSCH data that is soft - demapped for LDPC decoding, where at least one thread block is used to perform operations on code - block data elements (e.g., LLRs) in parallel. In at least one embodiment, a multi - core processor, such as multi - core processor 1305, executes a kernel launch function that passes parameters to at least one kernel on a graphics processor, such as GPU 1310, that layer demaps, descrambles, and de - rate matches data from 5G NR antennas in parallel.
[0213] FIG. 14 shows an exemplary integrated circuit and associated graphics processor that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general - purpose processor cores.
[0214] FIG. 14 is a block diagram showing an exemplary system-on-chip integrated circuit 1400 that may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the integrated circuit 1400 includes one or more application processors 1405 (e.g., CPUs), at least one graphics processor 1410, and may further include an image processor 1415 and / or a video processor 1420, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 1400 includes peripheral devices or bus logic including a USB controller 1425, a UART controller 1430, an SPI / SDIO controller 1435, and an I²S / I²C controller 1440. In at least one embodiment, the integrated circuit 1400 can include a display device 1445 coupled to one or more of a high-definition multimedia interface (HDMI™) controller 1450 and a mobile industry processor interface (MIPI) display interface 1455. In at least one embodiment, storage may be provided by a flash memory subsystem 1460 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1465 to access a SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1470.
[0215] In at least one embodiment, at least one component illustrated or described with respect to FIG. 14 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, the graphics processor 1410 is used to decode in parallel information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, the graphics processor 1410 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, the application processor 1405 performs a kernel launch function that passes parameters to at least one kernel on the graphics processor 1410 that layer demaps, descrambles, and de-rate matches data from a 5G NR antenna in parallel.
[0216] FIGS. 15A-15B illustrate an exemplary integrated circuit and associated graphics processor that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general purpose processor cores.
[0217] Figures 15A-15B are block diagrams showing exemplary graphics processors for use within a SoC, according to embodiments described herein. FIG. 15A shows an exemplary graphics processor 1510 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores, according to at least one embodiment. FIG. 15B shows a further exemplary graphics processor 1540 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 1510 of FIG. 15A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1540 of FIG. 15B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1510, 1540 can be a variation of the graphics processor 1410 of FIG. 14.
[0218] In at least one embodiment, the graphics processor 1510 includes a vertex processor 1505 and one or more fragment processors 1515A-1515N (e.g., 1515A, 1515B, 1515C, 1515D-1515N-1, and 1515N). In at least one embodiment, the graphics processor 1510 can execute different shader programs via separate logic, such that the vertex processor 1505 is optimized to execute operations for vertex shader programs, while the one or more fragment processors 1515A-1515N execute fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1505 implements the vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1515A-1515N use the primitives and vertex data generated by the vertex processor 1505 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1515A-1515N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.
[0219] In at least one embodiment, the graphics processor 1510 further includes one or more memory management units (MMUs) 1520A - 1520B, caches 1525A - 1525B, and circuit interconnects 1530A - 1530B. In at least one embodiment, one or more of the MMUs 1520A - 1520B provide a virtual - to - physical address mapping for the graphics processor 1510, which includes the vertex processor 1505 and / or the fragment processors 1515A - 1515N, and they may reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more of the caches 1525A - 1525B. In at least one embodiment, one or more of the MMUs 1520A - 1520B may be synchronized with one or more other MMUs in the system, which include one or more MMUs associated with one or more of the application processors 1405, image processors 1415, and / or video processors 1420 of FIG. 14, such that each of the processors 1405 - 1420 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more of the circuit interconnects 1530A - 1530B enable the graphics processor 1510 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0220] In at least one embodiment, the graphics processor 1540 includes one or more MMUs 1520A - 1520B, caches 1525A - 1525B, and circuit interconnects 1530A - 1530B of the graphics processor 1510 of FIG. 15A. In at least one embodiment, the graphics processor 1540 includes one or more shader cores 1555A - 1555N (e.g., 1555A, 1555B, 1555C, 1555D, 1555E, 1555F - 1555N - 1, and 1555N), which provide an integrated shader core architecture where a single core, or type, or core can execute all types of programmable shader code including vertex shader, fragment shader, and / or compute shader program code. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1540 includes a core - to - core task manager 1545 that acts as a thread dispatcher to dispatch execution threads to one or more of the shader cores 1555A - 1555N, and a tiling unit 1558 for accelerating tiling operations for tile - based rendering where the rendering operation of a scene is subdivided in the image space, for example, to utilize local - space coherence within the scene or to optimize the use of internal caches.
[0221] In at least one embodiment, at least one component illustrated or described with respect to FIGS. 15A and 15B is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics processor 1510 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 1510 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft-demapped for LDPC decoding, where at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0222] FIGS. 16A and 16B illustrate further exemplary graphics processor logic according to embodiments described herein. FIG. 16A shows a graphics core 1600 that may be included in the graphics processor 1410 of FIG. 14 in at least one embodiment and may be integrated shader cores 1555A-1555N as in FIG. 15B in at least one embodiment. FIG. 16B shows a highly parallel general-purpose graphics processing unit 1630 suitable for introduction into a multi-chip module in at least one embodiment.
[0223] In at least one embodiment, the graphics core 1600 includes a shared instruction cache 1602, a texture unit 1618, and a cache / shared memory 1620, which are common to the execution resources within the graphics core 1600. In at least one embodiment, the graphics core 1600 can include a plurality of slices 1601A - 1601N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 1600. The slices 1601A - 1601N can include support logic including local instruction caches 1604A - 1604N, thread schedulers 1606A - 1606N, thread dispatchers 1608A - 1608N, and sets of registers 1610A - 1610N. In at least one embodiment, the slices 1601A - 1601N can include a set of additional functional units (AFU 1612A - 1612N), floating point units (FPU 1614A - 1614N), integer arithmetic logic units (ALU 1616 - 1616N), address calculation units (ACU 1613A - 1613N), double precision floating point units (DPFPU 1615A - 1615N), and matrix processing units (MPU 1617A - 1617N).
[0224] In at least one embodiment, the FPUs 1614A to 1614N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPU 1615A to 1615N perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1616A to 1616N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1617A to 1617N can also be configured for mixed-precision matrix operations including half-precision floating-point and 8-bit integer operations. In at least one embodiment, the MPUs 1617 to 1617N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for the acceleration of general matrix-matrix multiplication (GEMM). In at least one embodiment, the AFUs 1612A to 1612N can perform additional logical operations not supported by a floating-point unit or an integer unit, including trigonometric operations (e.g., sine, cosine, etc.).
[0225] In at least one embodiment, at least one component illustrated or described with respect to FIG. 16A is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics processor 1600 is used to parallelly decode information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 1600 is used for layer demapping, descrambling, and rate matching of 5G NR PUSCH data that is soft-demapped for LDPC decoding, wherein at least one thread block is used to parallelly perform operations on code block data elements (e.g., LLRs).
[0226] FIG. 16B shows a general-purpose processing unit (GPGPU) 1630, which can be configured to perform high-parallel computing operations by an array of graphics processing units in at least one embodiment. In at least one embodiment, GPGPU 1630 can be directly linked to other instances of GPGPU 1630 to create a multi-GPU cluster to improve the training speed of a deep neural network. In at least one embodiment, GPGPU 1630 includes a host interface 1632 that enables connection to a host processor. In at least one embodiment, host interface 1632 is a PCI Express interface. In at least one embodiment, host interface 1632 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 1630 receives commands from a host processor and uses a global scheduler 1634 to distribute the execution threads associated with these commands to a set of compute clusters 1636A-1636H. In at least one embodiment, compute clusters 1636A-1636H share a cache memory 1638. In at least one embodiment, cache memory 1638 can act as a high-level cache for cache memory within compute clusters 1636A-1636H.
[0227] In at least one embodiment, GPGPU 1630 includes memories 1644A-1644B coupled to compute clusters 1636A-1636H via a set of memory controllers 1642A-1642B. In at least one embodiment, memories 1644A-1644B can include various types of memory devices, including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM), or graphics random access memory, including graphics double data rate (GDDR) memory.
[0228] In at least one embodiment, each of compute clusters 1636A - 1636H includes a set of graphics cores, such as graphics core 1600 of FIG. 16A, and this set of graphics cores can include multiple types of integer and floating - point logic units that can perform computational operations with various precisions, including those suitable for machine - learning computations. For example, in at least one embodiment, at least a subset of the floating - point units in each of compute clusters 1636A - 1636H can be configured to perform 16 - bit or 32 - bit floating - point operations, while another subset of the floating - point units can be configured to perform 64 - bit floating - point operations.
[0229] In at least one embodiment, multiple instances of the GPGPU 1630 can be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 1636A-1636H varies across embodiments. In at least one embodiment, multiple instances of the GPGPU 1630 communicate through the host interface 1632. In at least one embodiment, the GPGPU 1630 includes an I / O hub 1639 that couples the GPGPU 1630 to a GPU link 1640 that enables direct connection to other instances of the GPGPU 1630. In at least one embodiment, the GPU link 1640 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of the GPGPU 1630. In at least one embodiment, the GPU link 1640 is coupled to a high-speed interconnect for transmitting and receiving data to / from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1630 are located within separate data processing systems and communicate via a network device accessible via the host interface 1632. In at least one embodiment, the GPU link 1640 can be configured to connect to the host processor in addition to, or instead of, the host interface 1632.
[0230] In at least one embodiment, the GPGPU 1630 can be configured to train a neural network. In at least one embodiment, the GPGPU 1630 can be used within an inference platform. In at least one embodiment where the GPGPU 1630 is used for inference, the GPGPU may include fewer compute clusters 1636A - 1636H than when the GPGPU is used for neural network training. In at least one embodiment, the memory technology associated with memories 1644A - 1644B may be different for the inference configuration and the training configuration, and high - bandwidth memory technology is applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1630 can support inference - specific instructions. For example, in at least one embodiment, the inference configuration can support dot product instructions for one or more 8 - bit integers, which may be used during the inference operation of a pre - trained neural network.
[0231] In at least one embodiment, at least one component illustrated or described with respect to FIG. 16B is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one GPGPU 1630 is used to parallel - decode information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one GPGPU 1630 is used for layer demapping, descrambling, and rate - matching of 5G NR PUSCH data that is soft - demapped for LDPC decoding, where at least one thread block is used to perform operations on code - block data elements (e.g., LLRs) in parallel.
[0232] FIG. 17 is a block diagram showing a computing system 1700 according to at least one embodiment. In at least one embodiment, the computing system 1700 includes a processing subsystem 1701 having one or more processors 1702 and a system memory 1704 that communicate via an interconnect path that may include a memory hub 1705. In at least one embodiment, the memory hub 1705 may be a separate component within a chipset component or may be integrated within one or more processors 1702. In at least one embodiment, the memory hub 1705 is coupled to an I / O subsystem 1711 via a communication link 1706. In at least one embodiment, the I / O subsystem 1711 includes an I / O hub 1707 that enables the computing system 1700 to receive input from one or more input devices 1708. In at least one embodiment, the I / O hub 1707 can enable a display controller, which may be included in one or more processors 1702, to provide output to one or more display devices 1710A. In at least one embodiment, one or more display devices 1710A coupled to the I / O hub 1707 can include local, internal, or embedded display devices.
[0233] In at least one embodiment, the processing subsystem 1701 includes one or more parallel processors 1712 coupled to a memory hub 1705 via a bus or other communication link 1713. In at least one embodiment, the communication link 1713 may be one of a number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 1712 form a parallel or vector processing system focused on computing that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1712 form a graphics processing subsystem that can output pixels to one of one or more display devices 1710A coupled via an I / O hub 1707. In at least one embodiment, the one or more parallel processors 1712 can also include a display controller and display interface (not shown) that enable direct connection to one or more display devices 1710B.
[0234] In at least one embodiment, the system storage unit 1714 can be connected to the I / O hub 1707 to provide a storage mechanism for the computing system 1700. In at least one embodiment, an I / O switch 1716 can be used to interface with the I / O hub 1707 and other components such as a network adapter 1718 and / or a wireless network adapter 1719 that may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 1720, enabling communication with the I / O hub 1707. In at least one embodiment, the network adapter 1718 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1719 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0235] In at least one embodiment, the computing system 1700 can include other components not explicitly shown, including USB or other port connections, an optical storage drive, a video capture device, etc., which may also be connected to the I / O hub 1707. In at least one embodiment, the communication paths interconnecting the various components of FIG. 17 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces and / or protocols such as NV-Link high-speed interconnects, or interconnect protocols.
[0236] In at least one embodiment, one or more parallel processors 1712 incorporate circuitry optimized for graphics and video processing, including, for example, a video output circuit, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1712 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 1700 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1712, memory hub 1705, processor 1702, and I / O hub 1707 may be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 1700 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1700 may be integrated into a multi-chip module (MCM), and this module may be interconnected with other multi-chip modules to form a modular computing system.
[0237] In at least one embodiment, at least one component illustrated or described with respect to FIG. 17 is utilized to implement the techniques and / or functions described in relation to FIGS. 1-6. In at least one embodiment, at least one of processor 1702 and parallel processor 1712 is used to parallelly decode information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one parallel processor 1712 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to parallelly perform operations on code block data elements (e.g., LLRs). In at least one embodiment, processor 1702 executes a kernel startup function that passes parameters to at least one kernel on at least one parallel processor 1712 that parallelly performs layer demapping, descrambling, and derate matching of data from 5G NR antennas.
[0238] processor FIG. 18A shows a parallel processor 1800 according to at least one embodiment. In at least one embodiment, the various components of parallel processor 1800 may be implemented using one or more integrated circuit devices such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 1800 is a variant of one or more parallel processors 1712 shown in FIG. 17 according to an exemplary embodiment.
[0239] In at least one embodiment, parallel processor 1800 includes parallel processing units 1802. In at least one embodiment, parallel processing unit 1802 includes an I / O unit 1804 that enables communication with other devices including other instances of parallel processing unit 1802. In at least one embodiment, I / O unit 1804 may be directly connected to other devices. In at least one embodiment, I / O unit 1804 is connected to other devices through the use of a hub or switch interface such as memory hub 1705. In at least one embodiment, the connection between memory hub 1705 and I / O unit 1804 forms a communication link 1713. In at least one embodiment, I / O unit 1804 is connected to host interface 1806 and memory crossbar 1816, where host interface 1806 receives commands for performing processing operations and memory crossbar 1816 receives commands for performing memory operations.
[0240] In at least one embodiment, when host interface 1806 receives a command buffer via I / O unit 1804, host interface 1806 can direct a work operation for implementing these commands to front end 1808. In at least one embodiment, front end 1808 is coupled to scheduler 1810, and this scheduler is configured to distribute commands or other work items to processing cluster array 1812. In at least one embodiment, scheduler 1810 ensures that processing cluster array 1812 is properly configured and in an effective state before tasks are distributed to processing cluster array 1812 of processing cluster array 1812. In at least one embodiment, scheduler 1810 is implemented via firmware logic running on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 1810 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularities, enabling rapid preemption and context switching of threads running on processing array 1812. In at least one embodiment, host software can attest to the scheduling workload on processing array 1812 via one of a plurality of graphics processing doorbells. In at least one embodiment, then, the workload can be automatically distributed across processing array 1812 by scheduler 1810 logic within the microcontroller including scheduler 1810.
[0241] In at least one embodiment, the processing cluster array 1812 can include a maximum of "N" processing clusters (e.g., cluster 1814A, cluster 1814B ~ cluster 1814N). In at least one embodiment, each of the clusters 1814A - 1814N of the processing cluster array 1812 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 1810 can use various scheduling and / or work distribution algorithms to distribute work to the clusters 1814A - 1814N of the processing cluster array 1812, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 1810, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 1812. In at least one embodiment, the different clusters 1814A - 1814N of the processing cluster array 1812 can be distributed to process different types of programs or perform different types of calculations.
[0242] In at least one embodiment, the processing cluster array 1812 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 1812 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 1812 can include logic for executing processing tasks including filtering of video and / or audio data, performing modeling operations including physical operations, and performing data conversion.
[0243] In at least one embodiment, the processing cluster array 1812 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 1812 can include texture sampling logic for performing texture operations, as well as additional logic for supporting the execution of such graphics processing operations, including but not limited to mosaic logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1812 can be configured to execute graphics processing-related shader programs, such as but not limited to vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1802 can transfer data from the system memory for processing through the I / O unit 1804. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 1822) during processing and then written back to the system memory.
[0244] In at least one embodiment, when graphics processing is performed using the parallel processing unit 1802, the scheduler 1810 can be configured to divide the processing workload into tasks of approximately equal size so as to more effectively distribute the graphics processing operations to the plurality of clusters 1814A - 1814N of the processing cluster array 1812. In at least one embodiment, a portion of the processing cluster array 1812 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 1814A - 1814N can be stored in a buffer to enable transmission of the intermediate data between the clusters 1814A - 1814N for further processing.
[0245] In at least one embodiment, the processing cluster array 1812 can receive processing tasks to be executed via a scheduler 1810, which receives commands defining the processing tasks from a front end 1808. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., which program to execute). In at least one embodiment, the scheduler 1810 may be configured to fetch an index corresponding to the task or may receive the index from the front end 1808. In at least one embodiment, the front end 1808 can be configured to ensure that the processing cluster array 1812 is configured in an active state before the workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is started.
[0246] In at least one embodiment, one or more instances of the parallel processing unit 1802 can each be coupled to a parallel processor-memory 1822. In at least one embodiment, the parallel processor-memory 1822 can be accessed via a memory crossbar 1816, which can receive memory requests from the processing cluster array 1812 as well as the I / O unit 1804. In at least one embodiment, the memory crossbar 1816 can access the parallel processor-memory 1822 via a memory interface 1818. In at least one embodiment, the memory interface 1818 can include a plurality of partitioning units (e.g., partitioning unit 1820A, partitioning unit 1820B to partitioning unit 1820N), each of which can be coupled to a portion (e.g., a memory unit) of the parallel processor-memory 1822. In at least one embodiment, the number of partitioning units 1820A to 1820N is configured to be equal to the number of memory units, such that the first partitioning unit 1820A has a corresponding first memory unit 1824A, the second partitioning unit 1820B has a corresponding memory unit 1824B, and the Nth partitioning unit 1820N has a corresponding Nth memory unit 1824N. In at least one embodiment, the number of partitioning units 1820A to 1820N may not be equal to the number of memory devices.
[0247] In at least one embodiment, the memory units 1824A - 1824N can include various types of memory devices, including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory. In at least one embodiment, the memory units 1824A - 1824N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 1822, a render target, such as a frame buffer or texture map, can be stored across the memory units 1824A - 1824N, and the partition units 1820A - 1820N can be enabled to write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 1822 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.
[0248] In at least one embodiment, any one of clusters 1814A - 1814N of the processing cluster array 1812 can process data that is to be written to any one of memory units 1824A - 1824N within the parallel processor memory 1822. In at least one embodiment, the memory crossbar 1816 can be configured to transfer the output of each of clusters 1814A - 1814N to any partition unit 1820A - 1820N that can perform further processing operations on the output, or to another cluster 1814A - 1814N. In at least one embodiment, each of clusters 1814A - 1814N can communicate with the memory interface 1818 through the memory crossbar 1816 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 1816 has a connection to the memory interface 1818 for communicating with the I / O unit 1804, as well as a connection to a local instance of the parallel processor memory 1822, enabling processing units within different processing clusters 1814A - 1814N to communicate with the system memory or other memory not local to the parallel processing unit 1802. In at least one embodiment, the memory crossbar 1816 can use virtual channels to separate traffic streams between clusters 1814A - 1814N and partition units 1820A - 1820N.
[0249] In at least one embodiment, multiple instances of the parallel processing unit 1802 can be provided on a single add-in card or multiple add-in cards can be interconnected. In at least one embodiment, different instances of the parallel processing unit 1802 can be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 1802 can include a higher precision floating point unit compared to other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 1802 or the parallel processor 1800 can be implemented in various configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0250] FIG. 18B is a block diagram of a partition unit 1820 according to at least one embodiment. In at least one embodiment, the partition unit 1820 is an instance of one of the partition units 1820A-1820N of FIG. 18A. In at least one embodiment, the partition unit 1820 includes an L2 cache 1821, a frame buffer interface 1825, and a ROP: raster operations unit 1826. The L2 cache 1821 is a read / write cache configured to perform load and store operations received from the memory crossbar 1816 and the ROP 1826. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 1821 to the frame buffer interface 1825 for processing. In at least one embodiment, updates can also be sent to the frame buffer via the frame buffer interface 1825 for processing. In at least one embodiment, the frame buffer interface 1825 interfaces with one of the memory units of the parallel processor memory, such as the memory units 1824A-1824N (e.g., within the parallel processor memory 1822) of FIG. 18.
[0251] In at least one embodiment, the ROP 1826 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, the ROP 1826 then outputs the processed graphics data stored in the graphics memory. In at least one embodiment, the ROP 1826 includes compression logic that compresses depth or color data written to the memory and decompresses depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The type of compression performed by the ROP 1826 can be changed based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0252] In at least one embodiment, the ROP 1826 is included within each processing cluster (e.g., clusters 1814A-1814N of FIG. 18) rather than within the partition unit 1820. In at least one embodiment, read and write requests for pixel data rather than pixel fragment data are transmitted through the memory crossbar 1816. In at least one embodiment, the processed graphics data may be displayed on a display device such as one of the one or more display devices 1710 of FIG. 17, routed so as to be further processed by the processor 1702, or routed so as to be further processed by one of the processing entities within the parallel processor 1800 of FIG. 18A.
[0253] FIG. 18C is a block diagram of a processing cluster 1814 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 1814A-1814N of FIG. 18. In at least one embodiment, the processing cluster 1814 can be configured to execute multiple threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, a single instruction multiple data (SIMD) instruction issue technique is used to support parallel execution of multiple threads without providing a plurality of independent instruction units. In at least one embodiment, a single instruction multiple thread (SIMT) technique is used to support parallel execution of a plurality of threads that are overall synchronized using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0254] In at least one embodiment, the operation of processing cluster 1814 can be controlled via a pipeline manager 1832 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1832 receives instructions from the scheduler 1810 of FIG. 18 and manages the execution of these instructions via a graphics multiprocessor 1834 and / or a texture unit 1836. In at least one embodiment, the graphics multiprocessor 1834 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 1814. In at least one embodiment, one or more instances of the graphics multiprocessor 1834 can be included within the processing cluster 1814. In at least one embodiment, the graphics multiprocessor 1834 can process data and use a data crossbar 1840 to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, the pipeline manager 1832 can facilitate the distribution of the processed data by specifying the destination of the processed data to be distributed via the data crossbar 1840.
[0255] In at least one embodiment, each graphics multiprocessor 1834 within the processing cluster 1814 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction has completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be performed by leveraging the hardware of the same function units, and any combination of function units may exist.
[0256] In at least one embodiment, the instructions sent to processing cluster 1814 configure threads. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 1834. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 1834. In at least one embodiment, if the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycles in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 1834. In at least one embodiment, if the thread group includes more threads than the number of processing engines within graphics multiprocessor 1834, processing can be performed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 1834.
[0257] In at least one embodiment, the graphics multi-processor 1834 includes an internal cache memory that performs load and store operations. In at least one embodiment, the graphics multi-processor 1834 can forego the internal cache and use the cache memory (e.g., L1 cache 1848) within the processing cluster 1814. In at least one embodiment, each graphics multi-processor 1834 can also access the L2 cache within a partition unit (e.g., partition units 1820A - 1820N of FIG. 18), which are shared among all processing clusters 1814 and may be used to transfer data between threads. In at least one embodiment, the graphics multi-processor 1834 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 1802 may be used as global memory. In at least one embodiment, the processing cluster 1814 includes multiple instances of the graphics multi-processor 1834 that can share common instructions and data, which may be stored in the L1 cache 1848.
[0258] In at least one embodiment, each processing cluster 1814 may include an MMU 1845 (memory management unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 1845 may be within the memory interface 1818 of FIG. 18. In at least one embodiment, the MMU 1845 includes a set of page table entries (PTEs) that are used to map virtual addresses to the physical addresses of tiles (tiling is described in detail) and optionally cache line indices. In at least one embodiment, the MMU 1845 may include an address translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 1834 or L1 cache, or within the processing cluster 1814. In at least one embodiment, the physical addresses are processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, the cache line index may be used to determine whether a cache line request is a hit or a miss.
[0259] In at least one embodiment, each graphics multiprocessor 1834 is coupled to a texture unit 1836 such that the processing cluster 1814 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multiprocessor 1834 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multiprocessor 1834 outputs the processed tasks to a data crossbar 1840 to provide the processed tasks to another processing cluster 1814 for further processing, or stores the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 1816. In at least one embodiment, a pre-ROP 1842 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 1834 and direct the data to a ROP unit, which may be disposed within a partitioning unit (e.g., partitioning units 1820A - 1820N of FIG. 18) as described herein. In at least one embodiment, the pre-ROP 1842 unit can perform optimizations for color blending, organize pixel color data, and perform address translation.
[0260] In at least one embodiment, at least one component illustrated or described with respect to FIGS. 18A - 18C is utilized to implement the techniques and / or functions described in relation to FIGS. 1 - 6. In at least one embodiment, at least one parallel processor 1800 is used to parallelly decode information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one parallel processor 1800 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to parallelly perform operations on code block data elements (e.g., LLRs).
[0261] FIG. 18D shows a graphics multiprocessor 1834 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1834 is coupled to a pipeline manager 1832 of a processing cluster 1814. In at least one embodiment, the graphics multiprocessor 1834 has an execution pipeline that includes, but is not limited to, an instruction cache 1852, an instruction unit 1854, an address mapping unit 1856, a register file 1858, one or more general-purpose graphics processing unit (GPGPU) cores 1862, and one or more load / store units 1866. The GPGPU cores 1862 and the load / store units 1866 are coupled to cache memory 1872 and shared memory 1870 via a memory and cache interconnect 1868.
[0262] In at least one embodiment, the instruction cache 1852 receives a stream of instructions to be executed from the pipeline manager 1832. In at least one embodiment, the instructions are cached in the instruction cache 1852 and dispatched to be executed by the instruction unit 1854. In at least one embodiment, the instruction unit 1854 can dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within the GPGPU core 1862. In at least one embodiment, the instructions can access any of the local, shared, or global address spaces by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1856 can be used to translate an address in the unified address space to an individual memory address that the load / store unit 1866 can access.
[0263] In at least one embodiment, the register file 1858 provides a set of registers to the functional units of the graphics multiprocessor 1834. In at least one embodiment, the register file 1858 provides temporary storage for operands connected to the data paths of the functional units of the graphics multiprocessor 1834 (e.g., the GPGPU core 1862, the load / store unit 1866). In at least one embodiment, the register file 1858 is divided among each of the functional units such that each functional unit is allocated a dedicated portion of the register file 1858. In at least one embodiment, the register file 1858 is divided among different warps being executed by the graphics multiprocessor 1834.
[0264] In at least one embodiment, each GPGPU core 1862 can include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) that are used to execute instructions of the graphics multiprocessor 1834. The GPGPU cores 1862 may have the same architecture or different architectures. In at least one embodiment, a first portion of the GPGPU core 1862 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU can implement the IEEE 754-2008 standard for floating-point operations or enable variable-precision floating-point operations. In at least one embodiment, the graphics multiprocessor 1834 can further include one or more fixed-function units or special-function units that perform specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores can also include fixed or special-function logic.
[0265] In at least one embodiment, the GPGPU core 1862 includes SIMD logic that can execute a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 1862 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for the GPGPU core can be generated at compile time by a shader compiler or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logical unit.
[0266] In at least one embodiment, the memory and cache interconnect 1868 is an interconnect network that connects each functional unit of the graphics multiprocessor 1834 to the register file 1858 and the shared memory 1870. In at least one embodiment, the memory and cache interconnect 1868 is a crossbar interconnect that enables the load / store unit 1866 to implement load and store operations between the shared memory 1870 and the register file 1858. In at least one embodiment, the register file 1858 can operate at the same frequency as the GPGPU core 1862, and thus the data transfer between the GPGPU core 1862 and the register file 1858 is very low latency. In at least one embodiment, the shared memory 1870 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 1834. In at least one embodiment, the cache memory 1872 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 1836. In at least one embodiment, the shared memory 1870 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 1862 can store data programmatically in the shared memory in addition to automatically cached data stored in the cache memory 1872.
[0267] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core through a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into the same package or chip as the core and communicatively coupled to the core through an internal (i.e., inside the package or chip) processor bus / interconnect. In at least one embodiment, regardless of the method of connecting the GPU, the processor core may distribute work to the GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0268] In at least one embodiment, at least one component illustrated or described with respect to FIG. 18D is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics multiprocessor 1834 is used to concurrently decode information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one graphics multiprocessor 1834 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that has been soft demapped in preparation for LDPC decoding, wherein at least one thread block is used to concurrently perform operations on code block data elements (e.g., LLRs).
[0269] FIG. 19 shows a multi-GPU computing system 1900 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1900 can include a processor 1902 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 1906A-D via a host interface switch 1904. In at least one embodiment, the host interface switch 1904 is a PCI Express switch device that couples the processor 1902 to a PCI Express bus, through which the processor 1902 can communicate with the GPGPUs 1906A-D. The GPGPUs 1906A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 1916. In at least one embodiment, the GPU-to-GPU link 1916 is connected to each of the GPGPUs 1906A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 1916 enables direct communication between each of the GPGPUs 1906A-D without the need for the processor 1902 to communicate through the host interface bus 1904 to which it is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 1916, the host interface bus 1904 is kept available to access system memory or to communicate with other instances of the multi-GPU computing system 1900, for example, via one or more network devices. In at least one embodiment, the GPGPUs 1906A-D are connected to the processor 1902 via the host interface switch 1904, but in at least one embodiment, the processor 1902 includes direct support for the P2P GPU link 1916 and can be directly connected to the GPGPUs 1906A-D.
[0270] In at least one embodiment, at least one component illustrated or described with respect to FIG. 19 is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one GPGPU 1906 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one GPGPU 1906 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, where at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, processor 1902 executes a kernel launch function that passes parameters to at least one kernel on at least one GPGPU 1906 that layer demaps, descrambles, and derate matches data from 5G NR antennas in parallel.
[0271] FIG. 20 is a block diagram of a graphics processor 2000 according to at least one embodiment. In at least one embodiment, the graphics processor 2000 includes a ring interconnect 2002, a pipeline front end 2004, a media engine 2037, and graphics cores 2080A - 2080N. In at least one embodiment, the ring interconnect 2002 couples the graphics processor 2000 to other graphics processors or other processing units including one or more general purpose processor cores. In at least one embodiment, the graphics processor 2000 is one of a number of processors integrated within a multi - core processing system.
[0272] In at least one embodiment, the graphics processor 2000 receives a batch of commands via the ring interconnect 2002. In at least one embodiment, incoming commands are interpreted by the command streamer 2003 of the pipeline front end 2004. In at least one embodiment, the graphics processor 2000 includes scalable execution logic that performs 3D geometry processing and media processing via the graphics cores 2080A - 2080N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2003 supplies the commands to the geometry pipeline 2036. In at least one embodiment, for at least some media processing commands, the command streamer 2003 supplies the commands to the video front end 2034, which is coupled to the media engine 2037. In at least one embodiment, the media engine 2037 includes a Video Quality Engine (VQE) 2030 for post - processing of video and images, and a Multi - Format Encode / Decode (MFX) 2033 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2036 and the media engine 2037 each generate execution threads for the thread execution resources provided by at least one graphics core 2080A.
[0273] In at least one embodiment, the graphics processor 2000 includes a scalable thread execution resource featuring modular cores 2080A - 2080N (also sometimes referred to as core slices), and each modular core has a plurality of sub - cores 2050A - 550N, 2060A - 2060N (also sometimes referred to as core sub - slices). In at least one embodiment, the graphics processor 2000 can have any number of graphics cores 2080A - 2080N. In at least one embodiment, the graphics processor 2000 includes a graphics core 2080A having at least a first sub - core 2050A and a second sub - core 2060A. In at least one embodiment, the graphics processor 2000 is a low - power processor having a single sub - core (e.g., 2050A). In at least one embodiment, the graphics processor 2000 includes a plurality of graphics cores 2080A - 2080N, each of which includes a first set of sub - cores 2050A - 2050N and a second set of sub - cores 2060A - 2060N. In at least one embodiment, each sub - core of the first set of sub - cores 2050A - 2050N includes at least execution units 2052A - 2052N and a first set of media / texture samplers 2054A - 2054N. In at least one embodiment, each sub - core of the second set of sub - cores 2060A - 2060N includes at least execution units 2062A - 2062N and a second set of samplers 2064A - 2064N. In at least one embodiment, each sub - core 2050A - 2050N, 2060A - 2060N shares a set of shared resources 2070A - 2070N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0274] In at least one embodiment, at least one component illustrated or described with respect to FIG. 20 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics processor 2000 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 2000 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0275] FIG. 21 is a block diagram showing a microarchitecture of a processor 2100 that may include a logic circuit for performing instructions, according to at least one embodiment. In at least one embodiment, processor 2100 may perform instructions including x86 instructions, ARM instructions, special instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, processor 2110 may include registers for storing packed data, such as 64-bit wide MMX (trademark) registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, Calif. In at least one embodiment, MMX registers available in both integer and floating point formats may operate on packed data elements with single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as “SSEx”) technologies may hold operands of such packed data. In at least one embodiment, processor 2110 may perform instructions that accelerate machine learning or deep learning algorithms, training, or inference.
[0276] In at least one embodiment, the processor 2100 includes an in-order front end (“front end”) 2101 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2101 may include several units. In at least one embodiment, an instruction prefetcher 2126 fetches instructions from memory and supplies the instructions to an instruction decoder 2128, which decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2128 decodes the received instruction into one or more operations called “microinstructions” or “micro-operations” (also called “micro-ops” or “uops”) that the machine may execute. In at least one embodiment, the instruction decoder 2128 parses the instruction into an opcode and corresponding data, as well as control fields, such that these are used by the microarchitecture and operations according to at least one embodiment may be performed. In at least one embodiment, a trace cache 2130 may assemble the decoded uops into a program-order sequence or trace in a uop queue 2134 for execution. In at least one embodiment, when the trace cache 2130 encounters a complex instruction, a microcode ROM 2132 provides the uops necessary for completion of the operation.
[0277] In at least one embodiment, there may be instructions that can be converted into a single micro-op, or there may be instructions that require several micro-ops to complete all operations. In at least one embodiment, if more than five micro-ops are required to complete an instruction, the instruction decoder 2128 may access the microcode ROM 2132 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2128. In at least one embodiment, if a large number of micro-ops are required to complete an operation, the instruction may be stored in the microcode ROM 2132. In at least one embodiment, the trace cache 2130 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry point programmable logic array (「PLA」) to complete one or more instructions from the microcode ROM 2132 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2132 finishes sequencing the micro-ops for an instruction, the front end 2101 of the machine may resume fetching micro-ops from the trace cache 2130.
[0278] In at least one embodiment, an out-of-order execution engine (the "out-of-order engine") 2103 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth the flow of instructions and change their order, optimizing performance when instructions are scheduled to flow down a pipeline and be executed. The out-of-order execution engine 2103 includes, without limitation, an allocator / register renamer 2140, a memory uop queue 2142, an integer / floating-point uop queue 2144, a memory scheduler 2146, a fast scheduler 2102, a slow / general-purpose floating-point scheduler (the "slow / general-purpose FP scheduler") 2104, and a simple floating-point scheduler (the "simple FP scheduler") 2106. In at least one embodiment, the fast scheduler 2102, the slow / general-purpose floating-point scheduler 2104, and the simple floating-point scheduler 2106 are also collectively referred to herein as the "uop schedulers 2102, 2104, 2106". The allocator / register renamer 2140 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2140 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2140 also allocates an entry for each uop to one of two uop queues, namely the memory uop queue 2142 for memory operations and the integer / floating-point uop queue 2144 for non-memory operations, ahead of the memory scheduler 2146 and the uop schedulers 2102, 2104, 2106. In at least one embodiment, the uop schedulers 2102, 2104, 2106 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2102 of at least one embodiment may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2104 and the simple floating-point scheduler 2106 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2102, 2104, 2106 arbitrate dispatch ports to schedule uops for execution.
[0279] In at least one embodiment, the execution block b11 includes, without limitation, the integer register file / bypass network 2108, the floating-point register file / bypass network (referred to herein as the "FP register file / bypass network") 2110, the address generation units (AGUs) 2112 and 2114, the high-speed arithmetic logic units (ALUs) (referred to herein as the "high-speed ALUs") 2116 and 2118, the low-speed arithmetic logic units (referred to herein as the "low-speed ALUs") 2120, the floating-point ALUs (referred to herein as the "FPs") 2122, and the floating-point move units (referred to herein as the "FP moves") 2124. In at least one embodiment, the integer register file / bypass network 2108 and the floating-point register file / bypass network 2110 are also referred to herein as the "register files 2108, 2110". In at least one embodiment, the AGUs 2112 and 2114, the high-speed ALUs 2116 and 2118, the low-speed ALUs 2120, the floating-point ALUs 2122, and the floating-point move units 2124 are also referred to herein as the "execution units 2112, 2114, 2116, 2118, 2120, 2122, and 2124". In at least one embodiment, the execution block b11 may include any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination without limitation.
[0280] In at least one embodiment, register files 2108, 2110 may be disposed between uop schedulers 2102, 2104, 2106 and execution units 2112, 2114, 2116, 2118, 2120, 2122, and 2124. In at least one embodiment, integer register file / bypass network 2108 performs integer operations. In at least one embodiment, floating point register file / bypass network 2110 performs floating point operations. In at least one embodiment, register files 2108, 2110 may each, without limitation, include a bypass network that may bypass or transfer recently completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register files 2108, 2110 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2108 may, without limitation, include two separate register files, namely, one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating point instructions typically have operands with a width of 64 to 128 bits, floating point register file / bypass network 2110 may, without limitation, include 128-bit wide entries.
[0281] In at least one embodiment, execution units 2112, 2114, 2116, 2118, 2120, 2122, 2124 may execute instructions. In at least one embodiment, register files 2108, 2110 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2100 may include any number and combination of execution units 2112, 2114, 2116, 2118, 2120, 2122, 2124 without limitation. In at least one embodiment, floating-point ALU 2122 and floating-point shift unit 2124 may execute floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, floating-point ALU 2122 may include a floating-point divider by 64 bits each and execute division, square root, and other micro-ops without limitation. In at least one embodiment, instructions containing floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2116, 2118. In at least one embodiment, fast ALUs 2116, 2118 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, since low-speed ALU 2120 may include integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing without limitation, most complex integer operations proceed to low-speed ALU 2120. In at least one embodiment, memory load / store operations may be executed by AGUs 2112, 2114. In at least one embodiment, fast ALU 2116, fast ALU 2118, and low-speed ALU 2120 may perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2116, fast ALU 2118, and low-speed ALU 2120 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc. In at least one embodiment, floating-point ALU 2122 and floating-point shift unit 2124 may be implemented to support a wide range of operands having various bit widths.In at least one embodiment, the floating point ALU 2122 and the floating point shift unit 2124 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.
[0282] In at least one embodiment, the uop schedulers 2102, 2104, 2106 dispatch dependent operations before the parent load has completed execution. In at least one embodiment, since uops may be scheduled and executed speculatively in the processor 2100, the processor 2100 may also include logic for handling memory misses. In at least one embodiment, when a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed through a scheduler with temporarily incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, dependent operations may need to be replayed and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0283] In at least one embodiment, the term "register" may refer to a storage location of an on-board processor that may be used as part of an instruction to identify an operand, or may be accessible from outside the processor (from the perspective of the programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0284] In at least one embodiment, at least one component illustrated or described with respect to FIG. 21 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one processor 2100 is used to decode in parallel information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one processor 2100 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft-demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0285] FIG. 22 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2200 includes one or more processors 2202 and one or more graphics processors 2208, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 2202 or processor cores 2207. In at least one embodiment, system 2200 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile device, a portable device, or an embedded device.
[0286] In at least one embodiment, system 2200 may include, or be incorporated within, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 2200 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 2200 may also include, be coupled to, or be integrated within wearable devices such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 2200 is a television or set-top box device having one or more processors 2202 and a graphical interface generated by one or more graphics processors 2208.
[0287] In at least one embodiment, each of one or more processors 2202 includes one or more processor cores 2207 that process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 2207 is configured to process a particular instruction set 2209. In at least one embodiment, the instruction set 2209 may facilitate computing via a complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, each processor core 2207 may process a different instruction set 2209, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, processor core 2207 may also include other processing devices such as a digital signal processor (DSP).
[0288] In at least one embodiment, processor 2202 includes a cache memory 2204. In at least one embodiment, processor 2202 can have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of processor 2202. In at least one embodiment, processor 2202 may also use an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared among processor cores 2207 using known cache coherence techniques. In at least one embodiment, a register file 2206 is further included in processor 2202, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 2206 may include general-purpose registers or other registers.
[0289] In at least one embodiment, one or more processors 2202 are coupled to one or more interface buses 2210 to transmit communication signals such as address, data, or control signals between the processor 2202 and other components within the system 2200. In at least one embodiment, the interface bus 2210 can be a processor bus such as a version of a Direct Media Interface (DMI) bus in one embodiment. In at least one embodiment, the interface 2210 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2202 includes an integrated memory controller 2216 and a platform controller hub 2230. In at least one embodiment, the memory controller 2216 facilitates communication between the memory device and other components of the system 2200, while the platform controller hub (PCH) 2230 provides connections to I / O devices via a local I / O bus.
[0290] In at least one embodiment, the memory device 2220 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device having suitable performance to serve as a process memory. In at least one embodiment, the memory device 2220 operates as a system memory for the system 2200 and can store data 2222 and instructions 2221 for use when one or more processors 2202 execute an application or process. In at least one embodiment, the memory controller 2216 is also coupled to an optional external graphics processor 2212, and this graphics processor may communicate with one or more graphics processors 2208 within the processor 2202 to perform graphics and media operations. In at least one embodiment, the display device 2211 can be connected to the processor 2202. In at least one embodiment, the display device 2211 can include one or more of an internal display device such as a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, the display device 2211 can include a head-mounted display (HMD) such as a stereoscopic display device for use in virtual reality (VR) applications or augmented reality (AR) applications.
[0291] In at least one embodiment, the platform controller hub 2230 enables peripheral devices to be connected to the memory device 2220 and the processor 2202 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2246, a network controller 2234, a firmware interface 2228, a wireless transceiver 2226, a touch sensor 2225, and a data storage device 2224 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 2224 can be connected via a storage interface (e.g., SATA) or a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2225 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2226 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2228 enables communication with system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2234 enables a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2210. In at least one embodiment, the audio controller 2246 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2200 includes an optional legacy I / O controller 2240 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 2230 can also be connected to one or more universal serial bus (USB) controller 2242 connected input devices, such as a combination of a keyboard and a mouse 2243, a camera 2244, or other USB input devices.
[0292] In at least one embodiment, instances of the memory controller 2216 and the platform controller hub 2230 may be integrated into a separate external graphics processor, such as the external graphics processor 2212. In at least one embodiment, the platform controller hub 2230 and / or the memory controller 2216 may be external to one or more processors 2202. For example, in at least one embodiment, the system 2200 can include an external memory controller 2216 and a platform controller hub 2230, which may be configured as a memory controller hub and a peripheral device controller hub in a system chipset that communicates with the processor 2202.
[0293] In at least one embodiment, at least one component illustrated or described with respect to FIG. 22 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics processor 2208 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 2208 is used for layer demapping, descrambling, and de-rate matching of 5G NR PUSCH data that is soft-demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, a processor core 2207 performs a kernel startup function that passes parameters to at least one kernel on at least one graphics processor 2208 that layer demaps, descrambles, and de-rate matches data from a 5G NR antenna in parallel.
[0294] FIG. 23 is a block diagram of a processor 2300 having one or more processor cores 2302A-2302N, an integrated memory controller 2314, and an integrated graphics processor 2308, according to at least one embodiment. In at least one embodiment, the processor 2300 can include a lesser number of additional cores including an additional core 2302N represented by the dashed rectangle. In at least one embodiment, each of the processor cores 2302A-2302N includes one or more internal cache units 2304A-2304N. In at least one embodiment, each processor core can also access one or more shared cache units 2306.
[0295] In at least one embodiment, internal cache units 2304A - 2304N, and shared cache unit 2306 represent a cache memory hierarchy within processor 2300. In at least one embodiment, cache memory units 2304A - 2304N may include at least one level of cache for instructions and data within each processor core, as well as one or more levels of shared intermediate - level cache such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest - level cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2306 and 2304A - 2304N.
[0296] In at least one embodiment, processor 2300 may also include a set of one or more bus controller units 2316 and system agent core 2310. In at least one embodiment, one or more bus controller units 2316 manage a set of peripheral buses such as one or more PCI or PCI Express buses. In at least one embodiment, system agent core 2310 provides management functions for various processor components. In at least one embodiment, system agent core 2310 includes one or more integrated memory controllers 2314 for managing access to various external memory devices (not shown).
[0297] In at least one embodiment, one or more of the processor cores 2302A-2302N include support for simultaneous multi-threading. In at least one embodiment, the system agent core 2310 includes components for coordinating and operating cores 2302A-2302N during multi-threaded processing. In at least one embodiment, the system agent core 2310 may further include a power control unit (PCU) that includes logic and components for adjusting the power state of one or more of the processor cores 2302A-2302N and the graphics processor 2308.
[0298] In at least one embodiment, the processor 2300 further includes a graphics processor 2308 for performing graphics processing operations. In at least one embodiment, the graphics processor 2308 is coupled to a shared cache unit 2306 and a system agent core 2310 that includes one or more integrated memory controllers 2314. In at least one embodiment, the system agent core 2310 also includes a display controller 2311 for causing the output of the graphics processor to be provided to one or more attached displays. In at least one embodiment, the display controller 2311 may also be a separate module coupled to the graphics processor 2308 via at least one interconnect, or may be integrated within the graphics processor 2308.
[0299] In at least one embodiment, a ring-based interconnect unit 2312 is used to couple the internal components of the processor 2300. In at least one embodiment, alternative interconnect units such as point-to-point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 2308 is coupled to the ring interconnect 2312 via an I / O link 2313.
[0300] In at least one embodiment, I / O link 2313 represents at least one of a variety of I / O interconnects including an on-package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 2318 such as an eDRAM module. In at least one embodiment, each of processor cores 2302A - 2302N and graphics processor 2308 uses embedded memory module 2318 as a shared last-level cache.
[0301] In at least one embodiment, processor cores 2302A - 2302N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, processor cores 2302A - 2302N are heterogeneous from the perspective of an instruction set architecture (ISA), where one or more of processor cores 2302A - 2302N execute a common instruction set, but one or more other cores of processor cores 2302A - 2302N execute a subset of the common instruction set, or a different instruction set. In at least one embodiment, processor cores 2302A - 2302N are heterogeneous from the perspective of a microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, processor 2300 can be implemented on one or more chips or as a SoC integrated circuit.
[0302] In at least one embodiment, at least one component illustrated or described with respect to FIG. 23 is utilized to implement the techniques and / or functions described in connection with FIGS. 1-6. In at least one embodiment, at least one graphics processor 2308 is used to cause information received from a plurality of 5G new radio antennas to be decoded in parallel by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 2308 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel. In at least one embodiment, at least one processor core 2302 executes a kernel launch function that passes parameters to at least one kernel on at least one graphics processor 2308 that layer demaps, descrambles, and derate matches data from 5G NR antennas in parallel.
[0303] FIG. 24 is a block diagram of a graphics processor 2400, which may be an individual graphics processing unit or a graphics processor integrated with a plurality of processing cores. In at least one embodiment, the graphics processor 2400 communicates with registers of the graphics processor 2400 using commands placed in memory via an I / O interface mapped to the memory. In at least one embodiment, the graphics processor 2400 includes a memory interface 2414 for accessing memory. In at least one embodiment, the memory interface 2414 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.
[0304] In at least one embodiment, the graphics processor 2400 also includes a display controller 2402 for driving display output data towards a display device 2420. In at least one embodiment, the display controller 2402 includes hardware for one or more overlapping planes for the display device 2420 and for the composition of multi-layered video or user interface elements. In at least one embodiment, the display device 2420 can be an internal or external display device. In at least one embodiment, the display device 2420 is a head-mounted display device such as a virtual reality (VR) display device or an augmented reality (AR) display device. In at least one embodiment, the graphics processor 2400 includes a video codec engine 2406 for encoding, decoding, or transcoding media in, from, or between one or more media coding formats including but not limited to Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4 AVC, as well as Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats such as JPEG and Motion JPEG (MJPEG).
[0305] In at least one embodiment, the graphics processor 2400 includes a block image transfer (BLIT) engine 2404 for performing two-dimensional (2D) rasterizer operations, such as bit boundary block transfers. However, in at least one embodiment, 2D graphics operations are performed using one or more components of the graphics processing engine (GPE) 2410. In at least one embodiment, the GPE 2410 is a compute engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0306] In at least one embodiment, the GPE 2410 includes a 3D pipeline 2412 for performing 3D operations, such as rendering 3D images and scenes, using processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 2412 includes programmable and fixed function elements that perform various tasks for the 3D / media subsystem 2415 and / or spawn execution threads. Although media operations can be performed using the 3D pipeline 2412, in at least one embodiment, the GPE 2410 also includes a media pipeline 2416 that is used to perform media operations, such as post-processing of video and image enhancement.
[0307] In at least one embodiment, media pipeline 2416 includes fixed function or programmable logic units for performing one or more special media operations such as video decode acceleration, video deinterlacing, and video encode acceleration, instead of or on behalf of video codec engine 2406. In at least one embodiment, media pipeline 2416 further includes a thread spawning unit for spawning threads for execution in 3D / media subsystem 2415. In at least one embodiment, the spawned threads execute computations for media operations on one or more graphics execution units included in 3D / media subsystem 2415.
[0308] In at least one embodiment, 3D / media subsystem 2415 includes logic for executing threads spawned by 3D pipeline 2412 and media pipeline 2416. In at least one embodiment, 3D pipeline 2412 and media pipeline 2416 send thread execution requests to 3D / media subsystem 2415, which includes thread dispatch logic for arbitrating various requests and dispatching them to available thread execution resources. In at least one embodiment, the execution resources include an array of graphics execution units for processing 3D and media threads. In at least one embodiment, 3D / media subsystem 2415 includes one or more internal caches for thread instructions and data. In at least one embodiment, subsystem 2415 also includes shared memory including registers and addressable memory for sharing data between threads and storing output data.
[0309] In at least one embodiment, at least one component illustrated or described with respect to FIG. 24 is utilized to implement the techniques and / or functions described in relation to FIGS. 1 - 6. In at least one embodiment, at least one graphics processor 2400 is used to decode in parallel information received from a plurality of 5G new radio antennas by a plurality of processor pipelines. In at least one embodiment, at least one graphics processor 2400 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, wherein at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0310] FIG. 25 is a block diagram of a graphics processing engine 2510 of a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processing engine (GPE) 2510 is one version of the GPE 2410 shown in FIG. 24. In at least one embodiment, the media pipeline 2516 is optional and may not be explicitly included within the GPE 2510. In at least one embodiment, a separate media and / or image processor is coupled to the GPE 2510.
[0311] In at least one embodiment, the GPE 2510 is coupled to or includes the command streamer 2503, which provides a command stream to the 3D pipeline 2512 and / or the media pipeline 2516. In at least one embodiment, the command streamer 2503 is coupled to a memory, which may be a system memory, or one or more of an internal cache memory and a shared cache memory. In at least one embodiment, the command streamer 2503 receives commands from the memory and transmits the commands to the 3D pipeline 2512 and / or the media pipeline 2516. In at least one embodiment, the commands are instructions, primitives, or micro-operations fetched from a ring buffer that stores commands for the 3D pipeline 2512 and the media pipeline 2516. In at least one embodiment, the ring buffer can further include a batch command buffer that stores batches of multiple commands. In at least one embodiment, the commands for the 3D pipeline 2512 can also include references to data stored in a memory such as, but not limited to, vertex and shape data for the 3D pipeline 2512 and / or image data and memory objects for the media pipeline 2516. In at least one embodiment, the 3D pipeline 2512 and the media pipeline 2516 process commands and data by performing operations or by dispatching one or more execution threads to the graphics core array 2514. In at least one embodiment, the graphics core array 2514 includes one or more blocks of graphics cores (e.g., graphics core 2515A, graphics core 2515B), and each block includes one or more graphics cores.In at least one embodiment, each graphics core includes a set of graphics execution resources including general-purpose and graphics-specific execution logic for performing graphics and compute operations, as well as fixed-function texture processing and / or machine learning, and artificial intelligence acceleration logic.
[0312] In at least one embodiment, the 3D pipeline 2512 includes fixed-function and programmable logic for processing one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 2514. In at least one embodiment, the graphics core array 2514 provides an integrated block of execution resources for use in processing shader programs. In at least one embodiment, the versatile execution logic (e.g., execution units) within the graphics cores 2515A-2515B of the graphics core array 2514 includes support for various 3D API shader languages and can execute multiple concurrent execution threads associated with multiple shaders.
[0313] In at least one embodiment, the graphics core array 2514 also includes execution logic for performing media functions, such as video and / or image processing. In at least one embodiment, the execution units further include programmable general-purpose logic for performing parallel general-purpose computing operations in addition to graphics processing operations.
[0314] In at least one embodiment, output data generated by a thread running on the graphics core array 2514 can output data to the memory of the unified return buffer (URB) 2518. The URB 2518 can store data for multiple threads. In at least one embodiment, the URB 2518 may be used to transmit data between different threads running on the graphics core array 2514. In at least one embodiment, the URB 2518 may further be used for synchronization between a thread on the graphics core array 2514 and fixed function logic within the shared function logic 2520.
[0315] In at least one embodiment, the graphics core array 2514 is scalable, whereby the graphics core array 2514 includes a variable number of graphics cores, and each graphics core has a variable number of execution units based on the power and performance levels targeted by the GPE 2510. In at least one embodiment, the execution resources are dynamically scalable, whereby the execution resources may be enabled or disabled as needed.
[0316] In at least one embodiment, the graphics core array 2514 is coupled to a shared function logic 2520 that includes a plurality of resources shared among the graphics cores of the graphics core array 2514. In at least one embodiment, the shared functions implemented by the shared function logic 2520 are embodied in hardware logic units that provide dedicated supplementary functions to the graphics core array 2514. In at least one embodiment, the shared function logic 2520 includes, but is not limited to, the logic of the sampler 2521, mathematics 2522, and inter-thread communication (ITC) 2523. In at least one embodiment, one or more caches 2525 are included in or coupled to the shared function logic 2520.
[0317] In at least one embodiment, when there is insufficient demand for a dedicated function and it cannot be included within the Graphics Core Array 2514, a shared function is used. In at least one embodiment, an instance of a dedicated function instantiated as one is used in the shared function logic 2520 and shared among other execution resources within the Graphics Core Array 2514. In at least one embodiment, specific shared functions within the shared function logic 2520 that are widely used by the Graphics Core Array 2514 may be included within the shared function logic 2516 within the Graphics Core Array 2514. In at least one embodiment, the shared function logic 2516 within the Graphics Core Array 2514 can include some or all of the logic within the shared function logic 2520. In at least one embodiment, all logical elements within the shared function logic 2520 may be replicated within the shared function logic 2516 of the Graphics Core Array 2514. In at least one embodiment, the shared function logic 2520 is advantageously excluded from the shared function logic 2516 within the Graphics Core Array 2514.
[0318] In at least one embodiment, at least one component illustrated or described with respect to FIG. 25 is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one Graphics Processing Engine 2510 is used to decode information received from a plurality of 5G New Radio antennas in parallel by a plurality of processor pipelines. In at least one embodiment, at least one Graphics Processing Engine 2510 is used for layer demapping, descrambling, and derate matching of 5G NR PUSCH data that is soft demapped for LDPC decoding, where at least one thread block is used to perform operations on code block data elements (e.g., LLRs) in parallel.
[0319] FIG. 26 is a block diagram of the hardware logic of a graphics processor core 2600 according to at least one embodiment described herein. In at least one embodiment, the graphics processor core 2600 is included within a graphics core array. In at least one embodiment, the graphics processor core 2600, which may also be referred to as a core slice, can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2600 is an example of one graphics core slice, and the graphics processor described herein may include multiple graphics core slices based on the targeted power and performance envelope. In at least one embodiment, each graphics core 2600 can include a fixed function block 2630 coupled to a plurality of sub - cores 2601A - 2601F, also referred to as sub - slices, which include modular blocks of general - purpose and fixed function logic.
[0320] In at least one embodiment, the fixed function block 2630 includes a geometry / fixed function pipeline 2636 that can be shared by all sub - cores within the graphics processor 2600, for example, in a low - performance and / or low - power graphics processor implementation. In at least one embodiment, the geometry / fixed function pipeline 2636 includes a 3D fixed function pipeline, a video front - end unit, a thread spawner and thread dispatcher, and an integrated return buffer manager that manages an integrated return buffer.
[0321] In at least one embodiment, the fixed function block 2630 also includes a graphics SoC interface 2637, a graphics microcontroller 2638, and a media pipeline 2639. The graphics SoC interface 2637 provides an interface between the graphics core 2600 and other processor cores within the system-on-chip integrated circuit. In at least one embodiment, the graphics microcontroller 2638 is a programmable sub-processor configurable to manage various functions of the graphics processor 2600, including thread dispatch, scheduling, and preemption. In at least one embodiment, the media pipeline 2639 includes logic to facilitate decoding, encoding, preprocessing, and / or postprocessing of multimedia data including image and video data. In at least one embodiment, the media pipeline 2639 implements media operations via requests to compute logic or sampling logic within sub-cores 2601-2601F.
[0322] In at least one embodiment, the SoC interface 2637 enables the graphics core 2600 to communicate with a general-purpose application processor core (e.g., a CPU) and / or other components within the SoC. Other components within the SoC include memory hierarchy elements such as a shared last-level cache memory, system RAM, and / or an embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2637 also enables communication with fixed-function devices within the SoC, such as a camera imaging pipeline, enables and / or implements the use of global memory atomics that can be shared between the graphics core 2600 and the CPU within the SoC, and / or implements power management control of the graphics core 2600 and enables interfacing between the clock domain of the graphics core 2600 and other clock domains within the SoC. In at least one embodiment, the SoC interface 2637 is configured to receive command buffers from a command streamer and global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, the commands and instructions can be dispatched to the media pipeline 2639 when media operations are executed, or to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 2636, geometry and fixed-function pipeline 2614) when graphics processing operations are executed.
[0323] In at least one embodiment, the graphics microcontroller 2638 can be configured to perform various scheduling and management tasks for the graphics core 2600. In at least one embodiment, the graphics microcontroller 2638 can execute graphics and / or compute workload scheduling in the execution unit (EU) arrays 2602A - 2602F, 2604A - 2604F within the sub - cores 2601A - 2601F. In at least one embodiment, host software running on the CPU core of the SoC including the graphics core 2600 can send a workload to one of a plurality of graphics processor doorbells, and this doorbell calls a scheduling operation for an appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload to execute next, sending the workload to the command streamer, pre - empting existing workloads running on the engine, managing the progress of the workload, and notifying the host software when the workload is complete. In at least one embodiment, the graphics microcontroller 2638 can also facilitate the low - power or idle state of the graphics core 2600 and provide the graphics core 2600 with the function of saving and restoring registers within the graphics core 2600 throughout the transition to the low - power state, independent of the operating system and / or the graphics driver software on the system.
[0324] In at least one embodiment, the graphics core 2600 may have more than, or less than, up to N modular sub-cores than the illustrated sub-cores 2601A - 2601F. For each set of N sub-cores, in at least one embodiment, the graphics core 2600 may also include shared function logic 2610, shared and / or cache memory 2612, geometry / fixed function pipeline 2614, and additional fixed function logic 2616 for accelerating various graphics and computing processing operations. In at least one embodiment, the shared function logic 2610 may include logic units (e.g., sampler, math, and / or inter-thread communication logic) that can be shared by each of the N sub-cores within the graphics core 2600. The shared and / or cache memory 2612 can be a last-level cache for the N sub-cores 2601A - 2601F within the graphics core 2600 and can also serve as a shared memory accessible by multiple sub-cores. In at least one embodiment, the geometry / fixed function pipeline 2614 may be included instead of the geometry / fixed function pipeline 2636 within the fixed function block 2630 and can include the same or similar logic units.
[0325] In at least one embodiment, the graphics core 2600 includes additional fixed function logic 2616 that can include various fixed function acceleration logic for use by the graphics core 2600. In at least one embodiment, the additional fixed function logic 2616 includes an additional geometry pipeline for use in position only shading. In position only shading, there are at least two geometry pipelines, but in the full geometry pipeline and the cull pipeline within the geometry / fixed function pipelines 2616, 2636, and this cull pipeline is an additional geometry pipeline that may be included within the additional fixed function logic 2616. In at least one embodiment, the cull pipeline is a scaled down version of the full geometry pipeline. In at least one embodiment, the full pipeline and the cull pipeline can execute different instances of an application, and each instance has a separate context. In at least one embodiment, position only shading can hide long cull runs of discarded triangles and can complete shading faster in some instances. For example, in at least one embodiment, the cull pipeline fetches and shades vertex position attributes without rasterizing and rendering pixels to the frame buffer, so the cull pipeline logic within the additional fixed function logic 2616 can execute the position shader in parallel with the main application and generate critical results overall faster than the full pipeline. In at least one embodiment, the cull pipeline can use the generated critical results to compute visibility information for all triangles, regardless of whether these triangles are culled. In at least one embodiment, the full pipeline (which may be called the replay pipeline in this instance) can consume the visibility information, skip culled triangles, and shade only visible triangles, and this visible triangle is ultimately passed to the rasterization phase.
[0326] In at least one embodiment, the additional fixed-function logic 2616 can also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, for implementations that include optimization of machine learning training or inference.
[0327] In at least one embodiment, within each of the graphics sub-cores 2601A - 2601F, there is a set of execution resources that may be used to perform graphics operations, media operations, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-cores 2601A - 2601F include a plurality of EU arrays 2602A - 2602F, 2604A - 2604F, thread dispatch and inter-thread communication (TD / IC) logic 2603A - 2603F, 3D (e.g., texture) samplers 2605A - 2605F, media samplers 2606A - 2606F, shader processors 2607A - 2607F, and shared local memory (SLM) 2608A - 2608F. Each of the EU arrays 2602A - 2602F, 2604A - 2604F includes a plurality of execution units that are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logical operations in the service of graphics operations, media operations, or compute operations including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 2603A - 2603F performs local thread dispatch and thread control operations for the execution units within the sub-core and facilitates communication between threads executing on the execution units of the sub-core. In at least one embodiment, the 3D samplers 2605A - 2605F can read texture or other 3D graphics-related data into memory. In at least one embodiment, the 3D sampler can read texture data in different ways based on the configured sample state and texture format associated with a given texture. In at least one embodiment, the media samplers 2606A - 2606F can perform similar read operations based on the type and format associated with media data.In at least one embodiment, each of the graphics sub-cores 2601A - 2601F can alternatively include a 3D and media integrated sampler. In at least one embodiment, threads executing on the execution units within each of the sub-cores 2601A - 2601F can utilize the shared local memories 2608A - 2608F within each sub-core so that threads executing within a thread group can execute using a common pool of on-chip memory.
[0328] In at least one embodiment, at least one component illustrated or described with respect to FIG. 26 is utilized to implement the techniques and / or functions described in connection with FIGS. 1 - 6. In at least one embodiment, at least one graphics processor core 2600 is used to decode information received from a plurality of 5G new radio antennas in parallel by a plurality of processor pipelines. In a...
Claims
1. A processor comprising one or more circuits that cause information received from a plurality of 5th generation (5G) new radio antennas to be decoded in parallel by corresponding ones of a plurality of processor pipelines.
2. The processor of claim 1, wherein the information received from the plurality of 5G new radio antennas is processed by soft demapping, and decoding of the information includes layer demapping, descrambling, and derate matching.
3. The processor of claim 1, wherein decoding of the information received from the plurality of 5G new radio antennas includes assigning the information to a plurality of groups of threads, wherein each group of threads among the plurality of groups of threads computes a corresponding code block to be decoded from a transfer block.
4. The processor of claim 1, wherein the information includes data in a first mapping configuration, and decoding of the information includes layer demapping and descrambling the data, and storing the layer demapped and descrambled data in a second mapping configuration.
5. The processor of claim 4, wherein the one or more circuits cause the plurality of processor pipelines to combine the layer demapped and descrambled data with previously received data in a hybrid automatic repeat request (HARQ) buffer.
6. The processor of claim 1, wherein the information received from the plurality of 5G new radio antennas is processed by soft demapping and stored as a value representing a log likelihood ratio, and the one or more circuits cause the plurality of processor pipelines to layer demap the value according to a mapping function.
7. The processor of claim 1, wherein the information received from the plurality of 5G new radio antennas is processed by soft demapping, and decoding of the information includes deinterleaving the information in parallel.
8. The processor of claim 2, wherein derate matching includes inserting filler bits.
9. A machine-readable medium storing a set of instructions that, when executed, cause a parallel processor to at least: A machine-readable medium that causes the parallel processor to decode information transmitted using a plurality of fifth-generation (5G) new radio signals by scheduling a plurality of thread groups each corresponding to at least one of the plurality of 5G new radio signals on the parallel processor.
10. The machine-readable medium of claim 9, wherein the information includes data being processed by soft demapping, and scheduling the plurality of thread groups includes scheduling the plurality of thread groups for layer demapping, descrambling, and rate matching of the data.
11. The machine-readable medium of claim 9, wherein the information includes data being processed by soft demapping, and scheduling the plurality of thread groups includes scheduling the plurality of thread groups to descramble the value by reading a value corresponding to a log-likelihood ratio from the data and multiplying the value according to a descrambling function.
12. The machine-readable medium of claim 11, wherein scheduling the plurality of thread groups includes scheduling the plurality of thread groups to combine the descrambled value with previously received data in a hybrid automatic repeat request (HARQ) buffer.
13. The machine-readable medium of claim 9, wherein scheduling the plurality of thread groups includes scheduling the plurality of thread groups to deinterleave code block data elements.
14. The machine-readable medium of claim 9, wherein scheduling the plurality of thread groups includes scheduling the plurality of thread groups to rate match a code block by inserting filler bits at least in part.
15. The machine-readable medium of claim 10, wherein scheduling the plurality of thread groups to layer demap the data is at least partially based on a mapping function that uses a thread block index as a parameter.
16. The machine-readable medium according to claim 15, wherein the mapping function also uses the transport block size and the number of multiple-input multiple-output (MIMO) layers as parameters.
17. One or more processors that decode in parallel information received from a plurality of fifth-generation (5G) new radio antennas by corresponding multiple processor pipelines, A system comprising one or more memories for storing data corresponding to the information.
18. The system according to claim 17, wherein the information received from the plurality of 5G new radio antennas is processed by soft demapping, and the decoding of the information includes layer demapping, descrambling, and rate dematching.
19. The system according to claim 17, wherein the decoding of the information received from the plurality of 5G new radio antennas includes allocating the information to a plurality of groups of threads, and each group of threads among the plurality of groups of threads computes a corresponding code block to be decoded from a transport block.
20. The one or more processors cause the information to be decoded in parallel by at least partially causing the data to be converted for decoding by a decoder and storing the converted data in the one or more memories, and causing the data to be converted includes causing the data to be rate dematched in parallel by the plurality of processor pipelines. The system according to claim 17.
21. The system according to claim 20, wherein storing the converted data in the one or more memories includes combining the converted data with the contents of a hybrid automatic repeat request (HARQ) buffer.
22. The system according to claim 17, wherein the one or more processors cause the information to be decoded in parallel by at least partially extracting a transport block from the data and extracting a code block from the transport block.
23. The system according to claim 17, wherein the data includes values corresponding to log likelihood ratios, and the one or more processors at least partially decode the information in parallel by descrambling the data based at least in part on multiplying the values according to a descrambling function.
24. The data is stored in the one or more memories of a first mapping configuration, and the one or more processors cause the plurality of processor pipelines to read the data in parallel from the one or more memories, perform layer demapping on the data in parallel, perform rate dematching on the data in parallel, perform descrambling on the data in parallel, and combine the data that has been layer demapped, rate dematched, and descrambled with the contents of a hybrid automatic repeat request (HARQ) buffer of a second mapping configuration.
25. A method comprising accessing data corresponding to information transmitted using a plurality of fifth generation (5G) new radio signals using a plurality of thread groups executed by a parallel processor, and decoding the data by the plurality of thread groups.
26. The method according to claim 25, wherein the data is processed by soft demapping, and decoding the data includes converting the data for decoding by a decoder.
27. The method according to claim 26, wherein the data includes data elements of a plurality of code blocks, and converting the data includes allocating one or more of the plurality of thread groups to handle each code block.
28. The method according to claim 26, wherein converting the data includes performing layer demapping on the data.
29. The method according to claim 26, wherein converting the data includes performing rate dematching on the data.
30. The method according to claim 26, wherein converting the data includes reading values corresponding to log likelihood ratios from the data and descrambling the values by multiplying the values according to a descrambling function.
31. The method of claim 30, further comprising combining the descrambled value with previously received data in a hybrid automatic repeat request (HARQ) buffer. **Claim 32** The method of claim 26, wherein converting the data comprises converting the data for decoding by a low density parity check (LDPC) decoder in a fifth generation (5G) new radio (NR) physical uplink shared channel (PUSCH) physical layer (PHY).
Citation Information
Patent Citations
Method and apparatus for controlling reliability of feedback signal in mobile communication system for supporting HARQ
JP2007053771A
Hybrid automatic repeat request buffer flushing mechanism
US20090086657A1