Rapid retraining of fully fused nerve transceiver assemblies

CN115733569BActive Publication Date: 2026-09-29NVIDIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211054013.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-24
Filing Date
2022-08-30
Publication Date
2026-09-29
Estimated Expiration
2042-08-30

Smart Images

  • Figure CN115733569B_ABST
    Figure CN115733569B_ABST
Patent Text Reader

Abstract

The present disclosure relates to fast retraining of fully fused transceiver components. Systems, apparatuses, and methods are provided for fast retraining of a fully fused neural network configured to implement at least a portion of a transceiver. At least one of a demapping module, an equalization module, or a channel estimation module can be implemented at least in part using the fully fused neural network. The neural network can be trained online during operation by using a plurality of received data frames to obtain a training dataset. Periodic retraining of the neural network is performed to adapt the neural network to changing channel characteristics. In various embodiments, a neural demapper, a neural channel estimator, and a neural receiver are disclosed to replace or augment one or more components of a transceiver. In another embodiment, an autoencoder can be implemented between a transmitter and a receiver to replace a majority of components of a transceiver, the autoencoder trained by an end-to-end learning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 238,496, filed August 30, 2021, entitled “FAST RETRAINING OF FULLY-FUSED NEURAL TRANSCEIVER COMPONENTS,” the entire contents of which are incorporated herein by reference. Background Technology

[0003] Neural networks (NNs) can be used to replace or add signal processing blocks in transceivers used in communication systems. However, in wireless communication, it is generally accepted that it is impossible to retrain such neural networks online due to complexity and latency. Therefore, neural network training can at most be performed offline before being deployed to a device.

[0004] For example, it is believed that it is difficult to accurately train neural networks to adapt to channel characteristics that can change rapidly and drastically, especially when the user equipment, including the transceiver, is a mobile device such as a cellular phone. Factors such as other sources of interference, relative motion between the receiver and transmitter, and the number of independent signal paths between the transmitter and receiver can change very rapidly in real-world environments. For example, a user using a cellular phone might suddenly walk into a subway station and board a train, which would then accelerate away from or towards the transmitter. To mitigate these problems and allow the neural network to respond to the widest possible range of predicted channel responses, the complexity of the neural network may increase and / or the size of the training dataset used during training may become very large. However, increasing the complexity of the neural network only negatively impacts the computational cost during inference on communication latency and energy efficiency, which may prove to be a limiting problem in telecommunications applications. Furthermore, increased complexity may also increase training time, which only exacerbates the online training problem. Therefore, these problems and / or other issues related to the prior art need to be addressed. Summary of the Invention

[0005] A framework for implementing fully fused neural networks in communication systems is provided. Various embodiments of this disclosure provide a method, system, and apparatus comprising at least one fully fused neural network.

[0006] According to a first aspect of this disclosure, a communication apparatus is provided for communicating via a channel. The communication apparatus includes a receiver. The communication apparatus also includes a processor configured to at least partially implement a first neural network to perform tasks associated with the receiver. The first neural network is periodically retrained using a training dataset comprising multiple data frames received by the receiver via the channel. The first neural network comprises a fully fused neural network.

[0007] In an embodiment of the first aspect, the processor includes a parallel processing unit configured to implement a fully fused neural network using at least one or more tensor cores.

[0008] In an embodiment of the first aspect, the first neural network includes a multilayer perceptron (MLP) comprising at least one hidden layer, an activation layer, and an output layer.

[0009] In an embodiment of the first aspect, the first neural network is configured to generate a log-likelihood ratio (LLR) of transmitted bits based on equalization symbols of an orthogonal frequency division multiplexing (OFDM) frame received by a receiver through a channel. An OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0010] In a first aspect embodiment, the first neural network includes a multilayer perceptron (MLP) configured to receive equalization symbols of OFDM frames. Noise variance estimation associated with each resource element of an OFDM frame The positional encoding p(i,j) of the index tuple (i,j) associated with each resource element is taken as input. The MLP is configured to generate m LLRs for each equalization symbol of an OFDM frame, where 2 m It is equal to the order of the constellation used in the quadrature amplitude modulation (QAM) modulation and coding scheme (MCS) for communication over a channel.

[0011] In an embodiment of the first aspect, the first neural network is periodically retrained based on a training dataset. For each of a plurality of OFDM frames, the training dataset comprises a set of tuples, each tuple consisting of an equalization symbol for a specific resource element of the OFDM frame and the corresponding valid bits of the equalization symbol. The corresponding valid bits are generated by encoding the information bits using forward error correction (FEC), and the information bits are derived by the channel decoder based on LLR and are confirmed as valid based on a cyclic redundancy check (CRC) using the last C bits of the information bits.

[0012] In an embodiment of the first aspect, the first neural network is periodically retrained using stochastic gradient descent every f OFDM frames, where f is a positive integer. In one embodiment, the first neural network is retrained at least once every 10 ms during receiver operation to receive multiple OFDM frames.

[0013] In one embodiment of the first aspect, the first neural network is configured to generate an estimate of the channel matrix based on the receive resource grid of the t-th orthogonal frequency division multiplexing (OFDM) frame received by the receiver through the channel. The OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0014] In an embodiment of the first aspect, the first neural network is configured to generate an estimate of a three-dimensional tensor representing information bits based on the received resource grid of the t-th orthogonal frequency division multiplexing (OFDM) frame received by the receiver through the channel. The OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0015] In an embodiment of the first aspect, the processor is further configured to at least partially implement a second neural network to perform a second task associated with the receiver. The first neural network performs a demapping task, and the second neural network performs a channel estimation task. The first neural network is configured to at least partially process the output of the second neural network as input to the first neural network.

[0016] In an embodiment of the first aspect, the communication device communicates with a second communication device via a channel. The second communication device includes a transmitter and a second processor, the second processor being configured to at least partially implement a second neural network to perform tasks associated with the transmitter of the second communication device.

[0017] In an embodiment of the first aspect, the transmitter of the second communication device is configured to transmit a known data sequence to the receiver of the communication device. The first neural network is trained based on a binary cross-entropy loss between the known data and the predicted log-likelihood ratio (LLR) generated by the first neural network.

[0018] According to a second aspect of this disclosure, a communication system is provided. The communication system includes a first communication device, which includes a receiver for communicating via a channel. The first communication device also includes a processor configured to at least partially implement a first neural network to perform tasks associated with the receiver. The first neural network is periodically retrained using a training dataset comprising multiple data frames received by the receiver via the channel. The first neural network comprises a fully fused neural network.

[0019] In an embodiment of the second aspect, the first neural network is configured to generate the log-likelihood ratio (LLR) of the transmitted bits based on the equalized symbols of an orthogonal frequency division multiplexing (OFDM) frame received by the receiver through the channel. An OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0020] In a second aspect embodiment, the first neural network includes a multilayer perceptron (MLP) configured to receive equalization symbols of OFDM frames. Noise variance estimation associated with each resource element of an OFDM frame The positional encoding p(i,j) of the index tuple (i,j) associated with each resource element is taken as input. The MLP is configured to generate m LLRs for each equalization symbol of an OFDM frame, where 2 m It is equal to the order of the constellation used in the quadrature amplitude modulation (QAM) modulation and coding scheme (MCS) for communication over a channel.

[0021] In a second aspect embodiment, the first neural network is periodically retrained based on training data. For each of the plurality of OFDM frames, the training data comprises a set of tuples, each tuple consisting of the equalization symbol of a specific resource element of the OFDM frame and the corresponding valid bits of the equalization symbol. The corresponding valid bits are generated by encoding the information bits using forward error correction (FEC), and the information bits are derived by the channel decoder based on LLR and are confirmed as valid based on the cyclic redundancy check (CRC) using the last C bits of the information bits.

[0022] In an embodiment of the second aspect, the first neural network is periodically retrained using stochastic gradient descent every f OFDM frames, where f is a positive integer. In one embodiment, the first neural network is retrained at least once every 10 ms during receiver operation to receive multiple OFDM frames.

[0023] In one embodiment of the second aspect, the first neural network is configured to generate an estimate of the channel matrix based on the receive resource grid of the t-th orthogonal frequency division multiplexing (OFDM) frame received by the receiver through the channel. The OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0024] In an embodiment of the second aspect, the first neural network is configured to generate an estimate of a three-dimensional tensor representing information bits based on the received resource grid of the t-th orthogonal frequency division multiplexing (OFDM) frame received by the receiver through the channel. The OFDM frame consists of N subcarriers and K symbols for each subcarrier.

[0025] In a second aspect embodiment, the communication system further includes a second communication device comprising a transmitter. The second communication device also includes a second processor configured to at least partially implement a second neural network to perform tasks associated with the transmitter of the second communication device. The transmitter of the second communication device is configured to transmit a known data sequence to a receiver of the communication device, and the first neural network is trained based on a binary cross-entropy loss between the known data and the predicted log-likelihood ratio (LLR) generated by the first neural network.

[0026] According to a third aspect of this disclosure, a method for training a transceiver including a neural network component is provided. The method includes: acquiring a training dataset based on one or more frames received by a receiver of a communication device. The training data includes a set of tuples, the set of tuples including equalization symbols for one or more frames and a corresponding valid bit sequence for each equalization symbol. The valid bit sequence is encoded by forward error correction applied to the information bits that have successfully passed Cyclic Redundancy Check (CRC) based on the last C bits of the information bits. The method further includes periodically retraining a first neural network using the training dataset for multiple frames, the first neural network being configured to perform a task associated with the receiver. The first neural network includes a fully fused neural network.

[0027] In one embodiment of the third aspect, retraining is performed in B iterations, each iteration corresponding to the training from... A batch of index tuples (i, j) are deterministically or randomly selected from the given set, where It is the set of all index tuples (i, j) corresponding to the resource elements that carry valid data in the frame.

[0028] According to a fourth aspect of this disclosure, a method for end-to-end training of an autoencoder is provided. The autoencoder includes a first neural network corresponding to a transmitter of a first communication device and a second neural network corresponding to a receiver of a second communication device. The method includes: initializing parameters of the first and second neural networks; transmitting a known data sequence to the receiver via the transmitter; generating a set of equalization symbols via the receiver in response to the transmitted known data sequence; training the second neural network based on a binary cross-entropy loss between the known data sequence and a predicted log-likelihood ratio (LLR) generated by the second neural network based on the set of equalization symbols; transmitting a second known data sequence via the transmitter, wherein the transmitter is configured to perturb the output of the transmitter according to perturbation noise; receiving a feedback loss signal from the receiver via the transmitter; calculating a second loss signal based on the feedback loss signal and the perturbation noise; and training the first neural network based on the second loss signal. The method also includes periodically retraining the first neural network, which is configured to perform a task associated with the receiver using a training dataset for multiple frames. The second neural network includes a fully fused neural network.

[0029] In an embodiment of the fourth aspect, the training of the second neural network and the first neural network is repeated periodically according to a termination criterion.

[0030] In an embodiment of the fourth aspect, the second neural network is a fully fused neural network. Attached Figure Description

[0031] The present system and method for implementing transceiver components using a fully fused neural network are described in detail below with reference to the accompanying drawings.

[0032] Figure 1A A block diagram of a communication system according to one embodiment is shown.

[0033] Figure 1B An orthogonal frequency division multiplexing (OFDM) frame according to one embodiment is shown.

[0034] Figure 2A A neural demapper, implemented for a transceiver of a communication device according to an embodiment, is shown.

[0035] Figure 2B This is a flowchart of a method for retraining the neural components of a transceiver according to an embodiment.

[0036] Figure 2C A neural channel estimator 230 according to an embodiment is shown.

[0037] Figure 2D A neural receiver 240 according to an embodiment is shown.

[0038] Figure 3A End-to-end learning between a pair of fully fused neural networks operating on channels respectively via a transmitter and a receiver, according to an embodiment, is illustrated.

[0039] Figure 3B This is a flowchart of a method 350 for end-to-end learning between a transmitter and a receiver according to an embodiment.

[0040] Figure 4 The illustration shows an example parallel processing unit suitable for implementing some embodiments of the present disclosure.

[0041] Figure 5A For the purpose of implementing some embodiments of this disclosure, the use of Figure 4 A conceptual diagram of the processing system implemented by the PPU.

[0042] Figure 5B The illustration shows an exemplary system in which various architectures and / or functions of various prior embodiments can be implemented.

[0043] Figure 5C The illustration shows components of an exemplary system that can be used to train and utilize machine learning in at least one embodiment.

[0044] Figure 6 This is an example system diagram of a game streaming system according to some embodiments of the present disclosure. Detailed Implementation

[0045] This paper proposes using a fully fused neural network to replace certain components of the transceiver. In such applications, using a fully fused neural network can be orders of magnitude faster than traditional methods that attempt to use more complex neural networks. Retraining the fully fused neural network using a training dataset derived from received data determined to be valid based on cyclic redundancy check (CRC) allows for real-time adaptation to changing channel conditions. In contrast, traditional techniques train neural networks offline on large datasets before deployment and do not perform additional training after deployment.

[0046] Fully fused neural networks can be used as a replacement for or enhancement of traditional transceiver components for physical layer processing. For example, fully fused neural networks can be used to replace or enhance signal processing modules (e.g., channel estimation, soft symbol demapping, etc.) in a communication system transceiver.

[0047] At least one of the demapping module, equalization module, or channel estimation module can be implemented, at least partially, using a fully fused neural network. The neural network can be trained online during operation by acquiring a training dataset using multiple received data frames. The neural network is periodically retrained to adapt it to changing channel characteristics. In various embodiments, a neural demapping unit, a neural channel estimator, and a neural receiver are disclosed to replace or add one or more components of the transceiver. In another embodiment, an autoencoder can be implemented between the transmitter and receiver to replace most components of the transceiver; the autoencoder is trained using an end-to-end learning algorithm.

[0048] Figure 1A A block diagram of a communication system 100 according to one embodiment is shown. The communication system 100 includes a transmitter 102 and a receiver 106 configured to communicate via a channel 104. In one embodiment, the communication system 100 may be implemented as a cellular communication network, such as a Long Term Evolution (LTE) wireless network (also known as a fourth-generation or 4G network), a New Radio (NR) wireless network (also known as a fifth-generation or 5G network), etc. Alternatively, the communication system 100 may be a Wi-Fi network according to the IEEE 802.11 standard. In such a wireless communication system, channel 104 is a communication medium having dynamic characteristics based on the objects between the transmitter and receiver in a real-world environment.

[0049] Transmitter 102 includes an encoding module 112, a mapping module 114, a pilot insertion module 116, a precoding module 118, and a modulation module 120. These modules can at least partially implement the physical layer of the telecommunications stack implemented by a user equipment (UE) including the transmitter. Encoding module 112 converts bit sequences into encoded sequences based on line codes. Mapping module 114 maps encoded symbols of the bit stream to various resource elements of a frame. Pilot insertion module 116 inserts pilot signals into a subset of the resource elements of the frame. Precoding module 118 processes signals from multiple antennas to increase the signal strength at receiver 106. Modulation module 120 generates radio frequency (RF) signals transmitted by the antennas of transmitter 102. Modulation module 120 can transmit different RF signals on multiple carrier frequencies using, for example, quadrature amplitude modulation (QAM).

[0050] Receiver 106 includes a synchronization module 122, a demodulation module 124, a channel estimation module 126, an equalization module 128, a demapping module 130, and a decoding module 132. These modules can at least partially implement the physical layer of a telecommunications stack implemented by a user equipment (UE) including the receiver. Synchronization module 122 uses different delay values ​​to synchronize signals received at different carrier frequencies. Demodulation module 124 extracts signals from carriers used for different carrier frequencies. Channel estimation module 126 uses pilot signals to estimate the characteristics of channel 104 in order to adjust the received signal to account for attenuation, phase shift, and / or noise at different carrier frequencies. Equalization module 128 adjusts the signal in the frequency domain based on the channel estimation. Demapping module 130 maps resource elements to a bitstream of coded symbols. Decoding module 132 converts the coded symbols back into the original bitstream.

[0051] As used herein, a module can refer to a set of functions performed by any combination of hardware or software. For example, modulation module 120 can be implemented in a device as a hardware chip, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), etc., configured to modulate a carrier signal to generate an RF signal transmitted by one or more antennas. Conversely, some modules can be implemented in software by a processor, such as a central processing unit (CPU), parallel processing unit (PPU), reduced instruction set computer (RISC), or other coprocessors or microcontrollers configured to execute a set of instructions to process signals or data. It should be understood that various modules can be implemented in the digital and / or analog domains. A wide range of variations exist in how communication hardware designers can choose to implement the various components of the transmitter or receiver generally described above.

[0052] Figure 1B An orthogonal frequency division multiplexing (OFDM) frame 150 according to one embodiment is shown. OFDM frame 150 includes K symbols transmitted over a period of time on each of N subcarriers. A subcarrier refers to a specific center frequency of an RF signal, wherein the bandwidth of the OFDM frame is distributed across multiple subcarriers frequency-separated according to subcarrier spacing. An OFDM frame can be described as containing multiple time-frequency resources, and each time-frequency resource may be referred to herein as a resource element (RE) 152. RE 152 includes a symbol representing multiple bits. A subset of RE 152 of OFDM frame 150 contains pilot signals 154, which... Figure 1B They are shown as shaded squares in the grid.

[0053] Compared to conventional methods for implementing the physical layer components of communication devices, one or more of the modules described above can be replaced or enhanced by neural networks. For example, in the various embodiments described herein, one or more of the demapping module 130, the channel estimation module 126, and / or the decoding module 132 can be replaced by neural networks that predict the outputs of these modules based on one or more inputs. Each module can be replaced or enhanced by a single and different neural network, or alternatively, multiple modules can be replaced or enhanced by a single neural network. In some embodiments, a large portion of the components of the transmitter and receiver in two communication devices can be replaced by a pair of neural networks, which are then jointly trained.

[0054] When channel characteristics are stable, the use of neural networks offers improvements compared to traditional PHY layer implementations of transceivers. In other words, neural networks can be initially trained using large datasets or simulated channel responses, and then deployed to the system without retraining. However, in practice, channel characteristics are rarely stable unless the communication equipment is stationary and the environment around the equipment remains unchanged. In most cases, periodic retraining or online training is required to improve equipment quality and reduce the bit error rate of transmission. One advantage of using online training is that neural networks can be smaller because there is no need to train the neural network to generalize a wide set of channel responses. Furthermore, online training can compensate for hardware impairments such as nonlinearities in the RF chain. However, online training is constrained by very strict latency constraints (i.e., low latency) and small training datasets that use as few computational resources as possible.

[0055] Neural networks can be fully fused, allowing them to run during inference and / or be trained at extremely high rates. This enables the neural network to quickly adapt to changing channel characteristics as the UE or other device moves around in real-world environments, both online and at runtime. As used in this paper, a fully fused neural network refers to a neural network implemented in a single core on a parallel processor. The core can include multiple threads or blocks of threads (e.g., 32 threads), each performing multiple operations on a different subset (e.g., batch) of the input data. In some cases, instructions in each thread can be executed on a tensor core configured to efficiently perform matrix multiplication operations. The performance of a fully fused neural network is improved by minimizing access to global memory (e.g., off-chip video random access memory (VRAM), high-level L2 cache, etc.) and fully utilizing faster on-chip memory (e.g., L1 cache, shared memory / register file, etc.). In other words, the core is designed so that only slow memory accesses read the neural network's input from global memory and write the neural network's output to global memory. All intermediate values ​​are operated on in high-speed on-chip memory (e.g., registers or other shared memory and low-level cache).

[0056] In one embodiment, the fully fused neural network is implemented by dividing the input vector into batches of a specific width (e.g., 128 elements wide), where each batch is processed by a corresponding thread block. Each thread in the thread block then computes the number of rows (e.g., a subset of rows) of the output of a particular layer of the neural network by loading a subset of weights from global memory into registers and computing a portion of the output for all columns of the input in one or more passes. In some embodiments, the row and element portions of the weights can be cut into, for example, 16x16 element blocks to be processed in the tensor kernel. For the remaining rows of the output, new weights can be loaded into registers, and the remaining portions of the output can be computed for the batch. Intermediate results after processing each layer are retained in fast on-chip shared memory as it is processed through multiple layers of the fully fused neural network, alternating between element-wise weight multiplication operations and element-wise application of activation functions. Once all layers of the neural network have been processed, the output can be written back from shared memory to global memory.

[0057] Examples of fully-fused neural networks are described in U.S. Patent Application No. 17 / 340,283, filed June 7, 2021, entitled “Fully-Fused Neural Network Execution,” and U.S. Patent Application No. 17 / 672,566, filed February 15, 2022, entitled “Multiresolution Hash Encoding for Neural Networks,” each of which is incorporated herein by reference in its entirety.

[0058] Figure 2A The neural component of a transceiver in a communication apparatus according to an embodiment is illustrated. A neural demapping unit 210 can be used to replace or enhance a conventional demapping module, and the neural demapping unit 210 converts equalized symbols into predictions of the log-likelihood ratio (LLR) of the transmitted bits for each equalized symbol of a frame. After synchronization, cyclic prefix removal, and OFDM demodulation (e.g., using a Fast Fourier Transform operation), the received resource grid... It can be written as:

[0059]

[0060] in It is the channel matrix. It is a resource grid for transmission. It is an element with independent and identically distributed (iid) distribution. Additive white Gaussian noise (AWGN), This is an element-wise multiplication operator. In some embodiments, N represents not only AWGN, but also other types of interference, such as, but not limited to, interference transmission, self-interference due to imperfect synchronization, and quantization. It can also be understood that the receiver can only access an estimate of H, denoted as... It is used to equalize the received resource grid. There are many well-known ways to implement an equalizer, and the neural demapper 210 is unaware of the specific implementation details of the equalizer. The desired result is that the selected particular equalizer generates a balanced resource grid. As shown below:

[0061]

[0062] in Residual noise and self-interference are represented for all indexes. in It is a collection of all indices for all resource elements. It should be understood that each index refers to a specific resource element in the OFDM frame.

[0063] Neural demapper 210, for For each index (i, j) tuple in the array, its corresponding resource element carries data (i.e., not pilot signals), and the corresponding bit b is calculated. i,j,m LLR, where m = 1, ..., M, where 2 M M is the order of the QAM constellation used to modulate the RF signal (e.g., for QAM-16, M equals 4).

[0064] In traditional and widely used demapping modules, the so-called Gaussian demapping module (e.g., the posterior probability (APP) demapping module derived for AWGN) assumes that... Make:

[0065]

[0066] in and It is the set of constellation symbols where the m-th bit is equal to 0 and 1 respectively. The LLR is then fed into the decoder to recover the transmitted bits. As shown in Equation 3, the equalization symbols... and the corresponding estimated variance for each equilibrium sign. The data is provided to a Gaussian demapper, which calculates the corresponding LLR according to a given formula.

[0067] In practice, it is difficult to determine the variance. The estimated value of Z is not good, and the distribution of the elements of Z may differ significantly from the Gaussian distribution assumed in Equation 3 above. Therefore, there may be a large mismatch between the true APPLLR and the LLR calculated according to Equation 3. Mismatch typically leads to a decrease in the coding error rate (BER). Instead, the neural demapper 210 learns the mapping from the balanced symbols to the LLR, and the periodic retraining of the fully fused neural network implemented therein will compensate for the mismatch to improve performance compared to the conventional Gaussian demapper.

[0068] In one embodiment, the neural demapper 210 implements a first neural network configured to predict LLR based on equalized symbols. In one embodiment, the first neural network includes equalized symbols configured to receive OFDM frames. The noise variance estimation associated with each resource element of an OFDM frame using a multilayer perceptron (MLP). And the position encoding p(i,j) of the index tuple (i,j) associated with each resource element as input, and generate m LLRs for each equalization symbol of the OFDM frame, where 2 m This is equal to the order of the constellation used in the Quadrature Amplitude Modulation (QAM) modulation and coding scheme (MCS) for communication over a channel. First, the neural network implements the function... like Figure 2A As shown, the position encoder 212 receives the index tuple (i, j) and generates an encoding for the index tuple. This index tuple, along with the balance symbol and the corresponding noise variance estimation The result of being passed to the neural demapper 210 is:

[0069]

[0070]

[0071] In one embodiment, the position encoder 212 generates frequency-type encodings for index tuples (i, j) of the following form:

[0072]

[0073] in:

[0074]

[0075] Z is any positive real number. In other embodiments, the position encoder 212 may use other encoding algorithms, such as parametric encoding, normalizing grid positions to the interval [0,1], or at least one of any other type of encoding and embedding techniques.

[0076] As described above, the neural demapper 210 benefits from online retraining. Even if the channel estimation task (performed by the channel estimation module) is not perfectly accurate, continuously adjusting the neural network helps improve the BER. The method 220 for retraining the neural demapper 210 is performed in two steps, as follows: Figure 2B As shown.

[0077] In 222, forward error correction (FEC) and cyclic redundancy check (CRC) are used to acquire training data. Let... This represents the equalization symbol grid for frame t. This represents the corresponding noise variance estimate. This represents the set of all indexed tuples (i, j) that carry data for the corresponding resource element. The neural demapper 210 is executed to target... Each symbol in the matrix generates an LLR prediction, as shown below:

[0078]

[0079] in It is the Exponential Moving Average (EMA), calculated as follows:

[0080]

[0081] Where η(t) = 1 - α t Furthermore, α∈[0,1] is a hyperparameter. A typical value for α can be 0.99 to achieve a good trade-off between fast adaptation and time stability.

[0082] Next, LLR of all index tuples (i, j) i,j The values ​​of (t) are stacked into a vector, which is provided to the decoding module to decode the information bits. in yes The cardinality (i.e., the number of data symbols) is M, where M is the number of bits per symbol, and r ∈ (0, 1-1) is the code rate. In one embodiment, the decoding module can implement low-density parity-check (LDPC) or polar decoding algorithm to obtain a vector of transmitted information bits.

[0083] Once the information bits have been decoded, a CRC check is performed to verify their validity. Assume the last C bits of u(t) are used as check bits. If the CRC is valid, the frame transmission is considered successful, and the result can be used for online training.

[0084] To populate the training dataset Q, the information bits u(t) are re-encoded using FEC to obtain the vector. Then it is divided into M-bit blocks b i,j(t)∈{0,1} M This corresponds to the bit on each symbol mapped to the (i, j)th resource element. Then, the set of tuples is stored. Used for training.

[0085] In 224, a neural network is trained using a training dataset based on stochastic gradient descent. In one embodiment, training is performed iteratively on mini-batch tuples, with each batch... A subset corresponds to this. More specifically, for b = 1, ..., B, training is performed as follows:

[0086]

[0087] in It refers to from The index tuple (i, j) of B-micro-batch, which is deterministically or randomly selected. This represents the average loss term l d (t) The Nambla operator for the gradient, where γ>0 is a hyperparameter representing the learning rate. In one embodiment, the loss term l d (t) is the binary cross-entropy loss, calculated as follows:

[0088]

[0089] in The output of the neural network after applying the sigmoid activation function is given below:

[0090]

[0091] In one embodiment, the initial weights θ(0) can be chosen randomly or deterministically based on a selected distribution, or they can be obtained through meta-optimization. Meta-optimization is described in more detail in Finn et al., “Model-agnostic meta-learning for fast adaptation of deep networks,” International Conference on Machine Learning, Vol. 70, pp. 1126-1135 (2017), the entire contents of which are incorporated herein by reference.

[0092] In some embodiments, training data may include valid tuples from multiple frames (e.g., the first r frames). That is, the training set Ideally, the retraining task should be run once every r frames, using the valid tuples collected in the last r frames as the training dataset. It's understood that the batch size and batch number can remain constant regardless of how many frames the training dataset contains. In other words, even if the training dataset contains a large number of tuples, only a small subset of the available training data can be used to retrain the neural network during a given retraining task.

[0093] It should be understood that the method disclosed above can be generalized to the case of multiple-input multiple-output (MIMO), in which case the equalization symbol... This corresponds to a single received stream.

[0094] In some embodiments, codewords (e.g., bit vectors passed to the decoding module) may be spread or interleaved across multiple frames, rather than being included in a single frame. However, for example, it is straightforward to retrieve codewords from multiple frames using the method described above before decoding and performing CRC checks.

[0095] Conceptually, the above method can also be implemented using single-carrier transmission. However, in this case, it might be necessary to send an additional pilot signal at the start of transmission to train the system using the pilot signal before transmitting data.

[0096] In one embodiment, for illustration, the communication system can be configured to transmit data on a carrier frequency of 3.5 GHz, where N = 128 subcarriers spaced at 30 kHz. An OFDM frame consists of K = 14 OFDM symbols with a cyclic prefix length of 8. Each frame carries 1792 QAM16 symbols, corresponding to a codeword with an information bit length of 7168 bits, at a code rate of r = 0.5. Encoding / decoding uses standard-compliant 5G LDPC codes.

[0097] In one embodiment, the neural demapper 210 can be implemented using parameters P=8 and Z=10. 4 The neural network is trained using positional encoding. It can be an MLP with one hidden layer of 32 neurons, followed by a rectified linear unit (ReLU) activation layer and four output neurons (corresponding to M=4) with no activation. The neural demapper can be trained sequentially on a dataset including a selected number of the last frames.210

[0098] The preceding description demonstrated the benefits of using a fully fused neural network to replace the traditional demapper that relies on compensation based on dynamic channel characteristics. However, other advantages can be obtained using the same framework. For example, at high carrier frequencies such as in terahertz communication, residual receiver phase noise can cause a significant degrade in achievable communication performance. As a result, the previously described conventional Gaussian demapper is known to be suboptimal in such cases. In this context, the fully fused neural demapper can implicitly learn the demapping task to continuously track changes in the noise distribution.

[0099] Another example is the demand for low-budget consumer hardware, which has driven the need for robust and adaptive signal processing components that can compensate for hardware impairments. These impairments can be caused by higher manufacturing tolerances, hardware changes, temperature drift, and so on in cheaper devices. One example is IQ imbalance, which leads to non-orthogonality between in-phase and quadrature signal components. Other examples are amplifier nonlinearity and / or carrier frequency offset. However, fully fused neural demappers can learn to compensate for these impairments, which may also change over time.

[0100] In yet another example, the channel may exhibit non-AWGN behavior (e.g., colored noise due to the channel's low-pass or band-pass behavior). A fully fused neural demapper can learn to continuously adapt to the current noise characteristics, rather than relying on the assumption of AWGN.

[0101] Figure 2C A neural channel estimator 230 according to an embodiment is shown. As described above, the channel estimation module can also be configured to replace and / or enhance the neural network for predicting channel characteristics. The neural channel estimator 230 can be used with a conventional demapper, such as a Gaussian demapper, or in combination with the neural demapper 210. In other words, a first neural network can be trained to perform a decoding task, and a second neural network can be trained to perform a channel estimation task.

[0102] like Figure 2C As shown, the neural network implemented by the neural channel estimator 230 should be continuously (e.g., periodically) retrained to learn the second-order properties of the changing channel (e.g., updated covariance matrix). Compared to the neural demapper 210, the neural channel estimator 230 suffers from the problem that the true value of the channel estimate is unavailable; therefore, the neural network is trained based on the estimate of the true channel estimate. However, by using a channel estimate that generates valid codewords (based on CRC), the training dataset will produce generally better sufficient results than conventional channel estimation algorithms.

[0103] In one embodiment, the neural channel estimator 230 receives the resource grid Y(t) of the received t-th frame. Let This indicates the transmission bits obtained through frame processing, and their validity has been verified using FEC and CRC, as described above. These bits are first mapped to the corresponding transmission symbol X used according to the constellation and bit markings (e.g., QAM with gray markings). i,j (t). These symbols are then used with resource elements carrying pilot signals to estimate the channel matrix H(t), using Residual noise variance estimation. From the channel matrix and noise variance The initial estimate derived can be processed as input to the demapping module (e.g., neural demapping module 210). Tuples This data is then used as the training dataset for the neural channel estimator 230. As previously mentioned, the training dataset can include data from multiple frames (e.g., r previous frames). The training dataset is then used for online retraining of the fully fused neural network.

[0104] Figure 2D A neural receiver 240 according to an embodiment is shown. As described above, the tasks of channel estimation, equalization, and demapping can also be performed jointly by a single neural network. The neural receiver 240 is a fully fused neural network that directly estimates the bit sequence based on the resource grid of the received frames.

[0105] In one embodiment, neural receiver 240 receives the resource grid Y(t) of the received t-th frame. Neural receiver 240 is configured to predict information bits b. i,j Let B(t) be a three-dimensional tensor of (t). This represents the transmitted bits obtained by processing the frame, and their validity has been verified using FEC and CRC, as described above. The tuple Q(t) = {(Y(t), B(t))} is then used as the training dataset for the neural receiver 240. As previously mentioned, the training dataset can include data from multiple frames (e.g., r previous frames). The training dataset is then used for online retraining of the fully fused neural network.

[0106] Figure 3A End-to-end learning between a pair of fully fused neural networks operating across a channel, respectively via a transmitter and a receiver, is illustrated according to an embodiment. The principle of end-to-end learning is to replace the traditional transmitter and receiver with a pair of fully fused neural networks and then jointly train them to communicate over an actual channel. Previously, these so-called autoencoders were trained offline using data from static channels or a predefined set of possible channel implementations. However, the online training speed of fully fused neural networks opens up a new area of ​​application where radio frequency waveforms adapt to new environments in milliseconds or less (e.g., within the channel's coherence time).

[0107] like Figure 3A As shown, a transmitter 302, including a first neural network 310, communicates with a receiver 306, including a second neural network 320, via a channel 304. The first neural network 310 has trainable parameters θ. tx The second neural network 320 has trainable parameters θ rx Therefore, the first neural network 310 implements the function. Where M represents the number of bits b mapped to symbol x, and F2 represents the binary field. Let f ch (x) represents the channel, which is usually a random mapping that depends on the radio environment, the speed of the receiver relative to the transmitter, etc.

[0108] The second neural network 320 implementation function Make:

[0109]

[0110] Here, b represents a vector of M transmitted information bits, and LLR represents the corresponding LLR estimate at the receiver. The training task is to minimize the binary cross-entropy between b and LLR.

[0111] At least one of the first neural network 310 or the second neural network 320 can be a fully fused neural network. Jointly training the two components allows the system 300 to learn customized transceiver components or adapt to the waveform of the actual channel during operation. Feedback latency (i.e., the required signal propagation time, including computation latency) and training latency become new limiting factors. In particular, if the transmitter / receiver distance is large, the speed of light can effectively limit the update rate. For this reason, using a fully fused neural network capable of very fast online training can be important for end-to-end learning applications.

[0112] The channel itself is a non-differentiable component. Therefore, the transmitter 302 requires a carefully tuned training process. Figure 3B This is a flowchart of a method 350 for end-to-end learning between a transmitter and a receiver according to an embodiment.

[0113] At 352, the trainable parameter θ tx and θ rx Initialization is performed. At 354, a series of known data is sent from transmitter 302 to receiver 306. At 356, the second neural network 320 is trained based on the received symbols, using the known data sequence as the ground truth output of the second neural network. Steps 354 and 356 can be repeated in multiple iterations using the same or different sets of known data. If performance does not improve or once the maximum number of iterations is reached, the training of the second neural network 320 is complete.

[0114] At 358, transmitter 302 adds small perturbation noise to the transmitter output. In other words, the output x is given as:

[0115]

[0116] in The variance is The exploration noise is an additional training hyperparameter. At 360, transmitter 302 receives the feedback loss signal from receiver 306, which may be a BCE. It is important to note that in this embodiment, both the means of transmitter 302 and receiver 306 must include transceivers (i.e., both transmitter and receiver) to enable bidirectional communication. At 362, transmitter 302 uses the feedback loss signal along with knowledge of the perturbation noise to compute a new loss function for training the first neural network 310 using stochastic gradient descent. If performance does not improve or once the maximum number of iterations is reached, training of the first neural network 310 is complete.

[0117] In one embodiment, steps 354-362 can be repeated continuously until a selected termination criterion is met.

[0118] In yet another embodiment, the fully fused neural network can be implemented by the communication system to predict whether channel decoding is likely to fail or succeed based on a given LLR implementation generated by the demapper. Such an embodiment can even trigger frame retransmission before the decoder has determined that decoding of the current frame has failed. This can help reduce the average decoding latency across multiple frames, as retransmission can begin for frames predicted as failed based on the generated LLR before decoding is complete.

[0119] The proposed technique for rapidly retraining neural transceiver components is not limited to the uses described above. Other applications in digital communications that can utilize fully fused neural networks include digital predistortion applications and / or full-duplex interference computation.

[0120] As described above, the methods and techniques can be implemented using any combination of hardware and software, including using one or more processors to at least partially implement a fully fused neural network. In one embodiment, the fully fused neural network can be implemented using a host processor and / or a parallel processing unit configured to execute multiple instructions to perform various operations to implement the functionality of the neural network. For example, the neural demapper 210 can be implemented at least partially by using a PPU 400 to execute a series of instructions, as described in more detail below. This can include dividing the layers of the neural network into portions assigned to different threads and / or cooperative thread arrays. The threads or cooperative thread arrays can include instructions executed using tensor cores or dedicated matrix multiplication units to accelerate the execution of instructions that implement the functionality of the neural network.

[0121] Now, based on the user's needs, further illustrative information will be provided regarding the various optional architectures and features that can implement the aforementioned framework. It should be strongly noted that the following information is for illustrative purposes only and should not be construed as limiting in any way. Any of the following features may be optionally combined, excluding or not excluding the other features described.

[0122] Parallel processing architecture

[0123] Figure 4 A parallel processing unit (PPU) 400 according to an embodiment is illustrated. According to an embodiment, the PPU 400 can be used to implement rapid retraining of a fully fused neural transceiver component. In one embodiment, the PPU 400 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 400 is a latency-hiding architecture designed to process many threads in parallel. A thread (e.g., an execution thread) is an instantiation of a set of instructions configured to be executed by the PPU 400. In one embodiment, the PPU 400 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device. In other embodiments, the PPU 400 can be used to perform general-purpose computing. Although an exemplary parallel processor is provided herein for illustrative purposes, it should be specifically noted that such a processor is illustrated for illustrative purposes only, and any processor may be employed to complement and / or replace this processor.

[0124] One or more PPU 400s can be configured to accelerate thousands of high-performance computing (HPC), data center, cloud computing, and machine learning applications. PPU 400s can be configured to accelerate numerous deep learning systems and applications used in autonomous vehicles, simulations, computational graphics such as ray or path tracing, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.

[0125] like Figure 4 As shown, PPU 400 includes an input / output (I / O) unit 405, a front-end unit 415, a scheduler unit 420, a job allocation unit 425, a hub 430, a crossbar (Xbar) 470, one or more general purpose processing clusters (GPCs) 450, and one or more memory partitioning units 480. PPU 400 can be connected to a host processor or other PPU 400 via one or more high-speed NVLink 410 interconnects. PPU 400 can be connected to a host processor or other peripheral devices via interconnect 402. PPU 400 can also be connected to local memory 404, which includes multiple memory devices. In one embodiment, local memory may include multiple dynamic random access memory (DRAM) devices. The DRAM devices may be configured as a high-bandwidth memory (HBM) subsystem, wherein multiple DRAM dies are stacked within each device.

[0126] The NVLink 410 interconnect enables the system to expand and include one or more PPUs 400 in conjunction with one or more CPUs, supporting cache coherency between the PPUs 400 and the CPU, as well as the CPU controller. Data and / or commands can be sent from or from the NVLink 410 to other units of the PPU 400 via hub 430, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). Figure 5B A more detailed description of the NVLink 410.

[0127] I / O unit 405 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) via interconnect 402. I / O unit 405 may communicate directly with the host processor via interconnect 402, or via one or more intermediate devices such as memory bridges. In one embodiment, I / O unit 405 may communicate with one or more other processors, such as one or more PPUs 400, via interconnect 402. In one embodiment, I / O unit 405 implements a Peripheral Component Interconnect High Speed ​​(PCIe) interface for communication via a PCIe bus, and interconnect 402 is a PCIe bus. In alternative embodiments, I / O unit 405 may implement other types of known interfaces for communication with external devices.

[0128] I / O unit 405 decodes data packets received via interconnect 402. In one embodiment, the data packets represent commands configured to cause PPU 400 to perform various operations. I / O unit 405 transmits the decoded commands to various other units of PPU 400 that these commands may specify. For example, some commands may be transmitted to front-end unit 415. Other commands may be transmitted to hub 430 or other units of PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, I / O unit 405 is configured to route communication between and among the various logical units of PPU 400.

[0129] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to the PPU 400 for processing. The workload may include instructions and data to be processed by those instructions. The buffer is an area of ​​memory accessible (e.g., read / write) by both the host processor and the PPU 400. For example, I / O unit 405 may be configured to access a buffer in system memory connected to interconnect 402 via a memory request transmitted through interconnect 402. In one embodiment, the host processor writes a command stream to the buffer and then transmits a pointer to the start of the command stream back to the PPU 400. Front-end unit 415 receives pointers to one or more command streams. Front-end unit 415 manages the one or more streams, reads commands from these streams, and forwards the commands to the respective units of the PPU 400.

[0130] Front-end unit 415 is coupled to scheduler unit 420, which configures various GPCs 450 to process tasks defined by the one or more streams. Scheduler unit 420 is configured to track status information related to the various tasks managed by scheduler unit 420. Status can indicate which GPC 450 a task is assigned to, whether the task is active or inactive, the priority associated with the task, etc. Scheduler unit 420 manages the execution of multiple tasks on the one or more GPCs 450.

[0131] Scheduler unit 420 is coupled to job allocation unit 425, which is configured to dispatch tasks for execution on GPC 450. Job allocation unit 425 can track several scheduled tasks received from scheduler unit 420. In one embodiment, job allocation unit 425 manages a pending task pool and an active task pool for each GPC 450. When GPC 450 completes the execution of a task, the task is evicted from the active task pool of GPC 450, and one of the other tasks from the pending task pool is selected and scheduled for execution on GPC 450. If an active task on GPC 450 is idle, for example, while waiting for data dependencies to be resolved, then the active task can be evicted from GPC 450 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 450.

[0132] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API), enabling one or more applications executing on the host processor to be scheduled for operations to be performed on the PPU 400. In one embodiment, multiple computing applications are executed concurrently by the PPU 400, and the PPU 400 provides isolation, Quality of Service (QoS), and independent address spaces for the multiple computing applications. Applications may generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks to be executed by the PPU 400. The driver kernel outputs the tasks to one or more streams being processed by the PPU 400. Each task may include one or more associated thread groups, referred to herein as warps. In one embodiment, a warp includes 32 associated threads that can execute in parallel. Cooperative threads may refer to multiple threads that include instructions for performing tasks and can exchange data via shared memory. These tasks may be assigned to one or more processing units within the GPC 450, and instructions are scheduled for execution by at least one warp.

[0133] The work allocation unit 425 communicates with one or more GPCs 450 via an XBar 470. The XBar 470 is an interconnect network that couples a plurality of units of the PPU 400 to other units of the PPU 400. For example, the XBar 470 can be configured to couple the work allocation unit 425 to a specific GPC 450. Although not explicitly shown, one or more other units of the PPU 400 may also be connected to the XBar 470 via a hub 430.

[0134] Tasks are managed by scheduler unit 420 and dispatched to GPC 450 by job allocation unit 425. GPC 450 is configured to process tasks and generate results. Results can be consumed by other tasks within GPC 450, routed to different GPC 450 via XBar 470, or stored in memory 404. Results can be written to memory 404 via memory partitioning unit 480, which implements a memory interface for reading data from and writing data to memory 404. Results can be transferred to another PPU 400 or CPU via NVLink 410. In one embodiment, PPU 400 includes U number of memory partitioning units 480, equal to the number of individual and distinct memory devices coupled to memory 404 of PPU 400. Each GPC 450 may include a memory management unit to provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit provides one or more translation back buffers (TLBs) for performing virtual address to physical address translation in memory 404.

[0135] In one embodiment, memory partitioning unit 480 includes a raster operation (ROP) unit, a secondary (L2) cache, and a memory interface coupled to memory 404. The memory interface can implement 32-bit, 64-bit, 128-bit, or 1024-bit data buses for high-speed data transfer. PPU 400 can connect to up to Y memory devices, such as high-bandwidth memory stacks or graphics dual data rate, version 5, synchronous dynamic random access memory, or other types of persistent storage devices. In one embodiment, the memory interface implements an HBM2 memory interface, and Y equals half of U. In one embodiment, the HBM2 memory stack is located on the same physical package as PPU 400, providing significant power and area savings compared to conventional GDDR5 SDRAM systems. In one embodiment, each HBM2 stack includes four memory dies and Y equals 4, wherein each HBM2 stack includes two 129-bit channels per die, for a total of eight channels, and the data bus width is 1024 bits.

[0136] In one embodiment, memory 404 supports single error correction double detection (SECDED) error correction codes (ECC) to protect data. ECC provides a high level of reliability for computing applications sensitive to data corruption. Reliability is particularly important where the PPU 400 is handling very large datasets and / or in large-scale cluster computing environments where applications run for extended periods.

[0137] In one embodiment, PPU 400 implements a multi-level memory hierarchy. In one embodiment, memory partitioning unit 480 supports unified memory that provides a single, unified virtual address space for the CPU and PPU 400 memory, allowing data sharing between virtual memory systems. In one embodiment, the frequency with which PPU 400 accesses memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 400 that is accessing these pages more frequently. In one embodiment, NVLink 410 supports address translation services, allowing PPU 400 to directly access the CPU's page table and providing PPU 400 with full access to the CPU's memory.

[0138] In one embodiment, the replication engine transfers data between multiple PPUs 400 or between a PPU 400 and a CPU. The replication engine can generate page faults for addresses that are not mapped to a page table. The memory partitioning unit 480 can then repair the page faults, mapping these addresses to the page table, after which the replication engine can perform the transfer. In conventional systems, memory is fixed (e.g., non-pageable) for multiple replication engine operations across multiple processors, significantly reducing available storage. In the event of a hardware page fault, addresses can be passed to the replication engine without concern for whether memory pages reside, and the replication process is transparent.

[0139] Data from memory 404 or other system memory can be fetched by memory partitioning unit 480 and stored in an on-chip L2 cache 460 shared among the various GPCs 450. As shown, each memory partitioning unit 480 includes a portion of the L2 cache associated with the corresponding memory 404. Low-level caches can then be implemented in various units within the GPC 450. For example, each processing unit within the GPC 450 can implement a Level 1 (L1) cache. The L1 cache is a private memory dedicated to a specific processing unit. The L2 cache 460 is coupled to memory interface 470 and XBar 470, and data from the L2 cache can be fetched and stored in each of the L1 caches for processing.

[0140] In one embodiment, the processing unit within each GPC 450 implements a SIMD (Single Instruction Multiple Data) architecture, where each thread in a thread group (e.g., a thread bundle) is configured to process a different data set based on the same set of instructions. All threads in the thread group execute the same instructions. In another embodiment, the processing unit implements a SIMT (Single Instruction Multiple Thread) architecture, where each thread in the thread group is configured to process a different data set based on the same set of instructions, but individual threads in the thread group are allowed to diverge during execution. In one embodiment, maintaining a program counter, call stack, and execution state for each thread bundle allows for concurrency between thread bundles and serial execution within a thread bundle when threads diverge. In another embodiment, maintaining a program counter, call stack, and execution state for each individual thread allows for equal concurrency among all threads within and between thread bundles. When maintaining an execution state for each individual thread, threads executing the same instructions can aggregate and execute in parallel for maximum efficiency.

[0141] Cooperative groups are a programming model for organizing groups of communicating threads. They allow developers to express the granularity at which threads are communicating, enabling richer and more efficient expressions of parallel decomposition. The Cooperative Startup API supports synchronization between blocks of threads executing parallel algorithms. Conventional programming models provide a single, simple construct for synchronizing cooperative threads: a barrier across all threads in a thread block (e.g., the `syncthreads()` function). However, programmers often prefer to define thread groups smaller than thread blocks in the form of a collective group-wide function interface, and synchronize within the defined group to allow for greater performance, design flexibility, and software reuse.

[0142] Collaboration groups enable programmers to explicitly define thread groups (as small as a single thread) at the sub-block and multi-block granularity, and perform collective operations such as synchronization on threads within the collaboration group. This programming model supports clean composition across software boundaries, allowing libraries and utility functions to be safely synchronized within their local contexts without having to make assumptions about aggregation. The collaboration group primitive allows for the implementation of new cooperative parallelism patterns, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire mesh of thread blocks.

[0143] Each processing unit comprises a large number (e.g., 128, etc.) of different processing cores (e.g., functional units), which may be fully pipelined, single-precision, double-precision, and / or mixed-precision, and include floating-point arithmetic logic units and integer arithmetic logic units. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core comprises 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0144] Tensor cores are configured to perform matrix operations. Specifically, tensor cores are configured to perform deep learning matrix arithmetic, such as GEMM (matrix-matrix multiplication), for convolution operations during neural network training and inference. In one embodiment, each tensor core operates on a 4x4 matrix and performs matrix multiplication and accumulation operations, D = A × B + C, where A, B, C, and D are 4x4 matrices.

[0145] In one embodiment, matrix multiplication inputs A and B can be integer, fixed-point, or floating-point matrices, while accumulation matrices C and D can be integer, fixed-point, or floating-point matrices of equal or higher bit width. In one embodiment, the Tensor Core operates on 1-bit, 4-bit, or 8-bit integer input data using 32-bit integer accumulation. An 8-bit integer matrix multiplication requires 1024 operations and results in a full-precision product, which is then accumulated with other intermediate multiplications using 32-bit integer addition for an 8x8x16 matrix multiplication. In one embodiment, the Tensor Core operates on 16-bit floating-point input data using 32-bit floating-point accumulation. A 16-bit floating-point multiplication requires 64 operations and results in a full-precision product, which is then accumulated with other intermediate multiplications using 32-bit floating-point addition for a 4x4x4 matrix multiplication. In practice, the Tensor Core is used to perform operations on much larger two-dimensional or higher-dimensional matrices composed of these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix loading, matrix multiplication and accumulation, and matrix storage operations for efficient use with Tensor Core from CUDA-C++ programs. At the CUDA level, the thread bundle-level interface takes a 16x16 matrix spanning all 32 threads of the thread bundle.

[0146] Each processing unit may also include M Special Function Units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, an SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, an SFU may include a texture unit configured to perform texture map filtering operations. In one embodiment, a texture unit is configured to load texture maps (e.g., a 2D texture array) from memory 404 and sample these texture maps to produce sampled texture values ​​for use by a shader program executed by the processing unit. In one embodiment, the texture maps are stored in shared memory that may include or contain an L1 cache. The texture units use mip maps (e.g., texture maps with varying levels of detail) to implement texture operations such as filtering. In one embodiment, each processing unit includes two texture units.

[0147] Each processing unit also includes N Load Memory Units (LSUs) that implement load and store operations between shared memory and the register file. Each processing unit includes an interconnect network that connects each of the cores to the register file and connects the LSUs to the register file and the shared memory. In one embodiment, the interconnect network is a cross switch that can be configured to connect any of the cores to any register in the register file and connect the LSUs to memory locations in the register file and the shared memory.

[0148] Shared memory is an on-chip memory array that allows data storage and communication between processing units and between threads within a processing unit. In one embodiment, the shared memory includes 128KB of storage capacity and is located on the path from each of the processing units to memory partition unit 480. The shared memory can be used for caching reads and writes. One or more of the shared memory, L1 cache, L2 cache, and memory 404 serve as a backup cache.

[0149] Combining data caching and shared memory functionality into a single memory block provides optimal overall performance for both types of memory access. This capacity can be used as a cache by programs that do not use shared memory. For example, if shared memory is configured to use half its capacity, then texture and load / store operations can use the remaining capacity. Integration into shared memory allows it to function as a high-throughput conduit for streaming data, while providing high-bandwidth and low-latency access to frequently reused data.

[0150] When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics computing. Specifically, bypassing fixed-function graphics processing units (GPUs) creates a much simpler programming model. In this general-purpose parallel computing configuration, the work allocation unit 425 directly dispatches and assigns thread blocks to processing units within the GPC 450. Threads execute the same program using a unique thread ID during computation to ensure that each thread uses the executor program and the processing unit performing the computation, the shared memory for communication between threads, and the LSU (Local Subsystem for Memory) for reading and writing global memory via the shared memory and memory partitioning unit 480 to generate unique results. When configured for general-purpose parallel computing, processing units can also write commands, which the scheduler unit 420 can use to start new work on the processing unit.

[0151] Each of the PPUs 400 may include one or more processing cores and / or components thereof, such as a Tensor Core (TC), a Tensor Processing Unit (TPU), a Pixel Vision Core (PVC), a Ray Tracing (RT) Core, a Vision Processing Unit (VPU), a Graphics Processing Cluster (GPC), a Texture Processing Cluster (TPC), a Streaming Multiprocessor (SM), a Tree Traversal Unit (TTU), an Artificial Intelligence Accelerator (AIA), a Deep Learning Accelerator (DLA), an Arithmetic Logic Unit (ALU), an Application-Specific Integrated Circuit (ASIC), a Floating-Point Unit (FPU), Input / Output (I / O) Elements, Peripheral Component Interconnect (PCI) or Peripheral Component Interconnect High Speed ​​(PCIe) Elements and / or the like, and / or be configured to perform its functions.

[0152] The PPU 400 may be included in desktop computers, laptop computers, tablet computers, servers, supercomputers, smartphones (e.g., wireless, handheld devices), personal digital assistants (PDAs), digital cameras, vehicles, head-mounted displays, handheld electronic devices, and the like. In one embodiment, the PPU 400 is implemented on a single semiconductor substrate. In another embodiment, the PPU 400 is included in a system-on-a-chip (SoC) along with one or more other devices such as an additional PPU 400, memory 404, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), and the like.

[0153] In one embodiment, PPU 400 may be included on a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on a desktop computer's motherboard. In yet another embodiment, PPU 400 may be an integrated graphics processing unit (iGPU) or a parallel processor included in the chipset of the motherboard. In yet another embodiment, PPU 400 may be implemented in reconfigurable hardware. In yet another embodiment, a portion of PPU 400 may be implemented in reconfigurable hardware.

[0154] Exemplary computing system

[0155] As developers expose and leverage more parallelism in applications such as artificial intelligence computing, systems with multiple GPUs and CPUs are being used across various industries. High-performance GPU-accelerated systems with tens to thousands of compute nodes are being deployed in data centers, research facilities, and supercomputers to solve increasingly complex problems. With the increasing number of processing devices within high-performance systems, communication and data transmission mechanisms need to be scaled to support the increased bandwidth.

[0156] Figure 5A According to one embodiment, the use Figure 4 A conceptual diagram of a processing system 500 implemented by a PPU 400. The processing system 500 includes a CPU 530, a switch 510, multiple PPUs 400, and various memories 404.

[0157] The NVLink 410 provides a high-speed communication link between each of the PPUs 400. Although Figure 5B The diagram illustrates a specific number of NVLink 410 and interconnect 402 connections, but the number of connections to each PPU 400 and CPU 530 can vary. Switch 510 forms an interface between interconnect 402 and CPU 530. PPU 400, memory 404, and NVLink 410 can reside on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols to interface between various connections and / or links.

[0158] In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between each PPU 400 and CPU 530, and switch 510 forms an interface between interconnect 402 and each PPU 400. PPU 400, memory 404, and interconnect 402 may reside on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 402 provides one or more communication links between each PPU 400 and CPU 530, and switch 510 uses NVLink 410 to form an interface between each PPU 400 to provide one or more high-speed communication links between PPUs 400. In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between PPU 400 and CPU 530 via switch 510. In yet another embodiment (not shown), interconnect 402 directly provides one or more communication links between each PPU 400. One or more of the NVLink 410 high-speed communication links can be implemented as physical NVLink interconnects or on-chip or die-on interconnects using the same protocol as the NVLink 410.

[0159] In the context of this specification, a single semiconductor platform can refer to a single semiconductor-based integrated circuit fabricated on a bare die or chip. It should be noted that the term "single semiconductor platform" can also refer to a multi-chip module with increased connectivity, simulating on-chip operation and representing a significant improvement over conventional bus implementations. Of course, the various circuits or devices can also be located individually within the semiconductor platform or in various combinations thereof, as desired by the user. Alternatively, the parallel processing module 525 can be implemented as a circuit board substrate, and each PPU 400 and / or memory 404 can be a packaged device. In one embodiment, the CPU 530, switch 510, and parallel processing module 525 reside on a single semiconductor platform.

[0160] In one embodiment, the signaling rate of each NVLink 410 is 20-25 gigabits per second, and each PPU400 includes six NVLink 410 interfaces (e.g., Figure 5A As shown, each PPU 400 includes five NVLink 410 interfaces. Each NVLink 410 provides a data transfer rate of 25 gigabits per second in each direction, and the six links provide 400 gigabits per second. NVLink 410 can be used as follows: Figure 5A It is used exclusively for PPU-to-PPU communication, or for a combination of PPU-to-PPU and PPU-to-CPU when the CPU 530 also includes one or more NVLink 410 interfaces.

[0161] In one embodiment, NVLink 410 allows direct load / store / atomic access from CPU 530 to memory 404 of each PPU 400. In one embodiment, NVLink 410 supports coherent operation, allowing data read from memory 404 to be stored in the cache hierarchy of CPU 530, reducing cache access latency of CPU 530. In one embodiment, NVLink 410 includes support for Address Translation Service (ATS), allowing PPU 400 to directly access page tables within CPU 530. One or more of NVLink 410 may also be configured to operate in low-power mode.

[0162] Figure 5B An exemplary system 565 is illustrated in which various architectures and / or functions of various prior embodiments can be implemented. As shown, a system 565 is provided that includes at least one central processing unit 530 connected to a communication bus 575. The communication bus 575 may directly or indirectly couple one or more of the following devices: main memory 540, network interface 535, CPU 530, display device 545, input device 560, switch 510, and parallel processing system 525. The communication bus 575 may be implemented using any suitable protocol and may represent one or more links or buses, such as address bus, data bus, control bus, or combinations thereof. The communication bus 575 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect High Speed ​​(PCIe) bus, HyperTransport, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 530 may be directly connected to main memory 540. Furthermore, CPU 530 can be directly connected to parallel processing system 525. In cases where direct or point-to-point connections exist between components, communication bus 575 may include a PCIe link to implement that connection. In these examples, the PCI bus need not be included in system 565.

[0163] Although using lines Figure 5BThe different blocks are shown connected via a communication bus 575, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component such as a display device 545 can be considered an I / O component, such as an input device 560 (e.g., if the display is a touchscreen). As another example, the CPU 530 and / or the parallel processing system 525 may include memory (e.g., main memory 540 may represent storage devices other than the parallel processing system 525, the CPU 530, and / or other components). In other words, Figure 5B The term "computing device" is merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are expected to fall under [the relevant category]. Figure 5B Within the scope of computing devices.

[0164] System 565 also includes main memory 540. Control logic (software) and data are stored in main memory 540, which can take the form of a variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by system 565. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media can include computer storage media and communication media.

[0165] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory 540 may store computer-readable instructions such as an operating system (e.g., representing programs and / or program elements). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic cartridges, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that may be used to store desired information and that can be accessed by system 565. When used herein, computer storage media does not include the signal itself.

[0166] Computer storage media may contain computer-readable instructions, data structures, program modules, or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information transport medium. The term "modulated data signal" may refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0167] When executed, the computer program enables system 565 to perform various functions. CPU 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Each of CPUs 530 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing numerous software threads simultaneously. Depending on the type of system 565 implemented, CPU 530 may include any type of processor and may include different types of processors (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of system 565, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, system 565 may include one or more CPUs 530.

[0168] In addition to or alternatively to CPU 530, parallel processing module 525 may be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Parallel processing module 525 may be used by system 565 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, parallel processing module 525 may be used for general-purpose computing on a GPU (GPGPU). In embodiments, CPU 530 and / or parallel processing module 525 may execute any combination of the methods, processes, and / or portions thereof, discretely or jointly.

[0169] System 565 also includes input device 560, parallel processing system 525, and display device 545. Display device 545 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Display device 545 may receive data from other components (e.g., parallel processing system 525, CPU 530, etc.) and output that data (e.g., images, video, sound, etc.).

[0170] Network interface 535 enables system 565 to be logically coupled to other devices, including input device 560, display device 545, and / or other components, some of which may be embedded (e.g., integrated into) system 565. Illustrative input device 560 includes microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. Input device 560 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological input. In some instances, input can be transmitted to appropriate network elements for further processing. NUI can implement voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head-eye tracking, and touch recognition associated with the display of system 565 (described in more detail below). System 565 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. In addition, system 565 may include an accelerometer or gyroscope that allows motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by system 565 to render immersive augmented reality or virtual reality.

[0171] Furthermore, system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable network, etc.) via network interface 535 for communication purposes. System 565 can be included in a distributed network and / or cloud computing environment.

[0172] Network interface 535 may include one or more receivers, transmitters, and / or transceivers, enabling system 565 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. Network interface 535 may be implemented as a network interface controller (NIC) including one or more data processing units (DPUs) to perform operations such as (e.g., but not limited to) packet parsing and accelerating network processing and communication. Network interface 535 may include components and functions to enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0173] System 565 may also include an auxiliary storage device (not shown). The auxiliary storage device includes, for example, a hard disk drive and / or a removable storage drive representing a floppy disk drive, magnetic tape drive, compact disc drive, digital versatile disc (DVD) drive, recording device, Universal Serial Bus (USB) flash memory. The removable storage drive reads from and / or writes to the removable storage unit in a well-known manner. System 565 may also include a hard-wired power supply, a battery power supply, or a combination thereof (not shown). This power supply can supply power to System 565 to enable the components of System 565 to operate.

[0174] Each of the aforementioned modules and / or devices may even reside on a single semiconductor platform to form system 565. Alternatively, various different modules may be placed individually or located in various combinations of semiconductor platforms as desired by the user. Although various different embodiments have been described above, it should be understood that they are given by way of example only and without limitation. Therefore, the breadth and scope of preferred embodiments should not be limited to any of the exemplary embodiments described above, but should be defined only by the following claims and their equivalents.

[0175] Example network environment

[0176] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage devices (NAS), other back-end devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 5A Processing system 500 and / or Figure 5B Implemented on one or more instances of the exemplary system 565, for example, each device may include similar components, features and / or functions of the processing system 500 and / or the exemplary system 565.

[0177] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or a combination of both. A network can include multiple networks or a network of networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks—such as the Internet, and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (along with other components) can provide wireless connectivity.

[0178] A compatible network environment may include one or more peer-to-peer network environments—in which case the server may not be included in the network environment—and one or more client-server network environments—in which case one or more servers may be included in the network environment. In a peer-to-peer network environment, the functionality described herein regarding the server can be implemented on any number of client devices.

[0179] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework of one or more applications supporting a software layer and / or an application layer. The software or application may include web-based service software or applications, respectively. In embodiments, one or more client devices may use the web-based service software or application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software web application framework type that can be used for large-scale data processing (e.g., "big data").

[0180] A cloud-based network environment can provide cloud computing and / or cloud storage to implement the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., a central or core server in one or more data centers, which may be distributed across states, regions, countries, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, then the core server can assign at least a portion of the functions to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0181] Client devices may include Figure 5B Example processing system 500 and / or Figure 5C At least some of the components, features, and functions of the exemplary system 565. For example and without limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming equipment or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these defined devices, or any other suitable device.

[0182] Machine Learning

[0183] Deep neural networks (DNNs) developed on processors such as the PPU 400 have been used in a wide variety of use cases, from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technique that models the neural learning process of the human brain, which continuously learns, becomes smarter, and delivers more accurate results faster over time. Just as a child is initially taught by adults to correctly identify and classify various shapes, eventually becoming able to identify shapes without any guidance, a deep learning or neural learning system needs to be trained in object recognition and classification so that it becomes smarter and more efficient at identifying basic objects, occluded objects, and so on, while also attaching context to objects.

[0184] At its simplest level, neurons in the human brain receive various inputs, assigning a level of importance to each of these inputs, and the output is passed to other neurons to make a response. Artificial neurons, or perceptrons, are the most basic model of neural networks. In one example, a perceptron can receive one or more inputs representing various features of objects that the perceptron is being trained to recognize and classify, and each of these features is assigned a weight based on its importance in defining the shape of the object.

[0185] Deep neural network (DNN) models consist of multiple layers of numerous connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained on massive amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into different segments and searches for basic patterns such as lines and angles. The second layer assembles these lines to find higher-level patterns such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the final few layers generate labels for the input image that identify the model of a specific car brand.

[0186] Once trained, a DNN can be deployed and used to identify and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include identifying handwritten digits on a check deposited into an ATM, identifying images of friends in a photograph, delivering movie recommendations to over 50 million users, identifying and classifying different types of cars, pedestrians, and road hazards in self-driving cars, or translating human language in real time.

[0187] During training, data flows through the DNN in the forward propagation phase until a prediction is produced indicating the label corresponding to the input. If the neural network does not correctly label the input, the error between the correct label and the predicted label is analyzed, and the weights are adjusted for each feature during the backpropagation phase until the DNN correctly labels the input as well as other inputs in the training dataset. Training complex neural networks requires significant parallel computing power, including floating-point multiplication and addition supported by a PPU400. Inference is less computationally intensive than training and is a latency-sensitive process where the trained neural network is applied to new inputs it has not seen before for tasks such as image classification, sentiment detection, label recommendation, language recognition and translation, and typically infers new information.

[0188] Neural networks rely heavily on matrix operations, and complex, multi-layered networks require massive amounts of floating-point performance and bandwidth for both efficiency and speed. Leveraging thousands of processing cores optimized for matrix operations and delivering tens to hundreds of TFLOPS of performance, the PPU 400 is a computing platform capable of providing the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0189] Furthermore, images generated using one or more of the techniques disclosed herein can be used to train, test, or certify DNNs for recognizing real-world objects and environments. Such images can include driveways, factories, buildings, urban environments, rural environments, humans, animals, and any other physical objects or scenes of real-world environments. Such images can be used to train, test, or certify DNNs used in machines or robots to manipulate, process, or modify real-world physical objects. Additionally, such images can be used to train, test, or certify DNNs used in autonomous vehicles to navigate and move vehicles in the real world. Furthermore, images generated using one or more of the techniques disclosed herein can be used to communicate information to users of such machines, robots, and vehicles.

[0190] Figure 5C Components of an example system 555, which can be used to train and utilize machine learning according to at least one embodiment, are illustrated. As will be discussed, various components can be provided by a single computing system or various combinations of computing devices and resources that can be controlled by a single entity or more entities. Furthermore, aspects can be triggered, initiated, or requested by different entities. In at least one embodiment, the training of the neural network can be guided by a vendor associated with vendor environment 506, while in at least one embodiment, training can be requested by a customer or other user who can access the vendor environment through client device 502 or other such resources. In at least one embodiment, training data (or data to be analyzed by the trained neural network) can be provided by a vendor, user, or third-party content provider 524. In at least one embodiment, client device 502 can be, for example, a vehicle or object to be navigated on behalf of a user, who can submit requests and / or receive instructions that aid in device navigation.

[0191] In at least one embodiment, a request can be submitted via at least one network 504 for receipt by a vendor environment 506. In at least one embodiment, the client device can be any suitable electronic and / or computing device that enables a user to generate and send such requests, such as, but not limited to, desktop computers, laptop computers, computer servers, smartphones, tablets, game consoles (portable or otherwise), computer processors, computing logic, and set-top boxes. One or more networks 504 can include any suitable network for transmitting requests or other such data, such as the Internet, intranet, Ethernet, cellular network, local area network (LAN), wide area network (WAN), personal area network (PAN), self-organizing network providing direct wireless connectivity between peers, etc.

[0192] In at least one embodiment, a request may be received at interface layer 508, which in this example may forward data to training and inference manager 532. Training and inference manager 532 may be a system or service including hardware and software for managing services and requests corresponding to data or content. In at least one embodiment, training and inference manager 532 may receive a request to train a neural network and may provide data for the request to training module 512. In at least one embodiment, if the request is not specified, training module 512 may select an appropriate model or neural network to use and may train the model using the associated training data. In at least one embodiment, training data may be a batch of data stored in training data repository 514, received from client device 502, or obtained from third-party vendor 524. In at least one embodiment, training module 512 may be responsible for training the data. The neural network may be any suitable network, such as a recurrent neural network (RNN) or convolutional neural network (CNN). Once the neural network is trained and successfully evaluated, the trained neural network may be stored in, for example, model repository 516, which may store different models or networks for users, applications, or services, etc. In at least one embodiment, there may be multiple models for a single application or entity, which can be utilized based on multiple different factors.

[0193] In at least one embodiment, at a subsequent point in time, a request for content (e.g., path determination) or data that is at least partially determined or influenced by a trained neural network can be received from client device 502 (or another such device). This request may include, for example, input data to be processed using the neural network to obtain one or more inference or other output values, classifications, or predictions. Alternatively, in at least one embodiment, the input data may be received by interface layer 508 and directed to inference module 518, although different systems or services may also be used. In at least one embodiment, if not already locally stored in inference module 518, inference module 518 may obtain a suitably trained network, such as a trained deep neural network (DNN) as discussed herein, from model repository 516. Inference module 518 may provide data as input to the trained network, which may then generate one or more inferences as outputs. This may, for example, include the classification of instances of input data. In at least one embodiment, the inference may then be transmitted to client device 502 for display to a user or for other communication with the user. In at least one embodiment, user context data may also be stored in a user context data repository 522, which may include data about the user that can be used as network input to generate inference or determine data returned to the user after obtaining an instance. In at least one embodiment, relevant data, including at least some of the input or inference data, may also be stored in a local database 534 for processing future requests. In at least one embodiment, the user may use account information or other information to access resources or functions of the vendor environment. In at least one embodiment, user data may also be collected and used to further train the model, if permitted and available, to provide more accurate inference for future requests. In at least one embodiment, requests to a machine learning application 526 executed on a client device 502 may be received via a user interface, and the results may be displayed via the same interface. The client device may include resources such as a processor 528 and a memory 562 for generating requests and processing results or responses, and at least one data storage element 552 for storing data for the machine learning application 526.

[0194] In at least one embodiment, processor 528 (or the processor of training module 512 or inference module 518) will be a central processing unit (CPU). However, as mentioned above, resources in such an environment can utilize GPUs to process data for at least some types of requests. GPUs such as the PPU 300 have thousands of cores and are designed to handle large amounts of parallel workloads, thus becoming popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building allows for faster training of larger, more complex models, offline prediction generation means that request-time input features cannot be used, or predictions must be generated for all features and stored in a lookup table for real-time service requests. If the deep learning framework supports CPU mode and the model is small and simple enough that the feedforward can be performed on the CPU with reasonable latency, then a service on a CPU instance can host the model. In this case, training can be done offline on the GPU and inference can be performed in real-time on the CPU. If the CPU approach is not feasible, the service can run on a GPU instance. However, due to the different performance and cost characteristics of GPUs compared to CPUs, running a service that offloads runtime algorithms to the GPU may require it to be designed differently from a CPU-based service.

[0195] In at least one embodiment, video data can be provided from client device 502 for enhancement in vendor environment 506. In at least one embodiment, the video data can be processed for enhancement on client device 502. In at least one embodiment, the video data can be streamed from third-party content provider 524 and enhanced by third-party content provider 524, vendor environment 506, or client device 502. In at least one embodiment, video data can be provided from client device 502 for use as training data in vendor environment 506.

[0196] In at least one embodiment, supervised and / or unsupervised training may be performed by client device 502 and / or vendor environment 506. In at least one embodiment, a set of training data 514 (e.g., classified or labeled data) is provided as input for use as training data. In at least one embodiment, the training data may include instances of at least one type of object for which the neural network is to be trained, and information identifying that object type. In at least one embodiment, the training data may include a set of images, each image including a representation of an object of a type, wherein each image also includes, or is associated with, tags, metadata, classification, or other information identifying or identifying the type of object represented in the corresponding image. Various other types of data may also be used as training data, which may include text data, audio data, video data, and so on. In at least one embodiment, training data 514 is provided as training input to training module 512. In at least one embodiment, training module 512 may be a system or service including hardware and software, such as one or more computing devices executing a training application for training a neural network (or other model or algorithm, etc.). In at least one embodiment, training module 512 receives instructions or requests indicating the type of model to be used for training. In at least one embodiment, the model can be any suitable statistical model, network, or algorithm useful for such a purpose, which may include artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, etc. In at least one embodiment, training module 512 may select an initial model or other untrained models from an appropriate repository and train the model using training data 514 to generate a trained model (e.g., a trained deep neural network) that can be used to classify similar types of data or generate other such inference. In at least one embodiment where training data is not used, an initial model can still be selected to train on the input data of each training module 512.

[0197] In at least one embodiment, the model can be trained in several different ways, which may depend in part on the type of model chosen. In at least one embodiment, a training dataset can be provided to a machine learning algorithm, wherein the model is a model artifact created through a training process. In at least one embodiment, each instance of the training data contains the correct answer (e.g., classification) that may be referred to as the target or target attribute. In at least one embodiment, the learning algorithm finds patterns in the training data that map input data attributes to the target—the answer to be predicted—and the machine learning model is the output that captures these patterns. In at least one embodiment, the machine learning model can then be used to obtain predictions for new data without a specified target.

[0198] In at least one embodiment, the training and inference manager 532 may select from a set of machine learning models, including binary classification, multi-class classification, generative, and regression models. In at least one embodiment, the type of model to be used may depend at least in part on the type of target to be predicted.

[0199] Images generated using one or more of the techniques disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device may be directly coupled to the system or processor that generates or renders the image. In other embodiments, the display device may be indirectly coupled to the system or processor, for example, via a network. Examples of such networks include the Internet, mobile telecommunications networks, Wi-Fi networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, images generated by the system or processor can be streamed to the display device over the network. Such streaming allows, for example, video games or other applications that render images to execute on servers, data centers, or cloud-based computing environments, and the rendered images are transmitted and displayed on one or more user devices (e.g., computers, video game consoles, smartphones, other mobile devices, etc.) physically separate from the server or data center. Therefore, the techniques disclosed herein can be applied to enhance streamed images and services that stream images, such as NVIDIA GeForce Now (GFN), Google Stadia, etc.

[0200] Example game streaming system

[0201] Figure 6 This is a schematic diagram of an example system 605 for a game streaming system according to some embodiments of the present disclosure. Figure 6 Including game server 603 (which may include and Figure 5A Example processing system 500 and / or Figure 5B (Similar components, features and / or functions to exemplary system 565), client 604 (which may include similar ... Figure 5A Example processing system 500 and / or Figure 5B The exemplary system 565 has similar components, features, and / or functions to the network 606 (which may be similar to the network described herein). In some embodiments of this disclosure, system 605 may be implemented.

[0202] In system 605, for a game session, client device 604 can simply receive input data in response to input from an input device, send the input data to game server 603, receive encoded display data from game server 603, and display the display data on monitor 624. In this way, computationally intensive computation and processing are offloaded to game server 603 (e.g., rendering of the game session's graphics output, especially ray or path tracing, is performed by the GPU of game server 603). In other words, the game session is streamed from game server 603 to client device 604, thereby reducing the demands on client device 604 for graphics processing and rendering.

[0203] For example, regarding the instantiation of a game session, client device 604 can display frames of the game session on display 624 based on display data received from game server 603. Client device 604 can receive input from one of the input devices and generate input data in response. Client device 604 can send the input data to game server 603 via communication interface 621 and through network 606 (e.g., the Internet), and game server 603 can receive the input data via communication interface 618. CPU can receive the input data, process the input data, and send the data to GPU, which causes GPU to generate a rendering of the game session. For example, the input data can represent the movement of a user character in the game, such as firing a weapon, reloading, passing a ball, turning a vehicle, etc. Rendering component 612 can render the game session (e.g., representing the result of the input data), and rendering capture component 614 can capture the rendering of the game session as display data (e.g., image data as frames of the captured game session rendering). The rendering of a game session may include lighting and / or shadow effects computed using one or more parallel processing units of the game server 603 (e.g., a GPU, which may further employ one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques). The encoder 616 can then encode the display data to generate encoded display data, which can be sent to the client device 604 via the network 606 through the communication interface 618. The client device 604 can receive the encoded display data via the communication interface 621, and the decoder 622 can decode the encoded display data to generate display data. The client device 604 can then display the display data via the display 624.

[0204] It should be noted that the techniques described herein can be contained in executable instructions stored in a computer-readable medium for use by or in conjunction with a processor-based instruction execution machine, system, apparatus, or device. Those skilled in the art will appreciate that, for some embodiments, various different types of computer-readable media may be included for storing data. When used herein, “computer-readable medium” includes one or more of any suitable media for storing executable instructions of a computer program, such that an instruction execution machine, system, apparatus, or device can read (or retrieve) the instructions from the computer-readable medium and execute those instructions to implement the described embodiments. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); flash memory devices; and optical storage devices, including portable compact discs (CDs), portable digital video discs (DVDs), and the like.

[0205] It should be understood that the arrangement of components shown in the accompanying drawings is for illustrative purposes, and other arrangements are possible. For example, one or more of the elements described herein may be implemented wholly or partially as electronic hardware components. Other elements may be implemented in software, hardware, or a combination of software and hardware. Moreover, some or all of these other elements may be combined, some may be omitted entirely, and additional components may be added while still achieving the functionality described herein. Therefore, the subject matter described herein can be implemented in many different variations, and all such variations are contemplated to be within the scope of the claims.

[0206] To facilitate understanding of the topics described herein, many aspects are described in sequence of actions. Those skilled in the art will recognize that various actions can be performed by dedicated circuitry or circuit systems, by program instructions executed by one or more processors, or by a combination of both. The description of any sequence of actions herein is not intended to imply that a particular order in which the actions described for execution must be followed. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context.

[0207] In the context of describing the subject matter (especially in the context of the claims below), the use of the terms “a,” “an,” “this,” and similar designations should be interpreted to cover both the singular and plural, unless otherwise specified herein or obviously contradicted by the context. The use of the term “at least one” (e.g., at least one of A and B) followed by a list of one or more items should be interpreted to mean one item selected from the listed items (A or B), or any combination of two or more of the listed items (A and B), unless otherwise specified herein or obviously contradicted by the context. Furthermore, the foregoing description is for illustrative purposes only and not for limiting purposes, as the scope of protection sought is defined by the claims set forth thereafter with their equivalents. The use of any and all example or exemplary language provided herein (e.g., “such as”) is intended merely to better illustrate the subject matter and does not constitute a limitation on the scope of the subject matter, unless otherwise stated. The use of “based on,” and other similar phrases indicating conditions leading to the result, in both the claims and the written description, is not intended to exclude any other conditions leading to that result. The language in the description should not be interpreted as indicating that any unclaimed element is essential for the implementation of the claimed invention.

Claims

1. A communication device comprising a receiver for communicating via a channel, wherein: The communication device further includes a processor configured to at least partially implement a first neural network to perform a task associated with the receiver, wherein the first neural network includes a fully fused neural network, and wherein the fully fused neural network refers to a neural network implemented in a single core on the processor; the core is designed such that only slow memory accesses read the input of the neural network from global memory and write the output of the neural network to global memory. All intermediate values ​​are operated on in high-speed on-chip memory; as well as The first neural network is periodically retrained using a training dataset, which includes multiple data frames received by the receiver through the channel.

2. The communication device as claimed in claim 1, wherein, The processor includes a parallel processing unit configured to implement the fully fused neural network using at least one or more tensor cores.

3. The communication device as claimed in claim 1, wherein, The first neural network includes a multilayer perceptron (MLP), which includes at least one hidden layer, an activation layer, and an output layer.

4. The communication device as claimed in claim 1, wherein, The first neural network is configured to generate a log-likelihood ratio (LLR) of transmitted bits based on the equalization symbols of Orthogonal Frequency Division Multiplexing (OFDM) frames received by the receiver through the channel, wherein the OFDM frames are generated by... N Each subcarrier and each subcarrier K It consists of several symbols.

5. The communication device as claimed in claim 4, wherein, The first neural network includes a multilayer perceptron (MLP) configured to receive the equalization symbols of the OFDM frame. Noise variance estimation associated with each resource element of the OFDM frame and the index tuple associated with each resource element. Position encoding As input, and generated for each equalization symbol of the OFDM frame. m One LLR, of which It is equal to the order of the constellation used in the Quadrature Amplitude Modulation (QAM) modulation and coding scheme MCS for communication through the channel.

6. The communication device as claimed in claim 4, wherein, The first neural network is periodically retrained based on the training dataset. For each OFDM frame in a plurality of OFDM frames, the training dataset comprises a set of tuples, each tuple consisting of an equalization symbol for a specific resource element of the OFDM frame and a corresponding valid bit of the equalization symbol. The corresponding valid bit is generated by encoding information bits using forward error correction (FEC), and the information bit is derived by the channel decoder based on the LLR and based on the last valid bit. C The cyclic redundancy check (CRC) of 1 bit was confirmed as valid.

7. The communication device as claimed in claim 6, wherein, The first neural network each f Each OFDM frame is periodically retrained using stochastic gradient descent, where f is a positive integer.

8. The communication device as claimed in claim 1, wherein, The first neural network is configured to be based on the first received by the receiver through the channel. t The channel matrix estimate is generated by the receive resource grid of each Orthogonal Frequency Division Multiplexing (OFDM) frame, wherein the OFDM frame is composed of... N Each subcarrier and each subcarrier K It consists of several symbols.

9. The communication device as claimed in claim 1, wherein, The first neural network is configured to be based on the first received by the receiver through the channel. t The received resource grid of each Orthogonal Frequency Division Multiplexing (OFDM) frame generates an estimate of a three-dimensional tensor representing information bits, wherein the OFDM frame is generated by... N Each subcarrier and each subcarrier K It consists of several symbols.

10. The communication device as claimed in claim 1, wherein, The processor is further configured to at least partially implement a second neural network to perform a second task associated with the receiver, wherein the first neural network performs a demapping task and the second neural network performs a channel estimation task, and wherein the first neural network is configured to at least partially process the output of the second neural network as the input of the first neural network.

11. The communication device as claimed in claim 1, wherein, The communication device communicates with a second communication device via a channel, wherein the second communication device includes a transmitter and a second processor, the second processor being configured to at least partially implement a second neural network to perform tasks associated with the transmitter of the second communication device.

12. The communication device as claimed in claim 11, wherein, The transmitter of the second communication device is configured to send a known data sequence to the receiver of the communication device, wherein the first neural network is trained based on the known data and a binary cross-entropy loss between the known data and the predicted log-likelihood ratio (LLR) generated by the first neural network.

13. A communication system, comprising: A first communication device includes a receiver for communicating via a channel, wherein: The first communication device further includes a processor configured to at least partially implement a first neural network to perform a task associated with the receiver, wherein the first neural network comprises a fully fused neural network, and wherein the fully fused neural network refers to a neural network implemented in a single core on the processor; the core is designed such that only slow memory accesses read the input of the neural network from global memory and write the output of the neural network to global memory; all intermediate values ​​are operated on in high-speed on-chip memory; and The first neural network is periodically retrained using a training dataset, which includes multiple data frames received by the receiver through the channel.

14. The communication system as claimed in claim 13, wherein, The first neural network is configured to generate a log-likelihood ratio (LLR) of transmitted bits based on the equalization symbols of orthogonal frequency division multiplexing (OFDM) frames received by the receiver through the channel, wherein the OFDM frames are generated by... N Each subcarrier and each subcarrier K It consists of several symbols.

15. The communication system as claimed in claim 14, wherein, The first neural network includes a multilayer perceptron (MLP) configured to receive equalization symbols of the OFDM frames. Noise variance estimate associated with each resource element of the OFDM frame and the index tuple associated with each resource element. Position encoding As input, and generated for each equalization symbol of the OFDM frame. m One LLR, of which It is equal to the order of the constellation used in the Quadrature Amplitude Modulation (QAM) modulation and coding scheme MCS for communication through the channel.

16. The communication system as claimed in claim 15, wherein, The first neural network is periodically retrained based on training data for each of a plurality of OFDM frames. The training data includes a set of tuples, each tuple consisting of an equalization symbol for a specific resource element of the OFDM frame and the corresponding valid bit of the equalization symbol. The corresponding valid bit is generated by encoding information bits using forward error correction (FEC), and the information bit is derived by the channel decoder based on the LLR and based on the last valid bit in the information bits. C The cyclic redundancy check (CRC) of 1 bit was confirmed as valid.

17. The communication system of claim 16, wherein, The first neural network each f Each OFDM frame is periodically retrained using stochastic gradient descent, where f is a positive integer.

18. The communication system as claimed in claim 13, wherein, The first neural network is configured to be based on the first received by the receiver through the channel. t The channel matrix estimate is generated by the receive resource grid of each Orthogonal Frequency Division Multiplexing (OFDM) frame, wherein the OFDM frame is composed of... N Each subcarrier and each subcarrier K It consists of several symbols.

19. The communication system as claimed in claim 13, wherein, The first neural network is configured to be based on the first received by the receiver through the channel. t The received resource grid of each Orthogonal Frequency Division Multiplexing (OFDM) frame generates an estimate of a three-dimensional tensor representing information bits, wherein the OFDM frame is generated by... N Each subcarrier and each subcarrier K It consists of several symbols.

20. The communication system of claim 13, further comprising a second communication device including a transmitter, wherein the second communication device further includes a second processor configured to at least partially implement a second neural network to perform a task associated with the transmitter of the second communication device, wherein the transmitter of the second communication device is configured to transmit a known data sequence to a receiver of the communication device, and wherein the first neural network is trained based on the known data and a binary cross-entropy loss between the known data and the predicted log-likelihood ratio (LLR) generated by the first neural network.

21. A method for training a neural network, comprising: A training dataset is obtained based on one or more frames received by a receiver of a communication device. The training data includes multiple tuples, each tuple comprising an equalization symbol for the one or more frames and a corresponding valid bit sequence for each equalization symbol. The valid bit sequence is applied to the last information bit-based frame. C The forward error correction of the information bits of the Cyclic Redundancy Check (CRC) is encoded; and A first neural network is periodically retrained, the first neural network being configured to perform a task associated with the receiver using the training dataset for multiple frames, wherein the first neural network comprises a fully fused neural network, and wherein the fully fused neural network refers to a neural network implemented in a single core on a processor; the core is designed such that only slow memory accesses read the input of the neural network from global memory and write the output of the neural network to global memory; all intermediate values ​​are operated on in high-speed on-chip memory.

22. The method of claim 21, wherein, The retraining is in B Executed on the next iteration, each iteration and from A batch of index tuples that are deterministically or randomly selected from the data. Correspondingly, among them It is all index tuples corresponding to the resource elements carrying valid data in the frame. A set of.

23. A method for end-to-end training of an autoencoder, wherein, The autoencoder includes a first neural network corresponding to a transmitter of a first communication device and a second neural network corresponding to a receiver of a second communication device, and the method includes: Initialize the parameters of the first neural network and the second neural network; A known data sequence is transmitted from the transmitter to the receiver; In response to the known data sequence transmitted, a set of equalization symbols is generated via the receiver; The second neural network is trained based on the known data sequence and the binary cross-entropy loss between the predicted log-likelihood ratio (LLR) generated by the second neural network based on the set of balanced symbols. The second neural network is a fully fused neural network, and the fully fused neural network refers to a neural network implemented in a single core on a processor. The core is designed so that only slow memory accesses read the neural network's input from global memory and write the neural network's output to global memory; all intermediate values ​​are operated on in high-speed on-chip memory. A second known data sequence is transmitted via the transmitter, wherein the transmitter is configured to perturb the output of the transmitter according to perturbation noise; The transmitter receives a feedback loss signal from the receiver; Calculate the second loss signal based on the feedback loss signal and the disturbance noise; and The first neural network is trained based on the second loss signal.

24. The method of claim 23, wherein, The training of the second neural network and the first neural network is repeated periodically according to the termination criteria.

Citation Information

Patent Citations

  • Fully-fused neural network execution

    US11631210B2

  • Multiresolution hash encoding for neural networks

    US12711373B2

  • An online learning method of an artificial intelligence assisted OFDM receiver

    CN109861942A

  • Techniques for transforming serial program code into kernels for execution on a parallel processor

    US20190278574A1