Using the signal-to-noise ratio to select a neural network
By selecting a neural network based on signal properties like SNR and resource blocks, the system addresses inaccuracies in channel estimation, enhancing wireless communication performance and reducing inefficiencies.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2025-01-17
- Publication Date
- 2026-06-11
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field
[0001] At least one embodiment relates to a processor that uses software to select a neural network, the selected neural network being used to perform wireless signal operations (for example, channel estimation). For example, a processor comprising circuitry uses signal-to-noise ratio (SNR) values of a wireless signal to select one or more neural networks to generate channel estimates. BACKGROUND
[0002] A base station or user device (such as a mobile phone) can transmit a wireless signal, but due to certain conditions, the wireless signal may contain noise, which is then received. A method for estimating noise in a channel involves channel estimation, which includes analyzing the received signal to determine the channel's characteristics, such as amplitude, phase, and interference patterns. By using channel estimation techniques, the noise component can be identified and quantified, enabling adjustments to improve signal quality or optimize subsequent transmissions. These adjustments might include changing the transmit power, applying adaptive filtering, or modifying modulation schemes to mitigate the effects of noise.While a neural network can be used for channel estimation, it can produce inaccurate channel estimates, leading to suboptimal or failed wireless communication. Such inaccuracies can result in increased retransmissions, reduced throughput, or higher power consumption. Therefore, ensuring robust training and validation of neural network-based channel estimation models is crucial. There is a need to address at least this technical problem and implement further improvements. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a block diagram showing a system for selecting a neural network to perform wireless signal operations according to at least one embodiment; Fig. 2 shows a system for training and deploying neural networks according to at least one embodiment; Fig. Figure 3 is a block diagram showing an exemplary system for selecting a neural network to generate a channel estimate according to at least one embodiment; Fig. 4 shows an example of combining one or more neural networks to perform wireless signal operations according to at least one embodiment; Fig. Figure 5 is a process flow diagram showing the selection of a neural network that is based at least partially on one or more properties of a signal, according to at least one embodiment; Fig. Figure 6 shows a processor and instructions according to at least one embodiment; Fig. Figure 7 is a block diagram showing a driver and / or runtime comprising one or more libraries to provide one or more application programming interfaces (APIs) according to at least one embodiment; Fig. Figure 8 shows an example of a data center according to at least one embodiment; Fig. Figure 9 shows a system-on-a-chip (SOC) according to at least one embodiment; Fig. 10A shows a parallel processor according to at least one embodiment; Fig. Figure 10B shows a processing cluster according to at least one embodiment; Fig. 10C shows a graphics multiprocessor according to at least one embodiment; Fig. Figure 11 shows an accelerator according to at least one embodiment; Fig. 12A shows a central unit according to at least one embodiment; Fig. Figure 12B shows a core of the central processing unit in Fig. 12A according to at least one embodiment; Fig. Figure 13 shows a further accelerator according to at least one embodiment; Fig. Figure 14 shows a neuromorphic processor according to at least one embodiment; Fig. Figure 15 shows a supercomputer according to at least one embodiment; Fig. Figure 16 shows a further accelerator according to at least one embodiment; Fig. 17 shows another processor according to at least one embodiment; Fig. Figure 18 shows a further accelerator according to at least one embodiment; Fig. 19 shows a tensor processing unit according to at least one embodiment; Fig. Figure 20 shows a RISC-V compatible processor according to at least one embodiment; Fig. 21A and Fig. Figure 21B shows a speech processing unit according to at least one embodiment; Fig. Figure 22 shows a software stack of a programming platform according to at least one embodiment; Fig. 23 shows software supported by a programming platform according to at least one embodiment; Fig. 24 shows the compilation of code for execution on programming platforms according to Fig. 23 according to at least one embodiment; Fig. Figure 25 shows an example of an autonomous vehicle and its system architecture according to at least one embodiment; Fig. 26A shows an inference and / or training logic according to at least one embodiment; Fig. 26B shows an inference and / or training logic according to at least one embodiment; Fig. Figure 26C shows the training and use of a neural network according to at least one embodiment; Fig. Figure 27 shows a network for data communication within a 5G radio communication network according to at least one embodiment; Fig. Figure 28 shows a network architecture for a 5G LTE radio network according to at least one embodiment; Fig. Figure 29 is a diagram showing some basic functions of a mobile network / system operating according to the LTE and 5G principles, according to at least one embodiment; Fig. Figure 30 shows a radio access network that can be part of a 5G network architecture, according to at least one embodiment; Fig. Figure 31 shows an example of a 5G mobile communication system in which a variety of different types of devices are used, according to at least one embodiment; Fig. Figure 32 shows an example of a high-level system according to at least one embodiment; Fig. Figure 33 shows an architecture of a network system according to at least one embodiment; Fig. 34 shows example components of a device according to at least one embodiment; Fig. 35 shows example interfaces of a baseband circuit according to at least one embodiment; Fig. Figure 36 shows an example of an uplink channel according to at least one embodiment; Fig. 37 shows an architecture of a network system according to at least one embodiment; Fig. Figure 38 shows a control plane protocol stack according to at least one embodiment; and Fig. Figure 39 shows a user-level protocol stack according to at least one embodiment. DETAILED DESCRIPTION
[0003] The following description presents numerous specific details to provide a more thorough understanding of at least one embodiment. However, it is obvious to a person skilled in the art that the inventive concepts can also be implemented without one or more of these specific details.
[0004] Fig. Figure 1 is a block diagram showing a system 100 for selecting a neural network to perform wireless signal operations 116 according to at least one embodiment. In at least one embodiment, the system 100 comprises one or more processors 104, a memory 106, software 108 for selecting one or more neural networks, one or more neural networks, one or more selected neural networks 114, one or more wireless signal operations 116 (for example, equalization, channel estimation 116A, beamforming 116B, and / or signal mapping 116C), and / or one or more system components described herein. In at least one embodiment, the system 100 receives one or more input signals 118, and a processor 104 uses software 108 to perform neural network selection 110.In at least one embodiment, the system 100 selects one or more neural networks that are at least partially based on one or more properties of an input signal (for example, SNR, site-specific, cell-specific, antenna information or other information corresponding to a wireless signal used in 5G, 6G or other wireless signal communication).
[0005] In at least one embodiment, a technical problem arises in that a single neural network can be used to generate a channel estimate over an arbitrary signal-to-noise ratio (SNR) (for example, a large SNR range). In at least one embodiment, an SNR can vary based on the characteristics of an input signal 118; for example, high-power signals can have a different signal-to-noise ratio than low-power signals. In at least one embodiment, a problem arises in that a neural network is trained to receive every SNR (for example, a large range of SNRs, from low SNR to high SNR). In at least one embodiment, the use of a single neural network across all SNR values can lead to a neural network producing inaccurate results for a channel, since a neural network can overfit an SNR.In at least one embodiment, during the training of a neural network, a loss function can cause the weighting of the NN to adjust based on the magnitude of the loss, with signals with a low SNR producing a greater loss (e.g., larger changes in weighting) than signals with a higher SNR, which produce a smaller loss (e.g., smaller changes in weighting). In at least one embodiment, using a single neural network to generate a channel estimate would therefore be inaccurate, as it is trained to cover a large range of SNRs.
[0006] In at least one embodiment, software 108 executed by a processor selects a specific neural network from a group of neural networks to perform signal operations 116 (for example, channel estimation 116A, beamforming 116B, or signal mapping 116C) using properties of an input signal 118 to make its selection 110. In at least one embodiment, the software 108 first receives properties (for example, power, reference signal information, a number of resource blocks, or time allocations) of a signal as input 102 and estimates an expected signal-to-noise ratio (SNR) of an input signal 118. In at least one embodiment, the software 108 then selects a neural network from a group of neural networks using an estimated SNR.In at least one embodiment, a neural network 114 is selected using an estimated SNR for an input signal 118, using a small SNR range (for example, a small dB range for SNR) with which a neural network has been trained, wherein the estimated SNR lies within the range of a neural network. In at least one embodiment, the software 108 can, in some cases, select a neural network using an SNR property of an input signal 118 in combination with other signal properties 102A. In at least one embodiment, the software 108 selects a neural network that is specific to the properties of an input signal 118, thereby increasing the performance compared to a generalized model.In at least one embodiment, the system 100 uses the software 108 to select one of many neural networks to perform a channel estimate 116A using properties of an input signal 118, for example, using one or more SNR values to select one or more neural networks to generate one or more channel estimates. In at least one embodiment, the software uses two different inputs, which can also be called dimensions, to select a neural network. For example, one input can be the SNR of a received signal, and another dimension can be the number of PRBs associated with a user device.In at least one embodiment, the software executed by a processor uses interpolation when selecting a neural network, wherein the interpolation is based on several neural networks that have been trained for specific SNR ranges and PRB values.
[0007] In at least one embodiment, the software 108 for selecting a neural network receives one or more inputs 102, such as signal properties 102A and / or neural network properties 102B. In at least one embodiment, the input 102 of the system 100 comprises an input signal 118, signal properties 102A (for example, power, reference signal information, resource block information, a number of resource blocks, and / or a time allocation), neural network properties 102B (for example, neural network training information, types and / or extent of training received according to a signal property, layer information, expected latency, expected accuracy, and / or other neural network information described herein), telemetry data, one or more user inputs, information represented as data, and / or other inputs described herein.In at least one embodiment, an input 102 comprises an input signal 118, for example, a wireless signal from a user device (e.g., a mobile phone) and / or a wireless signal from a base station. In at least one embodiment, a receiver of a wireless signal comprises a base station and / or a user device (e.g., a mobile device, for example, a telephone). In at least one embodiment, one or more inputs 102 are transmitted by a signal to a processor 104. In at least one embodiment, one or more inputs 102 are information represented as one or more data packets. In at least one embodiment, the input 102 is received by one or more hardware components, such as those used in conjunction with the... Fig. 8A - 42 are described.
[0008] In at least one embodiment, the system 100 comprises a user device (UE) that is registered in one or more tracking areas of a 5G NR network. In at least one embodiment, one or more tracking areas comprise one or more cells. In at least one embodiment, a tracking area (for example, of a 5G NR network) is a designated area in which one or more base stations provide wireless services. In at least one embodiment, a mobile phone (for example, UE and / or receiving wireless device) registers with a base station within a tracking area to receive services. In at least one embodiment, the system 100 is said to use software 108 to select a neural network that performs one or more wireless signaling operations 116 in conjunction with one or more UEs.
[0009] In at least one embodiment, the processor 104 comprises one or more processor cores for executing one or more modules, for example, one or more modules for performing neural network selection 110 stored in memory 106. In at least one embodiment, these processor cores are hardware capable of executing or performing software modules 108 contained in memory 106. In at least one embodiment, these software modules 108 comprise one or more modules for selecting a neural network to perform one or more wireless signal operations 116, such as a module for selecting one or more neural networks to perform channel estimation 116A, beam shaping 116B, and / or signal mapping 116C.In at least one embodiment, the processor 104 transmits or executes an instruction, an API and / or another software instruction to access the memory 106 and to execute or perform software modules 108 for selecting a neural network 110.
[0010] In at least one embodiment, the software 108 for selecting a neural network can use one or more operations to perform the selection of the neural network 110, for example by: selecting a neural network with the highest ranking; selecting a neural network with the highest accuracy; selecting a neural network with the best accuracy-latency ratio; generating one or more ratings associated with a neural network to select a neural network; generating an average value from one or more ratings corresponding to the properties of neural network 102B to select a neural network; selecting a neural network using one or more user-inputted preferences; selecting a neural network using one or more votes and / or points;Selection of a neural network using one or more regions of one or more properties of a neural network 102B (for example, region of PRB mapping, region of SNR estimation and / or one or more regions of one or more measurements of one or more properties 102B); other methods described herein for selecting a neural network; and / or combinations thereof.
[0011] In at least one embodiment, the software 108 for selecting a neural network generates one or more outputs 112, such as a selected neural network 114 and / or a specification of a neural network, an expected signal property (for example, generated SNR estimate 306 and / or PRB assignment 302B, see Fig. 3), an output of the selected neural network 114 (for example, an output of one or more wireless signal operations 116), neural network signal properties 102A and / or neural network properties 102B.In at least one embodiment, the output 112 of the system 100 comprises the selected neural network 114 and / or a neural network specification, a neural network identifier, a memory location in memory 106, a wireless signal, signal properties 102A (for example, power, reference signal information, resource block information, a number of resource blocks, and / or a time allocation), neural network properties 102B (for example, neural network training information, types and / or extent of training according to a property of the input signal 118, layer information, expected latency, expected accuracy, and / or other neural network information described herein), telemetry data, information represented as data, and / or other outputs described herein. In at least one embodiment, one or more outputs 112 are transmitted to a processor 104 by a signal.In at least one embodiment, one or more outputs 112 are pieces of information represented as one or more data packets. In at least one embodiment, outputs 112 are generated by one or more hardware components, such as those used in conjunction with the . Fig. 8A - 42 are described.
[0012] In at least one embodiment, System 100 comprises a collection of one or more hardware and / or software computing resources with instructions that, when executed, perform one or more communication processes, such as those described herein. In at least one embodiment, System 100 is a software program running on computer hardware, an application running on computer hardware, and / or variations thereof. In at least one embodiment, one or more processes of System 100 are executed by any suitable processing system or processing unit (for example, a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a parallel processing unit (PPU), a central processing unit (CPU)), a data processing unit (DPU), as described below, and in any suitable manner, including sequential, parallel, and / or variations thereof.In at least one embodiment, the system 100 uses a machine learning training framework such as PYTORCH, TENSORFLOW, BOOST, CAFFE, MICROSOFT COGNITIVE TOOLKIT / CNTK, MXNET, CHAINER, KERAS, DEEPLEARNING4J, and / or another training framework to implement and execute the operations described herein for selecting one or more neural networks to perform one or more wireless signal operations 116. In at least one embodiment, training a neural network model includes, for example, the use of a server (e.g., NVIDIA DGX server) which further includes at least one GPU (e.g., AMD MI200, VEGAL10, VEGO20, and ARCTURUS), an optimizer (e.g., ADAM OPTIMIZER), or a discriminator architecture (e.g., the discriminator architecture of face-vid2vid for GAN-loss training).
[0013] In at least one embodiment, the system 100 comprises one or more modules such that the system selects a neural network to perform one or more wireless signal operations 116. In at least one embodiment, a module comprises any combination of any types of logic (for example, software, hardware, firmware) and / or circuits configured to perform a described function. In at least one embodiment, a module comprises one or more circuits that are part of a larger system (for example, an integrated circuit (IC), a system-on-chip (SoC), a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), etc.).In at least one embodiment, a controller comprises any combination of any types of logic (for example, software, hardware, firmware) and / or circuits configured to perform a described function. In at least one embodiment, software comprises software packages, code, programming languages, drivers, instructions, instruction sets, or a combination thereof. In at least one embodiment, the hardware comprises hard-wired circuits, programmable circuits, state machine circuits, fixed-function circuits, execution unit circuits, firmware with stored instructions executed by programmable circuits, or a combination thereof.
[0014] In at least one embodiment, the system 100 comprises a logic unit. In at least one embodiment, a logic unit comprises firmware logic, hardware logic, or a combination thereof, configured to provide any function as further described herein. In at least one embodiment, a logic unit comprises circuitry that is part of a larger system (for example, IC, SoC, CPU, GPU, DPU). In at least one embodiment, a logic unit comprises logic circuitry for implementing firmware and / or hardware to select one or more neural networks to perform one or more wireless signaling operations 116.
[0015] In at least one embodiment, the system 100 comprises an engine. In at least one embodiment, an engine comprises a module and / or a logic unit, as further described herein. In at least one embodiment, a component comprises a module and / or a logic unit, as further described herein. In at least one embodiment, an engine comprises software logic, firmware logic, hardware logic, or a combination thereof, configured to provide any desired function, as further described herein. In at least one embodiment, a component comprises software logic, firmware logic, hardware logic, or a combination thereof, configured to provide any desired function, as further described herein.In at least one embodiment, operations performed by hardware and / or firmware can alternatively be implemented via a software module, which may be implemented as a software package, code, and / or instruction set. In at least one embodiment, a logic unit can also use a section of the software to implement its function.
[0016] In at least one embodiment, the system 100 comprises one or more processors 104 for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein. In at least one embodiment, the system 100 is configured in the Fig. The systems shown in Figures 1-7 include and / or comprise one or more signal-to-noise ratio (SNR) values to use one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein. In at least one embodiment, the system 100 introduces one or more of the following: Fig. The processes shown in Figures 1-7, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 100 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0017] Fig. Figure 2 shows a system 200 for training and deploying neural networks according to at least one embodiment. In at least one embodiment, the system 200 comprises a neural network training system 210 and / or a neural network inference system 220. For example, the neural network training system 210 comprises a training dataset 212, a first neural network 214, a loss function 216, and an optimizer 218. For example, the neural network inference system 220 comprises an inference dataset 222, neural network selection software 228, and / or a second neural network 224. In at least one embodiment, the training is performed using supervised learning, which comprises labeled training datasets of radio signals from a 5G network with specific SNRs, channel estimates, and PRB values.
[0018] In at least one embodiment, the neural network training system 210 comprises one or more first neural networks 214 to receive training, for example, training to perform one or more wireless signal operations 116. For example, one or more neural network properties 232 comprise features of the first neural network 214 (for example, features of one or more layers of the first neural network) and / or features of the training of the first neural network 214. In at least one embodiment, the neural network inference system 220 receives neural network properties 232, for example, by receiving an inference dataset comprising neural network properties 232. The neural network inference system 220 comprises neural network selection software 228, for example, the software 108 (see Fig. 1) to generate an output 230 for selecting a neural network. In at least one embodiment, the output 230 comprises one or more outputs 112. In at least one embodiment, the neural network selection software 228 generates one or more outputs 230 to include a selected neural network (for example, the selected neural network 114 and / or the second neural network 224) and / or one or more pieces of information about a selected neural network. In at least one embodiment, the second neural network 224, which was selected using an output 230 of the neural network selection software 228, generates one or more output predictions 226 (for example, a channel estimate and / or one or more outputs for the wireless signal operation).
[0019] In at least one embodiment, the system 200 comprises a distributed system, which may refer to a network of independent computers that coordinate to achieve a common functionality (for example, neural network training, neural network inference). In at least one embodiment, the system 200 comprises nodes connected via communication protocols, data distribution methods, and synchronization mechanisms. In at least one embodiment, one or more nodes execute processes simultaneously on different machines, exchange messages, and replicate data. In at least one embodiment, the system 200 performs load balancing and fault detection operations to manage resources and ensure the reliability of the system. Alternatively, in at least one embodiment, the system 200 comprises a single computer or a server that manages and controls all operations.
[0020] In at least one embodiment, the system 200 comprises a neural network training system 210 and a neural network inference system 220. In at least one embodiment, the neural network training system 210 relates to one or more of the components associated with Fig. 1. Software and hardware components described herein for training one or more neural networks described herein. In at least one embodiment, the neural network training system 210 comprises one or more frameworks, such as TensorFlow, PyTorch, Keras, MXNet, Caffe, Theano, etc. In at least one embodiment, the neural network training system 210 uses one or more hardware accelerators (e.g., GPUs) described herein to accelerate one or more sections for training a neural network, such as the first neural network 214. In at least one embodiment, the first neural network 214 comprises one or more neural networks that, in conjunction with Fig. 3 are described.
[0021] In at least one embodiment, the neural network training system 210 performs normalization and transformation of input data, such as the training dataset 212. In at least one embodiment, the neural network training system 210 performs data normalization processes that scale feature values to a standard range, such as min-max scaling or Z-score normalization. In at least one embodiment, the neural network training system 210 generates additional training examples that are added to the training dataset 212 by transformations such as rotation, mirroring, or cropping. In at least one embodiment, the neural network training system 210 performs feature extraction operations, extracting relevant attributes from raw data, and feature selection, determining the most important features for the first neural network 214 to identify.In at least one embodiment, the neural network training system 210 removes noise, corrects missing values and / or performs data cleaning tasks using the training data set 212.
[0022] In at least one embodiment, the neural network training system 210 defines one or more layers and connections of the first neural network 214. In at least one embodiment, the neural network training system 210 determines a type of each layer, for example, convolutional layers, recurrent layers, or fully connected layers, and specifies parameters such as a number of neurons or filters. In at least one embodiment, the neural network training system 210 assigns specific activation functions to each layer, such as ReLU or sigmoid, to introduce nonlinearity. In at least one embodiment, the neural network training system 210 specifies one or more connection patterns by configuring how layers interact to form a sequence, jump connections, or branching paths.In at least one embodiment, the neural network training system 210 defines one or more input layers and / or output layers to ensure adequate data flow through the first neural network 214. In at least one embodiment, the neural network training system 210 initializes weights and biases for each connection and sets initial values that influence a training process. In some examples, the initialization of one or more weights and biases includes (1) zero initialization (for example, setting all weights to zero); (2) random initialization (for example, setting weights to small random values); (3) Glorot initialization (for example, adjusting a scale of weights according to a number of input and output neurons); and / or (4) He initialization (for example, setting weights with a variance that scales with a number of input neurons).
[0023] In at least one embodiment, the first neural network 214 comprises one or more neural networks that are connected to Fig. 1 are described. In some examples, the first neural network 214 comprises an untrained neural network, for example, a network architecture that is initialized but does not yet contain any training data. In other examples, the first neural network 214 comprises pre-trained neural networks, such as VGG, ResNet, GoogleNet, EfficientNet, YOLO, BERT, GPT, T5, RoBERTa, XLNet, DeepSpeech, Wav2Vec, Jasper, AlphaZero, StyleGAN, etc. In other examples, the first neural network 214 comprises a second neural network 224 that is already trained.
[0024] In at least one embodiment, the training dataset 212 comprises a collection of labeled and / or unlabeled data used to train the first neural network 214. In at least one embodiment, the training dataset 212 comprises one or more input samples representing one or more features and / or assigned to one or more neural network processes with corresponding target outputs that the first neural network 214 is to predict. In at least one embodiment, the training dataset 212 comprises one or more batches and / or mini-batches.In at least one embodiment, the training dataset 212 comprises various data formats, such as one or more wireless signals, one or more signal properties, text, and / or numerical data, by structuring the data into one or more formats compatible with an input layer of the first neural network 214. In at least one embodiment, the training dataset 212 includes metadata providing information about one or more data sources, one or more labeling schemes, and all applied preprocessing steps, as mentioned above. In some examples, there may be one or more neural networks (separate from the first neural network 214) that generate the training dataset 212. For example, one or more neural networks may include generative adversarial networks (GANs) and / or variational autoencoders (VAEs) that mimic one or more properties of a real dataset.
[0025] In at least one embodiment, the neural network training system 210 performs a forward pass using the training dataset 212. In at least one embodiment, the system 210 performs a forward pass, for example, a process in which input data from the training dataset 212 is passed through the first neural network 214 to generate one or more output predictions. In at least one embodiment, a forward pass comprises feeding input examples into an input layer of the first neural network 214, sequentially passing data through each hidden layer of the first neural network 214 by applying one or more defined activation functions, and generating outputs in an output layer of the first neural network 214.In at least one embodiment, the neural network training system 210 processes the computations of each layer by performing matrix multiplications with weights, adding biases, and applying activation functions to introduce nonlinearity.
[0026] In at least one embodiment, the neural network training system 210 uses a loss function 216 to evaluate the discrepancy between one or more output predictions and actual target values from the training dataset 212 generated during a forward pass. In at least one embodiment, the loss function 216 includes mechanisms for calculating a difference using specific mathematical formulations, such as the mean squared error for regression tasks or the cross-entropy loss for classification tasks. In at least one embodiment, the loss function 216 includes aggregations of individual errors across one or more training samples to generate a scalar value representing an overall performance of the first neural network 214.
[0027] In at least one embodiment, the system 200 uses an optimizer 218, for example, a computational component, that adjusts the weights and biases of the first neural network 214 to minimize the loss function 216. In at least one embodiment, the optimizer 218 includes algorithms such as stochastic gradient decay (SGD), Adam, and RMSprop, each implementing specific strategies for updating parameters based on computed gradients. In at least one embodiment, the optimizer 218 computes gradients of the loss function 216 with respect to each parameter by applying backward propagation and specifying a direction and the magnitude of the required adjustments. In at least one embodiment, the optimizer 218 manages learning rates that control a step size for each update and integrates techniques such as momentum to accelerate convergence by considering past gradient information.In at least one embodiment, the optimizer 218 performs adaptive adjustments of the learning rate, enabling different parameters to be updated at different rates based on their individual gradient profiles. In at least one embodiment, the optimizer 218 executes iterative update rules during each training period and systematically refines the parameters of the first neural network 214 to progressively reduce loss, thereby improving the performance of the first neural network 214 on the training dataset 212.
[0028] In at least one embodiment, the neural network training system 210 performs the training in a supervised, partially supervised, and / or unsupervised manner. In at least one embodiment, the neural network training system 210 performs federated learning, wherein several decentralized devices or servers jointly train the first neural network 214, while the training data (for example, sections of the training dataset 212) remain localized.
[0029] In at least one embodiment, the neural network training system 210 performs fine-tuning of the first neural network 214. In at least one embodiment, the fine-tuning includes performing additional training with a new, often more specific dataset to adapt its parameters to a particular task. In at least one embodiment, the fine-tuning includes loading one or more pre-trained weights and biases into an architecture of the first neural network 214, selecting specific layers of the first neural network 214 to be updated while others are frozen to retain previously learned features.In at least one embodiment, the fine-tuning includes reinitializing certain layers of the first neural network 214, if necessary, and applying regularization techniques to prevent overfitting during one or more subsequent training phases.
[0030] In at least one embodiment, the fine-tuning includes configuring a lower learning rate to make subtle adjustments to one or more parameters of the first neural network 214 to ensure that existing knowledge is retained while new information is taken into account.
[0031] In at least one embodiment, the neural network training system 210 performs an iterative process until the first neural network 214 achieves a desired accuracy. For example, the neural network training system 210 can evaluate the first neural network 214 using a test or validation set, where the accuracy can be a ratio of correctly predicted labels. In some examples, the accuracy of the first neural network 214 can depend on a final loss in a test or validation set. In at least one embodiment, after it has been determined that a desired accuracy has been achieved, the first neural network 214 becomes the second neural network 224. In some examples, the second neural network 224 can refer to one or more neural networks that are associated with Fig. 1 are described.
[0032] In at least one embodiment, the neural network inference system 220 can refer to a framework that executes trained neural network models, such as the second neural network 224, to generate output predictions 226 based on new input data, such as an inference dataset 222. In at least one embodiment, the neural network inference system 220 loads and / or initializes parameters (e.g., weights, biases) of the second neural network 224 into a runtime environment. In at least one embodiment, the neural network inference system 220 feeds an inference dataset 222 into an input layer of the second neural network 224, where values are generated and passed through one or more layers of the second neural network 224, generating output predictions 226. In some examples, an inference dataset 222 includes images, videos, text, audio, etc.In at least one embodiment, an inference dataset 222 comprises synthetic data generated by neural networks other than the second neural network 224 (e.g., GAN).
[0033] In at least one embodiment, the neural network inference system comprises 220 cloud servers or edge devices for providing the second neural network 224. In at least one embodiment, the neural network inference system comprises 220 cores, devices, inference chips, and GPUs for generating activations in order to further generate output predictions 226.In at least one embodiment, the output predictions 226 comprise one or more wireless signal operating outputs, one or more wireless signal predictions, one or more classification labels, one or more probability distributions, one or more continuous numeric values, one or more sequences, one or more images, one or more translations, one or more embeddings, one or more actions, one or more structured data outputs, audio, one or more heatmaps, one or more attention maps, generative content and / or one or more outputs described herein.
[0034] In at least one embodiment, the system 200 comprises one or more processors for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 200 is configured in the Fig. The systems shown in Figures 1-7 include and / or comprise one or more signal-to-noise ratio (SNR) values to use one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 200 introduces one or more of the Fig. The processes shown in Figures 1-7, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 200 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0035] Fig. Figure 3 is a block diagram showing an exemplary system 300 for selecting a neural network to generate a channel estimate according to at least one embodiment. In at least one embodiment, the system 300 comprises one or more neural networks and / or software for performing an SNR model selection 304, a PRB model selection 308, and / or a model selection using one or more signal properties.
[0036] In at least one embodiment, the System 300 estimates a wireless channel for different signal-to-noise ratios (SNRs) and time and frequency resource allocations, at least partially based on the selection of a neural network (for example, a specialized neural network) using properties of a signal. In at least one embodiment, a physical resource block (PRB) is a minimum time-frequency allocation for a user. In at least one embodiment, a resource element (RE) is the smallest time and frequency unit over which a channel is estimated. In at least one embodiment, a PRB consists of many REs.
[0037] In at least one embodiment, channel estimation in wireless communication comprises receiving reference signals and attempting to estimate channel-induced changes in a transmitted signal. In at least one embodiment, a receiver of a wireless signal (for example, a user device and / or base station) can, by knowing these changes, determine an original signal that was transmitted by a received signal. In at least one embodiment, this is a fundamental process that represents a bottleneck in many wireless systems, such as 5G.
[0038] In at least one embodiment, a single neural network is typically used to learn the channel estimate for many conditions, such as a single neural network over a variable number of resource blocks (frequency / time assignments) and a large dynamic range of received signal SNR. However, a single neural network can lead to suboptimal performance across all variations. For example, a specialized model for a narrow set of conditions and signal characteristics may perform better than a more generalist model that must adapt to many conditions. In at least one embodiment, the system selects 300 models and generates overlapping model outputs to obtain a coherent estimate of a channel.
[0039] In at least one embodiment, the system 300 performs an SNR model selection 304, which is based at least partially on generating an SNR estimate 306 (for example, using the least squares (LS) method 306A and / or the minimum mean squares error (MMSE) method 306B) and selects one or more neural networks where a generated estimate (for example, 14 dB) falls within a range of their neural network properties (for example, trained to 10–20 dB) and / or is a neural network that has been selected to best perform the channel estimation. In at least one embodiment, software (for example, software 108 and / or 228) performs an SNR model selection 304, which is based at least partially on receiving one or more signal properties 302A and generating an SNR estimate 306 (for example, using LS 306A and / or MMSE 306B).In at least one embodiment, a processor executes LS 306A in one or more 5G-NR operations to generate an SNR estimate of a signal. In at least one embodiment, LS 306A comprises the use of reference signals to estimate a channel and subsequently to calculate an SNR by minimizing a sum of the squared differences between observed and predicted values. In at least one embodiment, a processor executes software that includes one or more MMSE 306B operations, for example, a 5G-NR operation, to generate an SNR estimate of a signal. In at least one embodiment, MMSE 306B comprises one or more operations for minimizing a mean squared error between estimated and actual signal values.
[0040] In at least one embodiment, the system 300 performs a PRB model selection 308. In at least one embodiment, the PRB model selection comprises training one or more models with different PRB sizes and selecting a model, at least partially based on a set of assigned PRBs. For example, the system 300 comprises software for performing the PRB model selection 308 using one or more inputs 302A, such as one or more signal properties 302A and / or a PRB assignment 302B. In at least one embodiment, a neural network is selected at least partially based on one or more SNR estimates and one or more PRB assignments using one or more neural network properties.In at least one embodiment, the System 300 selects models at the time of inference that are at least partially based on a PRB 302B assigned to that user. For example, the System 300 trains models with respect to training data associated with 1 PRB, 4 PRBs, and 16 PRBs, where 31 PRBs are assigned to a user, and the System 300 would use 1 x 16-PRB model, 3 x 4-PRB models, and 3 x 1-PRB models for inference.
[0041] In at least one embodiment, the system 300 performs model overlap, for example, to combine the responses of several models across the frequency by overlapping model predictions to ensure a smoother transition at an edge of the model. In at least one embodiment, the system 300 improves performance by overlapping an error in channel estimation across a frequency band. In at least one embodiment, peaks (for example, corresponding to reduced performance) in an MSE are displayed on a graph (for example, the one in Fig. Spikes (as shown in graph 4) at edges where models transition are undesirable. In at least one embodiment, to reduce spikes, an additional model was placed in the middle of a transition and an estimate for this model was included in an overall estimate.
[0042] In at least one embodiment, the System 300 performs model selection between signal-to-inference-plus-noise ratios (SINRs), for example, by using multiple networks for different SINRs. In at least one embodiment, the System 300 includes software for performing SINR selection of a neural network. In at least one embodiment, the System 300 includes model selection across subcarriers for interpolation purposes—since reference signal piles have half a dimension of the required channel estimates, the System 300 can process them using two models. In at least one embodiment, this differs from a solution that can use a fully connected network to perform the interpolation, which increases complexity and can impair performance.
[0043] In at least one embodiment, the System 300 performs model overlap at model edge transitions to further improve the performance of the model stitching, for example, by overlapping model estimates to form an average and maintain good performance across a frequency band. For example, the System 300 is included in and / or used in conjunction with a system to perform Aerial Omniverse Digital Twin and / or PyAerial.
[0044] In at least one embodiment, a 5G standard alternates between even and odd resource elements (REs). In at least one embodiment, even REs can be assigned to one user while odd REs are assigned to another; for example, every fourth RE could be assigned to one user. In at least one embodiment, this approach allows for the interleaving of different pilot signals over time and frequency. In at least one embodiment, the system can accurately decode 300 data points, where the channel conditions for each RE can be known, by specializing neural networks according to even or odd REs. For example, depending on the sparsity of the pilot signals in time and frequency, multiple models can be trained for different patterns, such as every fourth RE. In at least one embodiment, this strategy forms the basis for our approach to selecting the interpolation model.In at least one embodiment, the system 300 performs a selection of the interpolation model. For example, the system 300 performs a selection of the interpolation model that is at least partially based on the use of separate models for interpolating a channel on REs without pilot signals, since pilot reference signals used to measure a channel may not be present on every resource element.
[0045] For example, the System 300 can also select a neural network based on location-specific channels, such as a neural network that has been selected at least partially based on a stochastic channel, different selectivities, different delay variations, different flat channels and / or frequencies.
[0046] In at least one embodiment, the System 300 comprises one or more processors for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the System 300 is configured in the Fig. The systems shown in 1-7 include and / or comprise them to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 300 introduces one or more of the Fig. The processes shown in Figures 1-7, for example, to use one or more signal-to-noise ratio (SNR) values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 300 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0047] Fig. Figure 4 shows an example of combining one or more neural networks to perform wireless signal operations according to at least one embodiment. For example, a System 400 uses software to combine (for example, link) one or more neural networks to perform wireless signal operations. In at least one embodiment, the System 400 comprises one or more neural networks to generate an SNR estimate 306, for example, to determine an average MSE without merging (for example, -15.5 dB) and / or an average MSE with merging of one or more neural networks (for example, -16.4 dB).
[0048] In at least one embodiment, the system 400 performs model overlap, for example, to combine the responses of multiple models across the frequency range by overlapping model predictions to ensure a smoother transition at a model edge. In at least one embodiment, the system 400 improves performance by overlapping an error in channel estimation across a frequency band. In at least one embodiment, peaks (for example, corresponding to reduced performance) in an MSE are displayed on a graph (for example, the one shown in [reference]). Fig. (Diagram 4 shown) at edges where models transition are undesirable. In at least one embodiment, to reduce peaks, an additional model was placed in the middle of a transition, and an estimate for this model was included in an overall estimate. In at least one embodiment, the system 400 performs model overlap at model edge transitions to further improve the model joining performance, for example, by overlapping model estimates to average and maintain good performance across a frequency band.
[0049] In at least one embodiment, the System 400 comprises one or more processors for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein. In at least one embodiment, the System 400 is configured in the Fig. The systems shown in Figures 1-7 include and / or comprise one or more signal-to-noise ratio (SNR) values to use one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 400 introduces one or more of the following: Fig. The processes shown in Figures 1-7, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 400 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0050] Fig. Figure 5 is a process flowchart 500 that illustrates the selection of a neural network, at least partially, based on one or more properties of a signal according to at least one embodiment. In at least one embodiment, the process 500 comprises one or more steps for receiving 502 one or more inputs comprising properties of a signal, generating 510 an expected signal-to-noise ratio (SNR) of a signal, specifying 512 a user-assigned PRB quantity, specifying 514 any other property information to be used for selecting a neural network, selecting 516 a neural network using selected properties and / or expected properties of a signal, combinations thereof, and / or performing one or more operations described herein.In at least one embodiment, process 500 begins when it is called by one or more processors, for example using an API.
[0051] In at least one embodiment, a system executing process 500 receives one or more inputs (for example, input 102 and / or 302 and / or inference data set 222), for example, an input comprising one or more properties of a signal (for example, signal properties 102A). In at least one embodiment, a system executing process 500 uses one or more received inputs to determine which properties of a signal should be used in selecting a neural network. In at least one embodiment, after receiving one or more inputs, a system executing process 502 proceeds to decision blocks 504, 506, and 508.In at least one embodiment, a system performing the process 500 can, after receiving 502 one or more inputs, proceed with the decision blocks 504, 506 and 508, which are arranged in series, in parallel or in combinations thereof.
[0052] In at least one embodiment, the decision in decision block 504 is “YES” if the software is to select a neural network using SNR; otherwise, the decision is “NO”. In at least one embodiment, a system executing process 500 proceeds to generate an expected SNR of a signal (for example, using LS 306A and / or MMSE 306B, see Figure 510). Fig. 3) if a decision in decision block 504 is "YES". In at least one embodiment, after generating 510 an expected SNR, a processor proceeds to output an expected SNR to one or more operations in order to 516 select a neural network using one or more selected properties and / or expected properties of a signal. In at least one embodiment, if a decision in decision block 504 is "YES", a processor selects a neural network without using SNR to select a neural network, for example, by instead using one or more specified properties to select a neural network.
[0053] In at least one embodiment, the decision in decision block 506 is "YES" if software is to select a neural network using an assigned set of PRBs; otherwise, the decision is "NO". In at least one embodiment, if a decision in decision block 506 is "YES", a system executing process 500 proceeds by specifying 512 a set of assigned PRBs and / or that a neural network is to be selected using a set of PRBs. In at least one embodiment, after specifying 512 an assigned set of PRBs, a processor performs one or more operations to select a neural network using one or more selected properties and / or expected properties of a signal 516.In at least one embodiment, if a decision in decision block 506 is “YES”, a processor proceeds with the selection 516 of a neural network without using SNR to select a neural network that has not been specified, for example, using one or more specified signal properties.
[0054] In at least one embodiment, the decision in decision block 508 is "YES" if the software is to select a neural network using one or more other properties (for example, signal properties and / or properties of the neural network); otherwise, the decision is "NO". In at least one embodiment, a system executing process 500 proceeds to specify 514 other property information to be used for selecting a neural network if a decision in decision block 506 is "YES". In at least one embodiment, after specifying 514 other property information, a processor proceeds to perform one or more operations to select 516 a neural network using one or more selected properties and / or expected properties of a signal.In at least one embodiment, if a decision in decision block 508 is “YES”, a processor proceeds with the selection 516 of a neural network without using SNR to select a neural network that has not been specified, for example, using one or more specified signal properties.
[0055] In at least one embodiment, a system performing the process 500 selects 516 one or more neural networks using one or more selected properties and / or expected properties of a signal. For example, the selection 516 of one or more neural networks is based at least partially on one or more properties of a neural network, such as one or more conditions for training a neural network (for example, types of training datasets 212 and / or properties of neural networks 232, see Fig. 2) For example, a system uses software 108 and / or 228 to select 516 a neural network using one or more properties and / or expected properties of a signal. In at least one embodiment, a processor executing process 500 can select 516 one or more neural networks, specify this selection, execute one or more selected neural networks, perform one or more operations described herein, and / or terminate.
[0056] In at least one embodiment, other deployment properties can include, for example, the location of a base station, the number and configuration of antennas, the locations and characteristics of antennas, site-specific attributes (such as terrain, urban density, or sources of interference), and types of channels (such as different Doppler shifts). For example, neural networks can be pre-trained for specific locations using detailed site features, and input signals can be analyzed to use these site-specific properties to identify and select the most suitable neural network.In another embodiment, software can simulate a 5G network environment, generate different scenarios for training neural networks, and then deploy these networks in real base stations to optimize performance under different conditions.
[0057] A digital twin in 5G is a virtual representation of the physical network, its components, and their interactions. It operates in real time and reflects the state, configurations, and traffic patterns of the live network. By simulating various scenarios, such as load balancing, interference mitigation, or equipment failures, the digital twin enables network operators to test and optimize performance without impacting the operational network. It also facilitates predictive maintenance by identifying potential problems before they occur and supports training and development by providing a controlled environment for testing new technologies and algorithms.
[0058] In at least one embodiment, a digital twin of a processor is used for the planning, deployment, and optimization of 5G networks. A digital twin enables network operators, for example, to assess coverage, capacity, and performance for specific locations, taking into account factors such as antenna placement, user density, and urban environments. In at least one embodiment, a digital twin includes software that enables service backup and ensures consistent quality by virtually identifying and resolving potential problems before they occur in a live network. In at least one embodiment, the software disclosed herein can be used with digital twins to analyze historical data, predict network behavior, and enable dynamic adjustments to network configurations.
[0059] In at least one embodiment, the disclosed neural network and selection software can be integrated into a digital twin to simulate a deployment that utilizes properties of that location. In at least one embodiment, the integration enables the digital twin not only to simulate the network but also to dynamically select and deploy the most appropriate neural network based on simulated input. For example, during a simulation, a digital twin can use site-specific features or environmental conditions for identification and utilize the selection software to choose neural networks optimized for those scenarios.
[0060] In at least one embodiment, the software disclosed herein can use additional properties to select one or more neural networks, including antenna pairings and imaging requirements, environmental conditions and user-specific factors, as well as based on known channels at one antenna and desired estimates at another. In at least one embodiment, properties of an input signal can include environmental factors such as time of day or week, which can determine the selection of models optimized for high- or low-traffic scenarios. In at least one embodiment, properties for selecting a neural network can include interference caused by buildings or geographical features in an environment for simulation.In at least one embodiment, the user's location and the physical environment, such as urban buildings or open fields, can influence communication quality. These properties can also be used by the disclosed software to select neural networks optimized for specific areas within a cell. In at least one embodiment, the selection is dynamic to ensure robust and efficient network performance under varying operating conditions.
[0061] In at least one embodiment, the process 500 (or any other process described herein, or variations and / or combinations thereof) is carried out wholly or partly under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (for example, computer-executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions that can be executed by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-volatile computer-readable medium.In at least one embodiment, at least some computer-readable instructions that can be used to carry out Process 500 are not stored exclusively using transitory signals (for example, a propagating transient electrical or electromagnetic transmission). In at least one embodiment, a non-volatile, computer-readable medium does not necessarily comprise non-volatile data storage (for example, buffers, caches, and queues) within transient signal transmitters. In at least one embodiment, Process 500 is carried out at least partially on a computer system as described elsewhere in this disclosure. In at least one embodiment, logic (for example, hardware, software, or a combination of hardware and software) carries out Process 500.
[0062] In at least one embodiment, one or more processors use Process 500, for example, to use one or more signal-to-noise ratio (SNR) values, to select one or more neural networks, to generate one or more channel estimates, and / or to otherwise perform operations described herein. In at least one embodiment, for example, a machine-readable medium (for example, non-volatile) comprises a set of instructions which, when executed by one or more processors, cause one or more processors to execute Process 500, for example, to use one or more signal-to-noise ratio (SNR) values, to select one or more neural networks, to generate one or more channel estimates, and / or to otherwise perform the operations described herein. In at least one embodiment, Process 500 is contained in the Fig. The processes shown in Figures 1-7 include and / or comprise using one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, one or more of the Fig. The systems shown in Figures 1-7 utilize Process 500, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, one or more of the systems shown in the Fig. 8-42 Hardware shown in the process 500, for example to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0063] Fig. Figure 6 shows a processor and instructions according to at least one embodiment. The system 600 can comprise a memory 602 and one or more processors 608. The memory 602 can, for example, comprise main memory, a cache, or other memory, which is further described herein. The memory 602 can be separate from the processor(s) 608, or the memory 602 can be contained within the processor(s) 608 (for example, in the memory 612). In at least one embodiment, a software program 604 and / or software libraries (or instructions) 606 can be stored in a memory, cache, or other memory and made available to the processor(s) 608 to cause one or more circuits of the processor(s) 608 to perform the operations described herein.In at least one embodiment, the software program 604 and / or the software libraries (or instructions) 606 can be integrated into one or more circuits of the processor(s) 608. The software program 604, which can be used to perform one or more of the operations described herein, can be stored in memory 602.
[0064] In at least one embodiment, the software program 604 can comprise one or more software modules. In at least one embodiment, one or more modules comprise a module for selecting a neural network and / or a module for estimating the SNR of a neural network.
[0065] In at least one embodiment, as used in each implementation described herein, unless otherwise indicated by the context or expressly stated otherwise, a module refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software is embodied as a software package, code, and / or instruction set or instructions, and "hardware," as used in each implementation described herein, includes, for example, hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, and / or firmware that stores instructions executed by programmable circuitry, either individually or in any combination.In at least one embodiment, modules are implemented collectively or individually as circuits that are part of a larger system, for example, an integrated circuit (IC), a system-on-chip (SoC), etc. In at least one embodiment, a module executes one or more processes in conjunction with a suitable processing unit and / or a combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variations thereof, which are further described herein.
[0066] In at least one embodiment, the software program 604 may comprise a collection of software code, commands, instructions, or other text sequences for instructing a computer device to perform one or more arithmetic operations and / or to invoke one or more other instruction sets, such as APIs or API functions, or ISA-level (Instruction Set Architecture) commands, to be executed and / or otherwise performed. Commands (e.g., hardware commands) or microcode may include ISA-level commands, which may contain native ISA commands or non-native ISA commands. The software program 604 and / or the software libraries 606 (e.g., one or more modules) containing commands may be distributed among multiple processors that communicate via a bus, a network, by writing to common memory, and / or a suitable communication process as described herein.
[0067] In at least one embodiment, the system 600 may comprise one or more software libraries 606, which may, for example, provide one or more APIs and / or ISA commands. In at least one embodiment, one or more APIs and / or ISA commands may be used to utilize one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform operations described herein. In at least one embodiment, one or more software libraries 606 may be included in drivers and / or runtime environments.In at least one embodiment, software libraries 606 (for example, including one or more ISA instructions) can comprise sets of software instructions which, when executed and / or otherwise performed, cause the processor(s) 608 to perform one or more arithmetic operations, such as one or more of the operations described herein. In at least one embodiment, one or more APIs and / or ISA instructions can be distributed or otherwise provided as part of one or more software libraries 606, runtimes, drivers, and / or any other grouping of software and / or executable code further described herein. In at least one embodiment, one or more APIs and / or ISA instructions can perform one or more arithmetic operations in response to a call by a software program 604.
[0068] The processor(s) 608 may comprise any number of processors and any suitable processing unit and / or combination of processing units, such as, but not limited to, central processing units (“CPUs”), graphics processing units (“GPUs”), or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, parallel processors, GPGPUs, DPUs, and / or variations thereof, including those further described herein), and may include all processors described herein, such as, but not limited to, the processors in the Fig. 9-21. In at least one embodiment, the processor(s) 608 can retrieve or fetch instructions (for example, one or more APIs and / or ISA instructions) from memory 602 by using, for example, the instruction fetch 616 (for example, for an instruction fetch phase). Instructions can include instructions for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to otherwise perform operations described herein. In at least one embodiment, the processor(s) 608 can include memory 612 and an instruction queue 610 to store and queue instructions fetched from memory 602.In at least one embodiment, retrieved instructions can be decoded by a decoding unit 618 to determine which operation is to be performed by the processor(s) 608 (for example, in an instruction decoding stage). In at least one embodiment, the processor(s) 608 can retrieve additional operands (data) that can be used for instructions, and operands can be stored, for example, in registers or in memory 612. In at least one embodiment, microoperations 620 can perform operations on data stored in one or more registers or in memory 612. For example, each step of instructions retrieved by the processor(s) 608 can be decomposed during execution so that the processor(s) 608 can execute instructions step by step through a series of microoperations 620.In at least one embodiment, the program counter (PC) 614 can contain an address for a next instruction and be updated to point to a next instruction to be executed by the processor(s) 608.
[0069] In at least one embodiment, the processor(s) 608 can execute instructions (for example, during an execution phase). For example, the processor(s) 608 can perform an operation defined by one or more instructions, such as an arithmetic operation, a logical operation, or a data transfer. In at least one embodiment, the arithmetic unit(s) 622 can execute instructions to perform one or more of the operations described herein. In at least one embodiment, the arithmetic unit can comprise ALU(s) 624 (Arithmetic Logic Units) that can be used to perform arithmetic and logical operations. In at least one embodiment, the arithmetic unit can comprise FPU(s) (Floating Point Units) 626 that can be used to perform floating-point calculations.In at least one embodiment, other circuits 628 can be used to perform other operations, such as vector and / or scalar operations. In at least one embodiment, the accelerator(s) 630 can comprise one or more matrix multiplication accelerators, one or more parallel processing units (PPUs), such as GPUs, or any other accelerator or processor further described herein. In at least one embodiment, the software program 604 can use one or more APIs and / or ISA instructions to perform various arithmetic operations with the accelerator(s) 630, such as matrix multiplication, arithmetic operations, or any other arithmetic operation further described herein.In at least one embodiment, one or more computational operations using the accelerator(s) 630 may comprise at least one or more groups of computational operations which are to be accelerated by at least partial execution by the accelerator(s) 630, including the use of one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein.
[0070] In at least one embodiment, the System 600 can be used to execute one or more commands comprising functions or operations such as those associated with the Fig. 1-7 are described. In at least one embodiment, the system 600, comprising one or more processors, causes one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 600 is in the Fig. The systems shown in Figures 1-7 include and / or comprise them to cause one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 600 comprises one or more of the circuits shown in Figures 1-7 to cause one or more neural networks to use one or more channel estimates to use one or more neural networks to use one or more neural networks to use one or more neural networks to generate one or more channel estimates and / or to otherwise perform the operations described herein. In at least one embodiment, the system 600 comprises one or more of the circuits shown in Figures 1-7 to cause one or more neural networks to use one or more neural networks to use one or more channel estimates to use one or more neural networks to use one or more neural networks to use one or more neural networks to generate one or more channel estimates and / or to perform the operations described herein. Fig. 8-42 Hardware shown, for example, to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein.
[0071] In at least one embodiment, the System 600 comprises one or more processors for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein. In at least one embodiment, the System 600 is configured in the Fig. The systems shown in Figures 1-7 include and / or comprise them to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or to perform other operations described herein. In at least one embodiment, the system 600 introduces one or more of the following: Fig. The processes shown in Figures 1-7, for example, to use one or more signal-to-noise ratio (SNR) values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 600 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0072] Fig. Figure 7 is a block diagram showing a driver and / or runtime comprising one or more libraries to provide one or more application programming interfaces (APIs), according to at least one embodiment. In at least one embodiment, a software program 702 is a software module, such as those described in Fig. 6 are described. In at least one embodiment, a software program 702 comprises one or more software modules. In at least one embodiment, a software module is, as described in Fig. 6 further described, but not exclusively. In at least one embodiment, one or more APIs 710 are sets of software instructions which, when executed, cause one or more processors to perform one or more arithmetic operations. In at least one embodiment, one or more APIs 710 are distributed or otherwise provided as part of one or more libraries 706, runtime environments 704, drivers 704, and / or another grouping of software and / or executable code further described herein. In at least one embodiment, one or more APIs 710 perform one or more arithmetic operations in response to a call by software programs 702.In at least one embodiment, a software program 702 is a collection of software code, commands, instructions, or other text sequences for instructing a computer device to perform one or more arithmetic operations and / or to call one or more other instruction sets, such as APIs 710 or API functions 712, to be executed. In at least one embodiment, the functionality provided by one or more APIs 710 includes software functions 712, such as those that can be used to accelerate one or more sections of software programs 702 using one or more parallel processing units (PPUs), such as graphics processing units (GPUs). In at least one embodiment, a software program is a neural network. In at least one embodiment, a software program comprises one or more algorithms.
[0073] In at least one embodiment, APIs 710 are hardware interfaces to one or more circuits for performing one or more arithmetic operations. In at least one embodiment, one or more of the software APIs 710 described herein are implemented as one or more circuits to access one or more of the functions associated with the Fig. to execute the techniques described in 1-6. In at least one embodiment, one or more software programs comprise 702 instructions which, when executed, cause one or more hardware devices and / or circuits to perform one or more of the techniques described in connection with the Fig. to perform the techniques described in 1-6.
[0074] In at least one embodiment, software programs 702, such as user-implemented software programs, use one or more application programming interfaces (APIs) 710 to perform various computational operations, such as memory allocation, matrix multiplication, arithmetic operations, or any computational operations performed by parallel processing units (PPUs), such as graphics processing units (GPUs), as further described herein. In at least one embodiment, one or more APIs 710 provide a set of callable functions 712, referred to herein as APIs, API functions, and / or functions, which individually perform one or more computational operations, such as computational operations related to parallel computing.In at least one embodiment, one or more APIs 710 provide functions 712 to specify 716 and / or cause a neural network (for example, a selected neural network) to perform wireless signaling operations and / or other operations described herein. In at least one embodiment, one or more APIs 710 provide functions 712 to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks, specify 716, and / or cause them to generate one or more channel estimates.
[0075] In at least one embodiment, one or more software programs 702 interact with or otherwise communicate with one or more APIs 710 to perform one or more computational operations using one or more PPUs, such as GPUs. In at least one embodiment, one or more computational operations using one or more PPUs comprise at least one or more groups of computational operations that are to be accelerated, at least partially, by the execution of the one or more PPUs. In at least one embodiment, one or more software programs 702 interact with one or more APIs 710 to facilitate parallel computing using a remote or local interface.
[0076] In at least one embodiment, an interface is a software instruction that, when executed, provides access to one or more functions 712 provided by one or more APIs 710. In at least one embodiment, a software program 702 uses a local interface when a software developer compiles one or more software programs 702 in conjunction with one or more libraries 706 that include or otherwise provide access to one or more APIs 710. In at least one embodiment, one or more software programs 702 are statically compiled in conjunction with precompiled libraries 706 or uncompiled source code that includes instructions for executing one or more APIs 710.In at least one embodiment, one or more software programs 702 are dynamically compiled, and the one or more software programs use a linker to establish a connection to one or more precompiled libraries 706 comprising one or more APIs 710.
[0077] In at least one embodiment, a software program 702 uses a remote interface when a software developer executes a software program that uses or otherwise communicates with a library 706 comprising one or more APIs 710 over a network or other remote communication medium. In at least one embodiment, one or more libraries 706 comprising one or more APIs 710 are executed by a remote computing service, for example, a computing resource services provider. In another embodiment, one or more libraries 706 comprising one or more APIs 710 are executed by any other computer host that provides the one or more APIs 710 to one or more software programs 702.
[0078] In at least one embodiment, one or more software programs 702 use one or more APIs 710 to allocate and otherwise manage memory to be used by the software programs 702. In at least one embodiment, one or more software programs 702 use one or more APIs 710 to handle the allocation and other management of memory to be used by one or more sections of the software programs 702 for acceleration by one or more PPUs, such as GPUs or other accelerators or processors further described herein. In at least one embodiment, these software programs 702 instruct a neural network (for example, by specifying, initiating, and / or identifying it) to generate one or more channel estimates and / or to perform operations described herein.
[0079] In at least one embodiment, an API 710 is an API for facilitating parallel data processing. In at least one embodiment, an API 710 is any other API further described herein. In at least one embodiment, an API 710 is provided by a driver and / or a runtime environment 704. In at least one embodiment, an API 710 is provided by a CUDA user-mode driver. In at least one embodiment, an API 710 is provided by a CUDA runtime. In at least one embodiment, a driver 704 is data values and software instructions that, when executed, perform or otherwise facilitate the operation of one or more functions 712 of an API 710 during the loading and execution of one or more sections of a software program 702.In at least one embodiment, a runtime environment 704 comprises data values and software instructions that, when executed, perform or otherwise facilitate the operation of one or more functions 712 of an API 710 during the execution of a software program 702. In at least one embodiment, one or more software programs 702 use one or more APIs 710, implemented or otherwise provided by a driver and / or a runtime 704, to specify 716 and / or cause a neural network to generate one or more channel estimates, wherein the neural network is selected by one or more software programs 702 during execution by one or more PPUs, such as GPUs.
[0080] In at least one embodiment, one or more software programs 702 use one or more APIs 710 provided by a driver and / or a runtime environment 704 to specify 716 and / or cause a neural network to generate one or more channel estimates. In at least one embodiment, one or more APIs 710 provide software programs to cause a neural network 716 to specify 716 and / or generate one or more channel estimates via a driver and / or a runtime environment 704, as described above. In at least one embodiment, one or more software programs 702 use one or more APIs 710 provided by a driver and / or a runtime environment 704 to allocate or otherwise reserve one or more memory blocks 714 of one or more PPUs, such as GPUs.In at least one embodiment, one or more software programs 702 use one or more APIs 710 provided by a driver and / or a runtime environment 704 to allocate or otherwise reserve memory blocks. In at least one embodiment, one or more APIs 710 are used to specify a neural network 716 and / or to cause one or more channel estimates to be generated, as in conjunction with the . Fig. 1-7 are indicated.
[0081] In at least one embodiment, to improve the usability of software programs 702 and / or to optimize one or more sections of the software programs 702 that are to be accelerated by one or more PPUs, such as GPUs, in one embodiment provide one or more APIs 710 and one or more API functions 712 to specify 716 and / or to cause a neural network to generate one or more channel estimates as specified above, and furthermore in conjunction with the Fig. 1-7 specified. In at least one embodiment, an exemplary system 700 comprises a processor 608 (see Fig. 6), comprising one or more circuits for executing one or more software programs to perform one or more functions 712. In at least one embodiment, an exemplary block diagram depicts the system 700, which comprises one or more processors for executing one or more software programs to specify 716 and / or to induce a neural network to generate one or more channel estimates. In at least one embodiment, the processor uses an API 710 to specify 716 and / or to induce a neural network to generate one or more channel estimates and / or otherwise to perform the operations described herein.
[0082] In at least one embodiment, the System 700 comprises one or more processors for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the System 700 is configured in the Fig. The systems shown in Figures 1-7 include and / or comprise one or more signal-to-noise ratio (SNR) values to use one or more neural networks to generate one or more channel estimates and / or otherwise perform the operations described herein. In at least one embodiment, the system 700 introduces one or more of the following: Fig. The processes shown in Figures 1-7, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates, and / or to perform other operations described herein. In at least one embodiment, the system 700 comprises one or more of the Fig. 8-42 Hardware shown, for example, to use one or more SNR values, to select one or more neural networks, to generate one or more channel estimates and / or to perform other operations described herein.
[0083] The preceding and following descriptions detail various techniques. Specific configurations and details are presented for illustrative purposes, aiming to provide a comprehensive understanding of the possible implementations of these techniques. However, it is also clear that the techniques described below can be practiced in different configurations without specific details. Furthermore, known features can be omitted or simplified to avoid the described techniques altogether. DATA CENTER
[0084] Fig. Figure 8 shows an example of a data center 800 according to at least one embodiment. The data center 800 can have one or more rooms in which racks 802 and accessory equipment for housing one or more racks 802 and one or more baseboards 804 are used. The rack 802 can comprise one or more baseboards 804. The rack 802 can comprise an enclosure that accommodates and supports individual baseboards 804. The operating aspects of the rack 802 can be controlled, among other things, at the rack level, corresponding to a group of baseboards 804, or at the baseboard level, corresponding to individual baseboards 804. The rack 802 or the baseboards 804 can have specifically selected maximum operating parameters, such as, but not limited to, power consumption, operating frequencies, and others.The 800 data center can be supported by various cooling systems, such as, but not limited to, cooling towers, cooling circuits, pumps, and other support systems. Cooling systems can include sensors and controllers to monitor and manage the cooling characteristics of the 802 racks. The 804 baseboards within the 802 racks can obtain their operating power from one or more power distribution units (PDUs; not shown). PDUs can be located within the 802 racks, for example, between 802 racks containing 804 baseboards, or within 802 racks that also contain 804 baseboards.
[0085] Racks 802 and baseboards 804 can include subsystems, modules, add-in cards, and other semiconductor components. Baseboards 804 can include one or more processing units 806, which can include one or more processors 808, one or more memory units 810, and an interface controller 812. The processing units 806 can include any number of processors, such as, but not limited to, central processing units (“CPUs”), graphics processing units (“GPUs”), or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), including all processors described herein, such as, but not limited to, the processors in the Fig. 9-21. The 806 Computing Units can include one or more 810 Storage Devices (for example, dynamic read-only memory, solid-state storage, or hard disk drives), as well as network input / output devices (“NW I / O devices”), network switches, virtual machines (“VMs”), power supply modules, cooling modules, etc. One or more 806 Computing Units can constitute a server that includes one or more of the above-mentioned computing resources.
[0086] Computing Units 806 can comprise separate groupings of compute units located in one or more racks (not shown), or many racks located in data centers at different geographic locations (also not shown). Separate groupings of compute units can comprise grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. Multiple compute units (for example, comprising CPUs and / or other processors) can be grouped within one or more racks to provide compute resources to support one or more workloads. A Resource Orchestrator 814 can configure or otherwise control one or more Computing Units 806 or groups of compute units.The Resource Orchestrator 814 can include a Software Design Infrastructure Management Unit (“SDI”) for the Data Center 800. The Resource Orchestrator 814 can include hardware, software, or a combination of both.
[0087] The 800 data center can comprise any combination of a framework layer (820), a software layer (830), and an application layer (840). As in Fig. As shown in Figure 8, the framework layer 820 comprises a job scheduler 822, a configuration manager 824, a resource manager 826, and a distributed file system 828. The framework layer 820 can include a framework to support the software 832 of the software layer 830 and / or one or more applications 842 of the application layer 840. The software 832 or the applications 842 can each include web-based service software or applications, such as, but not limited to, those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 820 can be a type of free and open-source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter "Spark"), which can utilize the distributed file system 828 for processing large amounts of data (for example, "Big Data").The Job Scheduler 822 can include a Spark driver to facilitate the scheduling of workloads supported by various layers of the Data Center 800. The Configuration Manager 824 can configure different layers, such as, but not limited to, the Software Layer 830 and the Framework Layer 820, including Spark and the Distributed File System 828, to support the processing of large amounts of data. The Resource Manager 826 can manage clustered or grouped compute units 806 that are associated with, or allocated to support, the Distributed File System 828 and the Job Scheduler 822. The Resource Manager 826 can coordinate with the Resource Orchestrator 814 to manage these associated or allocated compute resources.
[0088] Software 832 can be contained in software layer 830 and may include software used by at least sections of a computing unit 806, one or more computing units 806, groups of computing units 806, and / or the distributed file system 828 of framework layer 820. One or more types of software may include, among others, internet web search software, email virus scanning software, database software, and streaming video content software.
[0089] Applications 842 can be contained in the application layer 840 and comprise one or more types of applications used by at least sections of a compute unit 806, one or more compute units 806, groups of compute units 806, and / or a distributed file system 828 of the framework layer 820. One or more types of applications can include, among others, any number of genomics applications, cognitive computing applications, and machine learning applications, including training or inference software, machine learning software (for example, PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0090] Each of the Configuration Manager 824, Resource Manager 826, and Resource Orchestrator 814 can implement any number and type of self-modifying actions based on any amount and type of data collected in a technically feasible manner. Self-modifying actions can relieve a Data Center 800 operator of potentially making poor configuration decisions and potentially avoid underutilized and / or underperforming sections of a data center.
[0091] The Data Center 800 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or deriving information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using software and computational resources described above in relation to the Data Center 800. Trained machine learning models corresponding to one or more neural networks may be used to derive or predict information using resources described above in relation to the Data Center 800 by using weight parameters calculated by one or more training techniques described herein.
[0092] The Data Center 800 can support CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware (for example, embodiments in the Fig. 9-21) to perform some or all of the processes and techniques described elsewhere herein, such as, but not limited to, training and / or inference using the resources described above. In addition, one or more of the software and / or hardware resources described above may be configured as a service to allow users to train or infer information, such as, but not limited to, image recognition, speech recognition, or other artificial intelligence services.
[0093] In at least one embodiment, the processor 808 may comprise one of the processors listed below and / or one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform one of the operations described above or elsewhere herein. In at least one embodiment, the processor 808 is configured by the software 832 to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform one of the operations described above or elsewhere herein. The data center 800 may include logic, CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware (for example, embodiments in the Fig. 9-21) to perform any of the operations described above or elsewhere herein. PROCESSORS
[0094] The following figures show, without limitation, example processors and processing systems that can be used to utilize one or more SNR values to select one or more neural networks, to generate one or more channel estimates, and / or to otherwise perform some or all of the processes, operations, and / or techniques described elsewhere herein. Example processors and processing systems can be configured by software to use one or more SNR values to select one or more neural networks, to generate one or more channel estimates, and / or to otherwise perform any of the operations described above or elsewhere herein. Processors and processing systems can be logic, central processing units (CPUs), application-specific integrated circuits (ASICs), graphics processing units (GPUs), field-programmable arrays (FPGAs), XPUs (i.e.,any computer architecture that best meets the requirements of an application) or other hardware (for example, embodiments in the . Fig. 9-21) to perform any of the operations described above, below, or elsewhere in this description. The processors and / or processing systems described herein may include one or more circuits that can be used to utilize one or more signal-to-noise ratios (SNR values) to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. As used herein, one or more circuits may be configured by software to utilize one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. Fig. 26A and Fig. Figure 26B shows the logic 2615, which, as described elsewhere herein, can be used in one or more devices to perform operations such as, but not limited to, those described herein, according to at least one embodiment. Logic may, for example, refer to any combination of software logic, hardware logic, and / or firmware logic to provide the functions and / or operations described herein, wherein the logic may be implemented collectively or separately as a circuit that is part of a larger system, such as an integrated circuit (IC), an application-specific integrated circuit (ASIC), a field-programmable array (FPGA), a system-on-a-chip (SoC), or one or more processors (e.g., CPU, GPU).
[0095] Fig. Figure 9 shows a processor, which is a System-on-a-Chip (SOC) 900 (which may be called a System-on-Chip, Superchip, or by another name), according to at least one embodiment. SOC 900 may comprise a processor complex 910 and a processor complex 940. SOC 900 may comprise any number of processor complexes 910 and / or processor complexes 940, which may comprise any number of processors described herein, such as, but not limited to, those in the Fig. 9-21, in any combination. For example, Processor 910 may include a central processing unit (CPU), and Processor 940 may include a graphics processing unit (GPU). Alternatively, Processor 910 may include a GPU, and Processor 940 may include a GPU. SOC 900 may include any number of Display Controllers 992, any number of Multimedia Engines 994, any number of I / O Interfaces 970, any number of Memory Controllers 980, and any number of Fabrics 960 in any combination. For explanatory purposes, multiple instances of similar objects are indicated here with reference numbers to identify the object and, where necessary, with numbers in parentheses to identify the instance. SOC 900 may include a processor made by Broadcom in Palo Alto, California.
[0096] The 910 processor can include a CPU, the 940 processor can include a GPU, and the 900 SOC can include a processing unit that integrates the 910 and 940 onto a single chip. Some tasks can be assigned to the 910 processor, and other tasks can be assigned to the 940 processor. The 910 processor can be configured to run the main control software associated with the 900 SOC, such as, but not limited to, an operating system. The 910 processor can be the main processor of the 900 SOC, controlling and coordinating the operations of other processors. The 910 processor can issue instructions that control the operation of the 940 processor to perform some or all of the operations described herein.The 910 processor can be configured to execute host-executable code derived from CUDA or other source code (for example, HIP source code), and the 940 processor can be configured to execute device-executable code derived from CUDA or other source code to perform any of the operations described herein.
[0097] The 910 processor can include 920(1)-920(4) cores and a cache (for example, an L3 cache) 930 for storing information to perform the operations described herein. The 910 processor can include any number of 920 cores and any number and type of caches in any combination. The 920 cores can be configured to execute instructions from a specific instruction set (“ISA”) to perform some or all of the operations described herein. Each 920 core can include one CPU core. The 920(1)-920(4) cores can be referred to as arithmetic units or computing units. The 900 system-on-a-chip (SoC) can include any number of 910 processor complexes, 960 architectures, 970 I / O interfaces, and 980 memory controllers.
[0098] Each Core 920 can include a Fetch Unit 922, an Integer Execution Engine 924, a Floating-Point Execution Engine 926, and an L2 Cache 928. The Fetch / Decode Unit 922 can fetch instructions to perform some or all of the operations described here (such as, but not limited to, an API compiled to instructions), decode such instructions, generate micro-operations, and forward separate micro-instructions to the Integer Execution Engine 924 and / or the Floating-Point Execution Engine 926. The Fetch / Decode Unit 922 can simultaneously send one micro-instruction to the Integer Execution Engine 924 and another micro-instruction to the Floating-Point Execution Engine 926. The Integer Execution Engine 924 can perform integer and memory operations. The 926 floating-point execution engine can perform floating-point and vector operations.The 922 retrieval / decoding unit can send microinstructions to one or more execution engines, replacing both the 924 integer execution engine and the 926 floating-point execution engine.
[0099] Each kernel 920(i), where i is an integer representing a specific instance of kernel 920, can access the L2 cache 928(i) that contains kernel 920(i). Each kernel 920 contained in a kernel complex 910(j), where j is an integer representing a specific instance of kernel complex 910, can be linked to other kernels 920 contained in kernel complex 910(j) via an L3 cache 930(j) contained within kernel complex 910(j). Kernels 920 contained in kernel complex 910(j), where j is an integer representing a specific instance of kernel complex 910, can access the entire L3 cache 930(j) contained within kernel complex 910(j). The L3 cache 930 can contain any number of slices.
[0100] The Processor 940 can be a graphics complex configured to perform computational operations (such as those involved in the operations described here) in a highly parallel manner. The Processor 940 can be configured to perform graphics pipeline operations, such as, but not limited to, drawing commands, pixel operations, geometric calculations, and other operations related to rendering an image on a display. The Processor 940 can be configured to perform non-graphics-related operations, such as, but not limited to, neural network training and / or simulations. The Processor 940 can be configured to perform both graphics-related and non-graphics-related operations.
[0101] The 940 processor complex can include any number of 950(1)-950(N) processing units, where N is any integer greater than 1, and an L2 cache 942. The 950 processing units can share the L2 cache 942, which can store information used to perform some or all of the operations described here. The L2 cache 942 can be partitioned. The 940 processor complex can include any number of 950 processing units and any number (including zero) and type of caches. The 940 processor complex can include any amount of dedicated graphics hardware.
[0102] Each arithmetic unit 950 can include any number of SIMD units 952(1)-952(N), where N is any integer greater than 1, and a shared memory 954. Each SIMD unit 952 can implement a SIMD architecture and be configured in parallel for some or all of the operations described here. Each arithmetic unit 950 can execute any number of thread blocks, but each thread block can be executed on a single arithmetic unit 950, although in some embodiments a thread block can be executed on multiple arithmetic units. A thread block can include any number of execution threads. A workgroup can be a thread block. Each SIMD unit 952 can execute a group of threads.A group of threads (for example, 16 threads), also known as a warp, subgroup, or wavefront (as used by AMD and Intel, for example), where each thread in the warp, wave, subgroup, or wavefront can belong to a single thread block and is configured to process a different data set based on a single instruction. Predication can be used to disable one or more threads in a warp, subgroup, or wavefront. A lane can be a thread. A work item can be a thread, for example, but is not limited to, for example, with OpenCL. Different warps, subgroups, or wavefronts in a thread block can be synchronized and communicate with each other via shared memory 954.Each Computing Unit 950 can comprise one or more thread block clusters, where a thread block cluster enables programmatic control of locality with a granularity greater than that of a single thread block on a single streaming multiprocessor (SM). Thread block clusters (also referred to simply as "clusters") allow multiple thread blocks running concurrently on streaming multiprocessors to synchronize and jointly retrieve, exchange, or otherwise use data.In at least one embodiment, streaming multiprocessors (“SMs”) can be streaming microprocessors, stream processors (“SPs”), stream processing units (“SPUs”), compute units (“CUs”), execution units (“EUs”) and / or slices, where a slice in this context can denote a section of the processing resources in a processing unit (for example, 16 cores, a ray tracing unit, a thread director or scheduler).
[0103] Structure 960 can be a system interconnect that facilitates data and control transfers across the processor complex 910, the processor complex 940, the I / O interfaces 970, the memory controllers 980, the display controller 992, and the multimedia engine 994 to perform, for example, some or all of the operations described herein. SOC 900 can, in addition to or instead of Structure 960, include any number and type of system interconnects that facilitate data and control transfers across any number and type of directly or indirectly connected components, which may be located inside or outside of SOC 900. The I / O interfaces 970 can represent any number and type of I / O interface (for example, PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). Various types of peripheral devices can be connected to the I / O interfaces 970.Peripheral devices that can be connected to I / O interfaces 970 include keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interfaces, etc.
[0104] The display controller 992 can display images on one or more display devices, such as, but not limited to, a liquid crystal display (“LCD”). The multimedia engine 994 can include any number and type of circuitry related to multimedia, such as, but not limited to, a video decoder, a video encoder, an image signal processor, etc. Memory controllers 980 can facilitate data transfers between the SOC 900 and a unified system memory 990. The processor 910 and the processor 940 can share the unified system memory 990. The unified system memory 990 can include various types of memory devices, including dynamic random-access memory (DRAM) or graphics random-access memory, such as, but not limited to, synchronous graphics random-access memory (SGRAM), which includes double-data-rate graphics memory (GDDR).The Unified System Memory 990 can include a 3D stacked memory, including but not limited to High Bandwidth Memory (HBM), HBM2e or HDM3.
[0105] SOC 900 can implement a memory subsystem comprising any number and type of memory controllers 980 and memory devices (for example, shared memory 954) that can be allocated to a component or shared by multiple components to perform any of the operations described herein. SOC 900 can implement a cache subsystem that can include one or more caches (for example, L2 caches 928, L3 cache 930, and L2 cache 942), each of which can be private to or shared by any number of components (for example, cores 920, core complex 910, SIMD units 952, arithmetic units 950, and processor complex 940).
[0106] In at least one embodiment, SOC 900 may include one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0107] Fig. Figure 10A shows a parallel processor 1000 according to at least one embodiment. The parallel processor 1000 can be implemented using one or more circuits and can be a programmable processor (for example, a CPU and / or GPU), logic, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other hardware (for example, embodiments in the Fig. 9-21) to perform any of the operations described above or elsewhere herein.
[0108] The parallel processor 1000 may include a parallel processing unit 1002 to perform any of the operations described above or elsewhere herein. The parallel processing unit 1002 may include an I / O unit 1004, which enables communication with other devices, including other instances of the parallel processing unit 1002. The I / O unit 1004 may be directly connected to other devices. The I / O unit 1004 may be connected to other devices via a hub or switch interface, such as, but not limited to, a memory hub 1005. Connections between the memory hub 1005 and the I / O unit 1004 may form a communication link 1013.The I / O unit 1004 can be connected to a host interface 1006 and a memory crosspoint 1016, with the host interface 1006 receiving commands to perform processing operations and the memory crosspoint 1016 receiving commands to perform memory operations as instructed.
[0109] When the host interface 1006 receives a command buffer via the I / O unit 1004, the host interface 1006 can instruct operations to execute these commands to a frontend 1008. The frontend 1008 can be coupled to a scheduler 1010 (which can also be called a sequencer) configured to distribute commands or other work items to a processing cluster array 1012. The scheduler 1010 can ensure that the processing cluster array 1012 is properly configured and in a valid state before tasks can be distributed to a cluster of the processing cluster array 1012. The scheduler 1010 can be implemented using firmware logic running on a microcontroller.The microcontroller-implemented Scheduler 1010 can be configured to perform complex scheduling and workload distribution operations with coarse and fine granularity, enabling fast preemption and context switching of threads running on the Processing Array 1012. Host software can check workloads for scheduling on the Processing Array 1012 via one of several graphics processing paths. The workloads can then be automatically distributed across the Processing Array 1012 by the Scheduler 1010 logic within a microcontroller that includes the Scheduler 1010.
[0110] The processing cluster array 1012 can perform any of the operations described above or elsewhere herein and can contain up to "N" processing clusters (for example, cluster 1014A, cluster 1014B to cluster 1014N), where "N" is a positive integer (which may be different from the integer "N" used in other figures). Each cluster 1014A-1014N of the processing cluster array 1012 can execute a large number of threads concurrently. The scheduler 1010 can allocate work to the clusters 1014A-1014N of the processing cluster array 1012 using various scheduling and / or workload distribution algorithms, which can vary depending on the workload for each type of program or calculation.Scheduling can be performed dynamically by the Scheduler 1010 or partially supported by the compiler logic during the compilation of the program logic configured for execution by the Processing Cluster Array 1012. Different clusters 1014A-1014N of the Processing Cluster Array 1012 can be assigned to process different program types or to perform different types of calculations.
[0111] The Processing Cluster Array 1012 can be configured to perform various types of parallel processing operations, such as, but not limited to, any of the operations described above or elsewhere herein. The Processing Cluster Array 1012 can also be configured to perform general-purpose parallel computing operations. For example, the Processing Cluster Array 1012 can include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.
[0112] The 1012 processing cluster array can be configured to perform parallel graphics processing operations. The 1012 processing cluster array can include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. The 1012 processing cluster array can also be configured to run graphics processing-related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The 1002 parallel processing unit can transfer data from system memory to the 1004 I / O unit for processing.During processing, the transferred data can be stored in an on-chip memory (for example, the parallel processor memory 1022) and then written back to the system memory.
[0113] When the parallel processing unit 1002 is used for graphics processing, the scheduler 1010 can be configured to divide a workload into approximately equal-sized tasks to improve the distribution of graphics processing operations across multiple clusters 1014A-1014N of the processing cluster array 1012. Sections of the processing cluster array 1012 can be configured to perform different types of processing. For example, a first section can be configured to perform vertex shading and topology generation, a second section can be configured to perform tessellation and geometry shading, and a third section can be configured to perform pixel shading or other screen-space operations to produce a rendered image for display.Intermediate data generated by one or more of the clusters 1014A-1014N can be stored in buffers so that intermediate data can be transferred between the clusters 1014A-1014N for further processing.
[0114] The processing cluster array 1012 can receive processing tasks to be executed via the scheduler 1010, which receives commands defining processing tasks from the frontend 1008. Processing tasks can include indices of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixels, as well as state parameters and commands that define how data is to be processed (for example, which program is to be executed). The scheduler 1010 can be configured to retrieve indices corresponding to the tasks, or it can receive indices from the frontend 1008. The frontend 1008 can be configured to ensure that the processing cluster 1012 is brought into a valid state before initiating a workload specified by incoming command buffers (for example, batch buffers, push buffers, etc.).
[0115] Each of the one or more instances of the parallel processing unit 1002 can be coupled to a parallel processor memory 1022 to perform one of the operations described above or elsewhere herein. The parallel processor memory 1022 can be accessed via a memory crossbar 1016, which can receive memory requests from both the processing cluster array 1012 and the I / O unit 1004. The memory crossbar 1016 can access the parallel processor memory 1022 via a memory interface 1018. The memory interface 1018 can include multiple partition units (for example, partition unit 1020A, partition unit 1020B through partition unit 1020N), each of which can be coupled to a section (for example, a memory unit) of the parallel processor memory 1022.A number of partition units 1020A-1020N can be configured to correspond to a number of storage units, such that a first partition unit 1020A has a corresponding first storage unit 1024A, a second partition unit 1020B has a corresponding storage unit 1024B, and an Nth partition unit 1020N has a corresponding Nth storage unit 1024N. A number of partition units 1020A-1020N need not be equal to a number of storage units.
[0116] Memory units 1024A-1024N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as, but not limited to, synchronous graphics random access memory (SGRAM), including double data rate graphics memory (GDDR). The memory units 1024A-1024N can also include 3D stacked memory, including, but not limited to, high bandwidth memory (HBM), HBM2e, or HDM3. Render targets, such as, but not limited to, frame buffers or texture maps, can be stored on the memory units 1024A-1024N, allowing the partition units 1020A-1020N to write sections of each render target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 1022.A local instance of the parallel processor memory 1022 can be excluded in favor of a unified memory design that utilizes system memory in conjunction with local cache memory.
[0117] Each of the 1014A-1014N clusters of the 1012 processing cluster array can process data written to one of the 1024A-1024N storage units within the 1022 parallel processor memory. The 1016 memory crosspoint switch can be configured to transfer an output from each 1014A-1014N cluster to any 1020A-1020N partition unit or to another 1014A-1014N cluster, which can then perform additional processing operations on the output. Each 1014A-1014N cluster can communicate with the 1018 memory interface via the 1016 memory crosspoint switch to read from and write to various external storage devices.The memory crosspoint 1016 can communicate with the I / O unit 1004 via a connection to the memory interface 1018, as well as with a connection to a local instance of the parallel processing unit memory 1022. This allows processing units within different processing clusters 1014A-1014N to communicate with system memory or other memory that is not local to the parallel processing unit 1002. The memory crosspoint 1016 can use virtual channels to separate data streams between the clusters 1014A-1014N and the partition units 1020A-1020N.
[0118] Multiple instances of the Parallel Processing Unit 1002 can be deployed on a single add-on card, or multiple add-on cards can be interconnected. Different instances of the Parallel Processing Unit 1002 can be configured to work together, even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, some instances of the Parallel Processing Unit 1002 can include higher-precision floating-point units compared to other instances.Systems containing one or more instances of the Parallel Processing Unit 1002 or the Parallel Processor 1000 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld PCs, servers, workstations, game consoles and / or embedded systems.
[0119] Fig. 10A further comprises a block diagram of a partition unit 1020 according to at least one embodiment. The partition unit 1020 is an instance of one of the partition units 1020A-1020N from Fig. 10A. Partition unit 1020 can include an L2 cache 1021, a frame buffer interface 1025, and a ROP 1026 (Raster Operations Unit). The L2 cache 1021 can be a read / write cache configured to perform load and store operations received from the memory crosspoint 1016 and the ROP 1026. Read errors and urgent write-back requests can be sent from the L2 cache 1021 to the frame buffer interface 1025 for processing. Updates can also be sent to a frame buffer for processing via the frame buffer interface 1025. The frame buffer interface 1025 can be connected to any of the memory units in the parallel processor memory, such as, but not limited to, memory units 1024A-1024N (represented as 1024) in Fig. 10A (for example, within the parallel processor memory 1022).
[0120] ROP 1026 can be a processing unit that performs raster operations such as, but not limited to, stenciling, Z-testing, blending, etc. ROP 1026 can then output processed graphics data stored in graphics memory. ROP 1026 can include compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. The compression logic can be lossless, employing one or more of several compression algorithms. The type of compression performed by ROP 1026 can vary based on the statistical properties of the data being compressed. For example, delta color compression is performed on depth and color data on a tile basis.
[0121] ROP 1026 can be used in any processing cluster (for example, cluster 1014A-1014N in Fig. 10A) instead of the partition unit 1020. Read and write requests for pixel data can be transferred via the memory crosspoint 1016 instead of pixel fragment data. Processed graphics data can be displayed on a screen that is stretched for further processing by one or more processors, or for further processing by one of the processing units within the parallel processor 1000. Fig. 10A is stretched.
[0122] In at least one embodiment, the parallel processor 1000 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0123] Fig. Figure 10B comprises a block diagram of a processing cluster 1014 within a parallel processing unit according to at least one embodiment. A processing cluster can be an instance of one of the processing clusters 1014A-1014N from Fig. 10A, which can be used to perform any of the operations described above or elsewhere herein. Processing cluster 1014 can be configured to run many threads in parallel, where "thread" denotes an instance of a particular program running on a specific set of input data. Single-instruction multiple data (SIMD) command-output techniques can be used to support the parallel execution of a large number of threads without providing multiple independent instruction units. Single-instruction multiple thread (SIMT) techniques can be used to support the parallel execution of a large number of generally synchronized threads, using a common instruction unit configured to issue instructions to a set of processing machines within each processing cluster.
[0124] The operation of the processing cluster 1014 can be controlled via a pipeline manager 1032, which distributes processing tasks to SIMT parallel processors. The pipeline manager 1032 can issue instructions from the scheduler 1010. Fig. Processing Cluster 10A receives and manages the execution of these instructions via a graphics multiprocessor 1034 and / or a texture unit 1036. The graphics multiprocessor 1034 can be an example of a SIMT parallel processor. However, various types of SIMT parallel processors with different architectures can be included in the processing cluster 1014. One or more instances of the graphics multiprocessor 1034 can be contained in a processing cluster 1014. The graphics multiprocessor 1034 can process data, and a data crosspoint 1040 can be used to distribute processed data to one of several possible destinations, which may include other shader units. The pipeline manager 1032 can facilitate the distribution of the processed data by specifying destinations for the processed data to be distributed via the data crosspoint 1040.
[0125] Each 1034 graphics multiprocessor within the 1014 processing cluster can include an identical set of functional execution logic (for example, arithmetic logic units, load / store units, etc.) to perform computations for each of the operations described above or elsewhere herein. The functional execution logic can be configured in a pipelined manner, allowing new instructions to be issued before previous instructions have completed. The functional execution logic can support a wide variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and the computation of various algebraic functions. The same functional unit hardware can be used to perform different operations, and any combination of functional units is possible.
[0126] Instructions transmitted to the 1014 processing cluster can form a thread, also known as a warp, subgroup, wave, or wavefront. A group of threads executed across a group of parallel processing engines can be called a thread group. A thread group can execute a common program for different input data. Each thread within a thread group can be assigned to a different processing engine within a 1034 graphics multiprocessor. A thread group can contain fewer threads than the number of processing engines within the 1034 graphics multiprocessor. If a thread group contains fewer threads than the number of processing engines, one or more processing engines may be idle during the cycles in which that thread group is being processed.A thread group can contain more threads than the number of processing engines within the 1034 graphics multiprocessor. If a thread group contains more threads than the number of processing engines within the 1034 graphics multiprocessor, processing can be performed across successive clock cycles. Multiple thread groups can run concurrently on a 1034 graphics multiprocessor.
[0127] The 1034 graphics multiprocessor includes an internal cache for performing load and store operations, such as, but not limited to, any of the operations described above or elsewhere herein. The 1034 graphics multiprocessor may forgo an internal cache and use a cache (for example, L1 cache 1048) within the 1014 processing cluster. Each 1034 graphics multiprocessor may also access L2 caches within partition units (for example, partition units 1020A-1020N in...). Fig. 10A), which can be shared by all processing clusters 1014 and used to transfer data between threads. The graphics multiprocessor 1034 can also access off-chip global memory, which may include one or more local parallel processor memories and / or system memory. Any memory outside the parallel processing unit 1002 can be used as global memory. The processing cluster 1014 can comprise multiple instances of the graphics multiprocessor 1034 and share common instructions and data, which may be stored in the L1 cache 1048.
[0128] Each 1014 processing cluster can include a 1045 MMU (memory management unit), which can be configured to map virtual addresses to physical addresses. One or more instances of the 1045 MMU can be located within the 1018 memory interface. Fig. The MMU 1045 can include a set of page table entries (PTEs) used to map a virtual address to a physical address of a tile and, optionally, a cache row index. The MMU 1045 can include translation lookaside buffers (TLBs) or caches, which may reside within the graphics multiprocessor 1034, the L1 cache 1048, or the processing cluster 1014. A physical address can be processed to distribute access to surface data locally, enabling efficient interleaving of requests between partition units. A cache row index can be used to determine whether a request for a cache row is a hit or a failure.
[0129] A processing cluster 1014 can be configured such that each graphics multiprocessor 1034 is coupled with a texture unit 1036 to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. Texture data can be read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1034 and retrieved as needed from an L2 cache, local parallel processing memory, or system memory. Each graphics multiprocessor 1034 can output processed tasks to the data crosspoint 1040 to make the processed task available to another processing cluster 1014 for further processing or to store the processed task in an L2 cache, local parallel processing memory, or system memory via the memory crosspoint 1016.A PreROP unit 1042 (Pre-Raster Operations Unit) can be configured to receive data from the graphics multiprocessor 1034 and direct data to ROP units, which may be located at partition units as described here (for example, partition units 1020A-1020N in . Fig. 10A). The PreROP unit 1042 can perform optimizations for color mixing, pixel color data organization, and address translations.
[0130] In at least one embodiment, the processing cluster 1014 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0131] Fig. Figure 10C shows a graphics multiprocessor 1034 according to at least one embodiment, for example, for performing one of the operations described above or elsewhere herein. The graphics multiprocessor 1034 can be coupled to the pipeline manager 1032 of the processing cluster 1014. The graphics multiprocessor 1034 can include an execution pipeline that may include, among other things, an instruction cache 1052 (which can, for example, store instructions such as compiled API instructions), an instruction unit 1054, an address allocation unit 1056, a register file 1058, one or more general-purpose graphics processing units (GPGPUs) 1062, and one or more load / store units 1066, wherein one or more load / store units 1066 can perform load / store operations to load / store instructions corresponding to the execution of an operation.GPGPU cores 1062 and load / store units 1066 can be coupled via a memory and cache interconnect 1068 with a cache 1072 and a shared memory 1070. GPGPU cores 1062 can be part of a SoC, such as, but not limited to, the integrated circuit 900 in [reference missing]. Fig. 9.
[0132] The instruction cache 1052 can receive a stream of instructions (for example, to perform one of the operations described above or elsewhere herein) from the pipeline manager 1032 for execution. Instructions can be cached in the instruction cache 1052 and forwarded to an instruction unit 1054 for execution. The instruction unit 1054 can send instructions as thread groups (for example, warps, subgroups, wavefronts, or waves), with each thread in the thread group assigned to a different execution unit within the GPGPU cores 1062. An instruction can access any local, shared, or global address space by specifying an address within a unified address space. The address mapping unit 1056 can be used to translate addresses in a unified address space into a unique memory address that load / store units 1066 can access.
[0133] Register file 1058 can provide a set of registers for functional units of the 1034 graphics multiprocessor. Register file 1058 can provide temporary storage for operands associated with data paths of functional units (for example, 1062 GPGPU cores, 1066 load / store units) of the 1034 graphics multiprocessor. Register file 1058 can be partitioned among the individual functional units, so that each functional unit is allocated a dedicated section of register file 1058. Register file 1058 can be partitioned among different warps (which may be referred to as wavefronts, subgroups, and / or waves or threads) executed by the 1034 graphics multiprocessor.
[0134] GPGPU 1062 cores can each include floating-point units (FPUs) and / or integer arithmetic logic units (ALUs) that can be used to execute instructions from the 1034 graphics multiprocessor. GPGPU 1062 cores can have a similar architecture or differ in their architecture. A first section of the GPGPU 1062 cores can include a single-precision FPU and an integer ALU, while a second section of the GPGPU cores can include a double-precision FPU. FPUs can implement floating-point arithmetic according to the IEEE 754-2008 standard or enable variable-precision floating-point arithmetic. The 1034 graphics multiprocessor can additionally include one or more fixed-function or special-purpose units to perform specific functions, such as, but not limited to, copy-rectangle or pixel-mixing operations.One or more of the GPGPU 1062 cores can also include fixed function logic or special function logic.
[0135] GPGPU Cores 1062 can include SIMD logic capable of applying a single instruction to multiple data sets. GPGPU Cores 1062 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. SIMD instructions for GPGPU Cores can be generated at compile time by a shader compiler or automatically when running programs written and compiled for Single Program Multiple Data (SPMD) or SIMT architectures. Multiple threads of a program can be configured for a SIMT execution model that can be executed via a single SIMD instruction. For example, eight SIMT threads performing the same or similar operations can run in parallel over a single SIMD8 logic unit.
[0136] The memory and cache interconnect 1068 can comprise an interconnection network that connects each functional unit of the graphics multiprocessor 1034 to the register file 1058 and the shared memory 1070. The memory and cache interconnect 1068 can be a crossbar interconnect, allowing the load / store unit 1066 to perform load and store operations between the shared memory 1070 and the register file 1058. The register file 1058 can operate at the same frequency as the GPGPU cores 1062, enabling very low latency data transfer between the GPGPU cores 1062 and the register file 1058. The shared memory 1070 can be used to facilitate communication between threads running on functional units within the graphics multiprocessor 1034.Cache 1072 can be used, for example, as a data cache to temporarily store texture data between functional units and texture unit 1036. Shared memory 1070 can also be used as a program-managed cache. Threads running on GPGPU cores 1062 can programmatically store data in shared memory in addition to the automatically cached data stored in cache 1072.
[0137] A parallel processor or GPGPU described here can be communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern matching operations, and various general-purpose GPU (GPGPU) functions. A GPU can be communicatively coupled to host processors / cores via a bus or other interconnect (for example, a high-speed connection such as PCIe or NVLink). A system-on-a-chip (SoC) can include a parallel processor or GPGPU as described here, with the parallel processor or GPGPU running on the SoC. A GPU can be integrated as cores on a package or chip and communicatively coupled to cores via an internal processor bus / interconnect within the package or chip. Regardless of how a GPU is connected, processor cores can assign work to it in the form of instruction sequences.The GPU can then use dedicated circuitry / logic to efficiently process these commands / instructions and perform one of the operations described above or elsewhere herein.
[0138] In at least one embodiment, the graphics multiprocessor 1034 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0139] Fig. Figure 11 shows a processor 1100 according to at least one embodiment. The processor 1100 may comprise a hybrid-architecture processor (for example, Lunar Lake or Meteor Lake) from Intel Corporation in Santa Clara, California, or another processor that shares at least some of the components described herein. The processor 1100 may include one or more processing units (CPU 1102), one or more graphics processing units (GPU 1106), and / or one or more neural processing units (NPU 1108), which may, for example, be a dedicated AI accelerator that offloads artificial intelligence (AI) workloads from the CPU 1102 and the GPU 1106. The processor 1100 may use instructions that, when executed, cause the processor 1100 and / or one of its components to execute some or all of the processes and techniques described herein.The 1100 processor can include any number of 1110 memory and cache units to facilitate interprocessing between different components of the 1100 processor. The memory and cache on the 1100 processor can include one or more cache levels (for example, L1, L2, L3, and / or last-level cache) and high-bandwidth memory (for example, HBM2e or HBM3) in any combination.With respect to the Processor 1100 and each of its components described above or elsewhere herein, one or more of the APIs described herein may, for example, be compiled into instructions that may be fetched by an instruction fetch logic or similar, decoded by a processor decoder or similar, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or similar, executed by an execution logic or similar, reordered, and then withdrawn by an execution logic or similar. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or inside the Processor 1100 (e.g., in the cache and / or memory). A result of the API(s) may then be stored in memory inside or outside the Processor 1100, including registers, DRAM, flash, SRAM, cache, or other memory.One or more of the APIs described here may contain a call.
[0140] The 1100 processor can include compute engines designated as 1102 CPUs and contain any number of cores, for example, up to 16 cores / 22 threads. The cores in the 1102 CPU can be classified as P cores (performance), E cores (efficiency), and LP-E cores (low-power efficient). Performance cores can be used for low-latency, high-intensity, single-threaded workloads, while efficiency cores can be used for lower-intensity, multi-threaded workloads. Low-power efficient cores can be used for scalable multi-threaded performance and offloading background tasks. P cores can be used for single- and limited-threaded performance, while E and LP-E cores can be used for multi-threaded throughput and energy efficiency.
[0141] The GPU 1106 can include any number of graphics engines, such as, but not limited to, Intel® Arc™ Graphics Engines (Xe LPG) with 8 Xe cores (up to 128 execution units or EUs). As shown in Fig. As shown in Figure 11, the GPU 1106 can include vector engines 1110 and matrix engines 1112, which can, for example, execute FP, INT, and matrix operations all concurrently, separately, or in batches. The GPU 1106 can include a load / store unit 1114, as well as other memory, such as, but not limited to, an instruction cache (I$) 1116 and an L1 cache / subsystem local memory (SLM) 1118, which can, for example, store instructions for performing any of the operations described above or elsewhere herein.
[0142] The NPU 1104 can include one or more integrated Intel® AI Boost neural processing units (NPUs). The NPU 1104 can be mapped to a host processor as an integrated PCIe device. The NPU 1104 can include one or more (e.g., two) Neural Compute Engine (NCE) tiles 1130. Each tile can be configured with any combination of (e.g., 2000) Multiplication-Accumulation Engines (MACs) 1134, a post-processing engine (not shown), an AI DSP processor (not shown), and memory (2 MB dedicated SRAM) per tile, as shown in Fig. Figure 11 shows the following. For general computing needs, Neural Compute Engines 1130 can include an interference pipeline 1132, an activation function (AF) 1136, a data conversion 1138, a load / save function 1140, and Streaming Hybrid Architecture Vector Engines (SHAVE) 1128 for high-performance parallel computing, which can include DMA (Direct Memory Access) Engines 1124 for transferring data between system memory DRAM (Dynamic Random Access Memory) 1126 and a software-controlled cache. The integrated MMU (Memory Management Unit) 1122 plus IOMMU (Input-Output Memory Management Unit) (not shown) can support multiple concurrent hardware contexts and provide security isolation between execution contexts according to the MCDM (Microsoft Compute Driver Model) architecture.The Processor 1100 may also include a media unit (not shown) that is contained in XCDs or other components of the Processor 1100, or separate from them, to enable video playback and video processing of compressed or uncompressed data, for example, using HEVC, AV1, VP9 and AVC hardware-accelerated decoding support and HEVC, VP9 and AVC hardware-accelerated encoding support.
[0143] An Intel® Thread Director, comprising the firmware integrated into the 1100 processor, can prioritize and manage workload distribution and send tasks to optimized cores. For example, Thread Director can combine P cores, E cores, and / or LP-E cores (as described above) with task scheduling capabilities and the ability to send less demanding tasks to E cores or LP-E cores. Intel® Deep Learning Boost (Intel® DL Boost) (not shown) can provide integrated AI acceleration for training and inference workloads and can include support for the VNNI (for CPU) and DP4a (for GPU) instruction sets. This instruction set can be optimized with the OpenVINO™ Toolkit and oneAPL to accelerate INT8 inference. A software stack, such as that described elsewhere herein, can be used to enable AI inference with the OpenVINO™ Toolkit.The 1100 processor can be configured to run an application, such as, but not limited to, a CUDA program.
[0144] In at least one embodiment, the processor 1100 may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0145] The 1100 processor can alternatively comprise a processor based on Qualcomm Corporation's AI Engine Direct architecture in Santa Clara, California, or another processor that shares at least some of the components described herein. It can include any number of NPUs, GPUs, CPUs, and other associated components, such as, but not limited to, the 1104 NPU as a Hexagon NPU, the 1106 GPU as an Adreno GPU, the 1102 CPU as a Kryo or Qualcomm Oryon CPU, a Qualcomm Sensing Hub (not shown), and a 1110 Memory Subsystem in any combination.The Hexagon NPU 1104 can include a power rail, a microtile inference unit, a hardware acceleration unit, a tensor unit, a scalar unit, and a vector unit (all not shown), which may have dedicated memory or shared memory (such as cache or storage, like HBM3) to store instructions for performing any of the operations described above or elsewhere herein. The Adreno GPU 1106 can provide graphics and parallel processing for AI in formats such as, but not limited to, 32-bit floating-point (FP32), 16-bit floating-point (FP16), and 8-bit integer (INT8). Kryo or Qualcomm Oryon CPUs 1102 can run AI workloads and handle contextualization for ubiquitous generative AI applications.The CPU 1102 may also include an instruction fetch unit, a rename and undo unit, a memory management unit, a vector execution unit, an integer execution unit, and a load / store unit for processing and instruction management. With respect to the processor 1100 and each of its components described above or elsewhere herein, one or more of the APIs described herein may, for example, be compiled into instructions that can be fetched by the instruction fetch unit, decoded by a processor decoder or equivalent, scheduled for execution (for example, sequentially or out of sequence) by a scheduler or equivalent, executed by execution logic or equivalent, reordered, and then undoed by the rename and undo unit.APIs (and / or compiled instructions including APIs) can be stored in any memory outside or inside the processor 1100 (for example, in the cache and / or memory). Any number of CPU cores 1102 can be comprised of any number of CPU clusters, which can be coupled to the memory and / or cache, such as, but not limited to, a shared L2 cache. The memory can be separate or shared; for example, CPU clusters of CPU cores 1102 can be coupled to a memory subsystem 1110, which can comprise a structure, a system-level cache, and any number of memory management units that can, for example, read and write memory (such as DRAM).The Qualcomm Sensing Hub (not shown) includes micro-NPUs, a power rail, and conventional sensors (a gyroscope, an accelerometer, even a barometer) carrying voice and data streams. The Memory Subsystem 1110 can include memory and cache on the Processor 1100, comprising one or more cache levels (for example, L1, L2, L3, and / or last-level cache) and high-bandwidth memory (for example, HBM2e or HBM3) in any combination, for example, to store information and / or instructions for performing any of the operations described above or elsewhere herein. All or part of the memory and / or cache in the Memory Subsystem 1110 can be shared or used individually by one or more components (for example, GPU 1106, NPU 1104, and CPU 1102) on the Processor 1100.
[0146] The Qualcomm AI Engine 1100 can be programmed and controlled with a software stack to perform some or all of the operations described here. This stack includes, for example, a Qualcomm® Neural Processing SDK for inference, with versions for Android, Linux, and Windows. Developer libraries and services support programming languages, virtual platforms, and compilers. At a lower level of the software stack, the system software includes a basic real-time operating system (RTOS), system interfaces, and drivers. The software stack supports various operating systems, including Android, Windows, Linux, and QNX, as well as deployment and monitoring infrastructures such as Prometheus, Kubernetes, and Docker. OpenCL and DirectML are supported for direct cross-platform access to the GPU 1106. For the CPU 1102, optimizations to the LLVM compiler infrastructure enable accelerated and efficient AI inference.With respect to the Qualcomm AI Engine 1100 and all its components described above or elsewhere herein, one or more of the APIs described herein can, for example, be compiled into instructions that can be fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then withdrawn by a withdrawal logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) can be stored in any memory outside or inside the Qualcomm AI Engine 1100 (e.g., in the cache and / or memory).A result from API(s) can then be stored in memory inside or outside the Qualcomm AI Engine 1100, including registers, DRAM, Flash, SRAM, cache or other memory.
[0147] In at least one embodiment, the Processor 1100 or the Qualcomm AI Engine 1100 may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0148] Fig. Figure 12A shows a processor 1200 according to at least one embodiment. The processor 1200 may comprise a processor from Intel Corporation's scalable family in Santa Clara, California, or another processor that shares at least some of the components described herein. The processor 1200 may comprise one or more cores 1212(1)-1212(N), where N is any integer greater than 1 capable of performing the operations described elsewhere herein. The cores 1212(1)-1212(N) may be interconnected via ring interconnects and / or grid interconnects. In a grid interconnect, an arrangement of vertical and horizontal communication paths may enable traversal from one core to another 1212(1)-1212(N) via a shortest path (jumping along the vertical path to the correct row and jumping along the horizontal path to the correct column).In mesh interconnects, a chip can house cores 1212(1)-1212(N) and comprise a grid of converged mesh stops (CMS) that can be associated (e.g., 1:1) with cores 1212(1)-1212(N). Each core can be associated with a slice 1214(1)-1214(N) of a lower-level cache (LLC), or cores 1212(1)-1212(N) can share a cache, such as a lower-level cache. The LLCs 1214(1)-1214(N) can be inclusive, containing blocks in a higher-level cache (e.g., L2 cache), or non-inclusive, containing blocks that may not be present in the higher-level cache. Each core and each LLC slice can include a Caching and Home Agent (CHA) (not shown) that can maintain cache coherence by providing resource scalability over grid connections for the Intel® Ultra Path Interconnect (Intel® UPI 1216) cache coherence functionality.UPI 1216 can provide coherent interconnection for scalable systems and allow multiple processors to share a single common address space via connections, such as, but not limited to, two or three UPI connections per processor.
[0149] The 1200 processor may also include a 1210 system agent, which can host and / or perform various functions, such as, but not limited to, memory management, display functions, and / or input / output (I / O) functions. For example, the 1200 processor may include one or more 1208 integrated memory controllers (IMCs). The 1208 IMC can control and manage memory, such as, but not limited to, different types of memory, such as DDR RAM, like DDR4, or others described elsewhere herein. The 1210 system agent may include a display controller (not shown) to support displays. The 1210 system agent may also include a 1204 PCIe (for example, up to 20 PCIe lanes), which can be connected to an external dedicated graphics connector, for example, via a 1206 DMI bus (for example, Intel's DMI 3.0 bus).The System Agent 1210 can include an integrated processing unit (IPU) (not shown) containing an on-chip image processor (ISP). Fabric 1202 can provide scalability for connecting to other nodes (for example, processors such as the Processor 1200) and can be used, for example, with Cornelis Networks, an element of the Intel® Scalable System Framework, which provides the performance for high-performance computing (HPC) workloads and the ability to scale to tens of thousands of nodes.
[0150] Fig. Figure 12B shows components within the core 1212 according to at least one embodiment. The core 1212 can comprise a frontend 1218, a backend or execution engine 1232, and a memory subsystem 1242. The frontend 1218 can supply the execution engine 1232 with operations (for example, operations described elsewhere herein) by decoding instructions stored in memory. For example, the frontend 1218 can include a micro-operation cache path (µOps) and / or a legacy path, as well as a branch prediction unit 1221 that can determine path instructions. A legacy path for instructions can include retrieving variable-length instructions (e.g., x86) from the L1 instruction cache 1220 with instruction retrieval and pre-decoding 1222, queuing the instructions into the instruction queue 1224, and decoding the instructions into µOps using the decoder 1226, which can then be delivered to the allocation queue 1228.Alternatively, a µOPs cache path can include a cache containing pre-decoded µOps (µOps 1230) that can be sent to the allocation queue 1228. The allocation queue 1228 can act as an interface between the frontend 1218 and the execution engine 1232, delivering instructions to the execution engine 1232. For example, one or more of the APIs described here can be compiled into instructions that can be stored, processed, and executed by the frontend 1218 and the execution engine 1232, and stored in the storage subsystem 1242.
[0151] The execution machine 1232 can receive micro-operations into the reorder buffer 1234, which can register, rename, and withdraw micro-operations. From the reorder buffer, micro-operations can be sent to the scheduler 1236, which can be connected to one or more different execution units 1238, which in turn can be connected to the address generation unit (AGU) 1240. The execution units 1238 can, for example, perform basic arithmetic logic unit (ALU) operations, such as multiplications, divisions, and / or more complex operations, such as various vector operations. The scheduler 1236 can queue the micro-operations for one or more execution units 1238, for example, depending on the operations to be executed.
[0152] The 1242 memory subsystem can handle load and store requests as well as sort operations. For example, µOPs can relate to memory accesses (such as load and store) and these can be sent to dedicated scheduler ports that can execute these memory operations. Store and load operations, for example, can be sent to the 1244 load and store buffers. The 1242 memory subsystem can also include a shared or separate L1 data and instruction cache 1246 and an L2 cache 1248 that can be used and shared with the L1 data and instruction cache 1246. As above for Fig. As described in 12A, each 1212 core can be connected to a slice of a third-level cache (for example, LLC 1214) that can be shared by all 1212 cores.
[0153] In at least one embodiment, the processor 1200 may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0154] Fig. Figure 13 shows a Kl accelerator 1300 according to at least one embodiment. The processor 1300 may comprise a processor with Kl accelerator architecture from Intel Corporation in Santa Clara, California, or another processor that shares at least some of the components described herein. The Kl accelerator 1300 may use instructions which, when executed by the Kl accelerator 1300, cause the Kl accelerator 1300 to execute some or all of the processes and techniques described elsewhere herein.With respect to the Kl Accelerator 1300 and each of its components described above or elsewhere herein, one or more of the APIs described herein can, for example, be compiled into instructions that can be fetched by an instruction fetch logic or similar, decoded by a processor decoder or similar, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or similar, executed by an execution logic or similar, reordered, and then withdrawn by a retract logic or similar. API(s) (and / or compiled instructions comprising API(s)) can be stored in any memory outside or inside the Kl Accelerator 1300 (e.g., in the cache and / or memory). A result of the API(s) can then be stored in memory inside or outside the Kl Accelerator 1300, including registers, DRAM, flash, SRAM, cache, or other memory.The Kl Accelerator 1300 can comprise one or more computing chips, which can include homogeneous or heterogeneous processors. Computing chips can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), or combinations of both.
[0155] In at least one embodiment, the computing units can include computation engines for performing AI calculations. In at least one embodiment, the computing units of the AI accelerator 1300 can be divided into any number (for example, four) of clusters, which can be designated as DCORE (Deep Learning Core) 1306 and can contain any number of matrix multiplication engines (MMEs) 1308, tensor processor cores (TPCs) 1310, memory management units 1312, and L2 cache 1314 in any combination. MME(s) 1308 can perform operations that use matrix multiplication, such as fully connected layers, convolutions, and bundled general matrix multiplications (GEMMs).MMEs 1308 can be equipped with multiplication-accumulation units (MACs) (not shown) that can perform general matrix multiplication (GEMM) operations, such as, but not limited to, AxB multiplication, where a tensor C[NxM] is generated from two input tensors, A[NxK] and B[KxN]. MMEs 1308 can be programmed with array dimensions, memory locations, data types, and various execution operands. MMEs 1308 can retrieve the tensors A and B from memory and pull them into their streaming buffers so that the matrix multiplication can be performed in parallel by MACs. MMEs 1308 can write the tensor C back into memory after completion.TPC(s) 1310 can include any number of scalar units for performing scalar operations, any number of vector units for performing vector operations, any number of register files or local memory units (for example, a vector local memory), and instruction load and store components that can be coupled to memory or cache (for example, HBM, L3 cache, and / or L2 cache) (all not shown). TPCs can support various types of parallel processing, such as Very Long Instruction Word (VLIW) and Single Instruction Multiple Data (SIMD), which supports data types such as FP32, BF16, FP16, and FP8 (both E4M3 and E5M2), UINT32, INT32, UINT16, INT16, UINT8, and INT8. Any number of compute chips can be interconnected.Interconnection that can link computing units can be achieved via an interposer bridge, which is transparent to software, for example.
[0156] The memory on the Kl Accelerator 1300 can include one or more cache levels (for example, L1, L2, L3, and / or last-level cache) and high-bandwidth memory (for example, HBM2e or HBM3) in any combination. Memory and / or cache systems can be unified or separate. The compute units of the Kl Accelerator 1300 can include on-die memory containing one or more cache levels (for example, two levels). The on-die SRAM, or other memory described elsewhere herein, can be used as a unified, accessible last-level cache (L3) or divided into L2 cache slices accessible to groups of MMEs 1308 and TPCs 1310. The use of on-die memory as L2 or L3 cache can be fully configured by software, which can dynamically determine the optimal cache allocation for each I / O tensor.The AI Accelerator 1300 can include one or more Memory Management Units (MMUs) 1322 for managing memory, for example to enable the memory subsystem of the AI Accelerator 1300 to operate in a virtual space when accessing VRAM.
[0157] The Kl Accelerator 1300 can include a communication port (for example, a PCIe Gen5 X16 port) 1302 for communication with a host and a scheduling and synchronization unit 1304. The Kl Accelerator 1300 can include a media unit 1316, which can include any number or combination of media decoder engines (DECs) 1320 and rotator engines (ROTs) 1318. The Kl Accelerator 1300 can include a network unit 1324, which can include any number or combination of network ports 1326 and associated RDMA engines 1328, L2 cache, and memory stacks (for example, HBM2e or HBM3). The Kl Accelerator 1300 can include a programmable control path unit (not shown) to manage the parallel and efficient execution of different engines.The control path can include transmission queues (SQs) that can be issued by the runtime system, completion queues (CQs) that can be used for reporting job completion, a programmable scheduling mechanism that can be used for task scheduling, a programmable hardware synchronization mechanism or "Sync Manager" (SM), and a programmable interrupt service mechanism or "Interrupt Manager (INTR)" that can allow the passing of asynchronous events to drivers.
[0158] The Kl Accelerator 1300 can include media decoding units that support video formats such as, but not limited to, HEVC, Progressive H.264, SVC Base Layer, MVC, VP9, JPEG, and Progressive JPEG. The Kl Accelerator 1300 can support post-processing of decoded media streams, such as, but not limited to, image resizing (changing the size of an image), vertical and horizontal scaling with different scaling ratios, image enlargement, image cropping, bilinear scaling, and Lancos scaling. The Kl Accelerator 1300 can implement two post-processing channels per decoder unit: one with scalar (up and down) and one for outputting the original image only.The Kl Accelerator 1300 can include a hardware rotation engine that performs the following transformations of an input image: 2D rotation, 3D rotation, projection, distortion and rectification of images, recalculation of input data at user-defined coordinates, and rescaling.
[0159] RDMA 1328 over Converged Ethernet on the AI Accelerator 1300 enables scaling from a single node (i.e., a single AI Accelerator 1300) to hundreds or thousands of nodes or AI Accelerators 1300. The NW subsystem 1324 can include an Intel® Gaudi® Communication Library (IGCL), a master conductor that coordinates data movement, and a programmable scheduler mechanism that enables smooth engine activation while maintaining task dependencies. An accelerator network subsystem can include Gigabit Ethernet NIC ports 1326, a Layer 2 MAC (not shown), and RDMA engines 1328. The AI Accelerator 1300 can include aggregation engines for performing summation activities. All engines in the 1300 processor can work in parallel; for example, MME(s) 1308, TPC(s) 1310 and NIC(s) 1326 can all work simultaneously.There can be a dependency between operations executed on different engines; for example, the output of one engine can be used as input for another, and / or MME, TPC, and NIC can be scheduled to run in parallel. Once one engine has completed its execution operation, another engine can be scheduled to begin the next operation (immediately after receiving its inputs).
[0160] The AI accelerator 1300 can be operated and controlled using a software layer 1328, which includes low-level components such as a graph compiler, an automatic kernel fuser, and a library of pre-compiled kernels, as well as integration with AI ecosystems such as PyTorch, DeepSpeed, Hugging Face, vLLM Ray, and others, or as described elsewhere in this document regarding software and programming platforms. Software layer 1328 can include implementations of algorithms such as, but not limited to, Paged Attention, Flash Attention, and others. Software layer 1328 can generate optimized binary code that implements a specific model topology, such as, but not limited to, operator fusion, data layout management, parallelization, pipelining, memory management, and graph-level optimizations.
[0161] In at least one embodiment, the Kl accelerator 1300 may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0162] A neuromorphic computing system is described that uses a multicore architecture, where each core includes computing elements, including neurons, synapses with on-chip learning capability, and local memory for storing synaptic weights and routing tables. Fig. Figure 14 is a simplified block diagram 1400 showing an example of at least one section of such a neuromorphic computer device 1405 according to at least one embodiment. The neuromorphic computer device 1405 may include a neuromorphic processor from Intel Corporation in Santa Clara, California, or another processor that shares at least some of the components described herein. As shown in this example, a device 1405 may be provided with a network 1410 of multiple neural network cores interconnected by an internal network, such that potentially several different connections between the cores can be defined. For example, a network 1410 of spiking neural network cores may be provided in the device 1405, each core being able to communicate via short packetized spike messages sent from core to core over network channels.Each core (for example, 1415) can have processing resources, memory resources, and logic to implement a certain number of primitive nonlinear temporal computational elements, such as, but not limited to, multiple (for example, 1000+) different artificial neurons (referred to here as "neurons"). For example, each core can be capable of implementing multiple neurons simultaneously, so neuromorphic cores using the device 1405 can implement a multiple of neurons.With respect to the neuromorphic computing device 1405 and each of its components described above or elsewhere herein, one or more of the APIs or equivalents described herein may, for example, be compiled into instructions or equivalents that may be fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then withdrawn by a withdrawal logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or within the neuromorphic computing device 1405 (e.g., in the cache and / or memory).A result of the API(s) can then be stored in a memory inside or outside the neuromorphic computing device 1405, which includes registers, DRAM, flash, SRAM, cache or other memory equivalents.
[0163] Continuing the example from Fig. 14. The neuromorphic computing device 1405 may additionally include a processor 1420 and system memory 1425 to implement one or more components for managing and providing the functionality of the neuromorphic computing device 1405. For example, a system manager 1430 may be provided to manage global attributes and operations of the neuromorphic computing device 1405 (for example, attributes affecting the network of cores 1410, multiple cores in the network 1410, the interconnections of the neuromorphic computing device 1405 with other devices, managing access to the global system memory 1425, among other possible examples). In one example, the system manager 1430 may, among other things, define and provide specific routing tables for different routers in the network 1410, orchestrate a network definition and attributes (for example, weights, decay rates, etc.).), which are to be used in the 1410 network, manage core synchronization and time-division multiplexing management, as well as the routing of inputs to suitable cores.
[0164] As a further example, the neuromorphic computing device 1405 can additionally include a programming interface 1435 through which a user or a system can specify a neural network definition to be applied (for example, via a routing table and individual neuron properties) that is implemented by the grid 1410 of neuromorphic cores. A software-based programming tool can be provided together with or separately from the neuromorphic computing device 1405, through which a user can provide a definition for a specific neural network to be implemented using the network 1410 of neuromorphic cores.The programming interface 1435 can accept input from a programmer, then generate corresponding routing tables and fill the local memory of individual neuromorphic cores (for example, 1415) with specified parameters to implement a corresponding, adapted network of artificial neurons, which is implemented by neuromorphic cores 1415.
[0165] In some cases, the neuromorphic computing device 1405 can advantageously serve as an interface to other devices, including general-purpose computing devices, and cooperate with them to implement specific applications and use cases. Accordingly, in some cases, an external interface logic 1440 may be provided to communicate with one or more other devices (for example, via one or more defined communication protocols). An external interface 1440 can be used to accept input data from another device or an external memory controller that serves as the input data source.The external interface 1440 can be used additionally or alternatively to deliver results or outputs of computations of a neural network implemented using the neuromorphic computing device 1405 to another device (for example, another general-purpose processor implementing a machine learning algorithm) to realize, among other things, additional applications and enhancements.
[0166] As in Fig. Figure 14 shows the network 1410, consisting of multiple neural network cores interconnected by an internal network, as a section of a network structure connecting multiple neuromorphic cores (e.g., 1415 ad). For example, multiple neuromorphic cores (e.g., 1415 ad) can be provided in a grid, each core being connected by a network comprising multiple routers (e.g., 1450). In one embodiment, each neuromorphic core (e.g., 1415 ad) can be connected to a single router (e.g., 1450), and the routers can be connected to at least one other router (as shown in Figure 1410). Fig. (14 shown). In one particular embodiment, for example, four neuromorphic cores (e.g., 1415 ad) can be connected to a single router (e.g., 1450), and each of the routers 1450 can be connected to two or more other routers to form a manycore grid, thereby enabling each neuromorphic core to be interconnected with every other neuromorphic core in the neuromorphic computer device 1405. Since each neuromorphic core can be configured to implement multiple different neurons, the router network of the neuromorphic computer device 1405 can similarly define connections or artificial synapses (or simply “synapses”) between any two of potentially many (e.g., 30,000+) neurons defined using the network of neuromorphic cores 1410 provided in the neuromorphic computer device 1405.
[0167] Fig. Figure 14 shows a block diagram illustrating the internal components of an example implementation of the neuromorphic kernel 1415. In this example, a single neuromorphic kernel can implement a certain number of neurons (for example, 1024) that share the architectural resources of the neuromorphic kernel 1415 in a time-division multiplexed manner. In this example, each neuromorphic kernel 1415 can comprise a processor block 1455 capable of performing arithmetic functions and routing in conjunction with the realization of a digitally implemented artificial neuron, such as, but not limited to, those described herein.Each neuromorphic kernel 1415 can additionally provide local memory in which a routing table for a neural network can be stored and retrieved, the accumulated potential of each soma of each neuron implemented by the kernel 1415 can be tracked, parameters of each neuron implemented by the kernel 1415 can be recorded, among other data and uses. Components or architectural resources of the neuromorphic kernel 1415 can further include an input interface 1465 for receiving input spike messages generated by other neurons on other neuromorphic kernels, and an output interface 1470 for sending spike messages to other neuromorphic kernels via the mesh network 1410. In some cases, the routing logic for the neuromorphic kernel 1415 can be implemented, at least partially, using the output interface 1470.Furthermore, in some cases, the core (for example, 1415) can implement multiple neurons within an exemplary SNN, and some of these neurons can be interconnected. In such cases, spike messages sent between neurons hosted on core 1415 can bypass communication via the routing structure of the neuromorphic computing device 1405 and instead be managed locally on a specific neuromorphic core 1415.
[0168] Each neuromorphic core can additionally include logic to implement an artificial dendrite 1480 and an artificial soma 1485 (referred to here simply as "dendrite" and "soma," respectively) for each neuron 1475. The dendrite 1480 can be a hardware-implemented process that receives pulses from the network 1410. The soma 1485 can be a hardware-implemented process that receives the cumulative neurotransmitter amounts of each dendrite for the current time and further develops the potential state of each dendrite and soma to generate outgoing pulse messages at appropriate times. Dendrite 1480 can be defined for each connection that receives input from another source (for example, another neuron). In one implementation, the dendrite process 1480 can receive and process spike messages when they arrive serially in a time-division multiplexed manner from the network 1410.When impulses are received, the activation of the neuron (tracked by soma 1485 and local memory 1460) can increase. If the neuron's activation exceeds a threshold set for neuron 1475, neuron 1475 can generate a pulse message, which is relayed via output interface 1470 to a fixed group of fanout neurons. The network distributes spike messages to all target neurons, and in response, these neurons can update their activations in a transient, time-dependent manner, and so on. This may cause the activation of some of these target neurons to also exceed the corresponding thresholds and trigger further spike messages, as in real biological neural networks.
[0169] As mentioned above, the neuromorphic computer device 1405 can reliably implement a spike-based model of neural computation. Such models can also be referred to as spiking neural networks (SNNs). In addition to neuronal and synaptic state, SNNs also integrate the concept of time. In an SNN, communication occurs, for example, via event-driven action potentials or spikes, which transmit no explicit information other than the spike time and an implicit source-target neuron pair corresponding to the transmission of the spike. Computation takes place in each neuron as a result of the dynamic, nonlinear integration of the weighted spike input. In some implementations, recursion and dynamic feedback can be integrated into an SNN computational model.Furthermore, various network connectivity models can be used to model different real-world networks or relationships, including fully connected (all-to-all) networks, feedforward trees, completely random projections, small-world networks, and other examples. A homogeneous, two-dimensional network of neuromorphic kernels, such as the example in [reference missing], is one such example. Fig. As shown in Figure 14, the neuromorphic computing device 1405 can advantageously support all these network models. Since some or all cores of the neuromorphic computing device 1405 can be connected, some or all neurons defined in the cores can therefore also be fully connected via a certain number of router hops. Furthermore, the neuromorphic computing device 1405 can include fully configurable routing tables to define a variety of different neural networks, allowing the neurons of each core to distribute their impulses to any number of cores in the grid 1410, thus realizing completely arbitrary graphs.
[0170] In an improved implementation of a system that can support SNNs, such as, but not limited to, a VLSI (Very Large Scale Integration) hardware device, as in the example of Fig. As shown in Figure 14, fast and reliable circuits can be provided to implement SNNs to model information processing algorithms such as those used by a brain, but in a more programmable manner. For example, while a biological brain can only implement a specific set of defined behaviors conditioned over years, a neuromorphic processor device can offer the ability to rapidly reprogram all neural parameters. Accordingly, a single neuromorphic processor can be used to realize a wider range of behaviors than what a single slice of biological brain tissue can offer. This difference can be achieved by using a neuromorphic processor with neuromorphic design implementations that differ significantly from those of neural circuits found in nature.
[0171] For example, a neuromorphic processor can use time-division multiplexing (TDM) in both a Spike communication network and the neuron machinery of the neuromorphic computer device 1405 to implement SNNs. Accordingly, the physical circuitry of the neuromorphic computer device 1405 can be shared by many neurons to achieve a higher neuron density. With TDM, a network can connect N cores with a total cabling length of O(N), whereas discrete point-to-point cabling would scale as O(N²), thus achieving a significant reduction in cabling resources to enable, among other things, planar and non-plastic VLSI cabling technologies.In neuromorphic nuclei, time-division multiplexing can be implemented through dense memory allocation, for example, using static random-access memory (SRAM) with shared buses, address decoding logic, and other multiplexed logic elements. The state of each neuron can be stored in the processor's memory, with the data describing the state of each neuron including, among other things, the state of each neuron's collective synapses, all currents and voltages across its membrane, and other information (such as, but not limited to, configuration and other information).
[0172] A neuromorphic processor can use a "digital" implementation, which differs from other processors that employ more "analog" or "isomorphic" neuromorphic approaches. For example, a digital implementation can integrate synaptic current using digital adder and multiply circuits, as opposed to analog isomorphic neuromorphic approaches that accumulate charge on capacitors in an electrically analogous way to how neurons accumulate synaptic charge on their lipid membranes. The accumulated synaptic charge can then be stored, for example, for each neuron in the local memory of a corresponding nucleus.Furthermore, at the architectural level of an exemplary digital neuromorphic processor, reliable and deterministic operation can be achieved by synchronizing time across a network of cores, ensuring that any two implementations of a design, under identical initial conditions and configurations, produce identical results. Asynchronicity can be maintained at the circuit level to allow individual cores to operate as quickly and freely as possible, while preserving determinism at the system level. Accordingly, the concept of time can be abstracted as a temporal variable in neural computations by separating it from the "wall clock time" used by the hardware to perform the computation. Consequently, some implementations may include a time synchronization mechanism that globally synchronizes neuromorphic cores at discrete time intervals.A synchronization mechanism enables neural computations to be performed as quickly as possible, where there is a divergence between the runtime and the biological time that models a neuromorphic system.
[0173] In operation, the neuromorphic computer device 1405 can begin in a sleep state in which all neuromorphic cores are inactive. As each core asynchronously traverses its neurons, it generates spike messages that a grid interconnection forwards to the corresponding target cores, which contain all the target neurons. The implementation of multiple neurons on a single neuromorphic core can be time-division multiplexed, and a time step can be defined in which all spikes involving multiple neurons can be processed and accounted for using shared resources of a corresponding core.Once each core has finished processing its neurons for a given time step, in some implementations the cores can communicate with neighboring cores (for example, using a handshake), employing synchronization messages to clear a grid of all spike messages in transmission. This allows the cores to reliably determine that all spikes for a given time step have been processed. At this point, all cores can be considered synchronized, allowing them to continue their time step, return to an initial state, and begin the next time step.
[0174] Against this background and as introduced above, a device (for example 1405) can be provided that implements a network 1410 of interconnected neuromorphic cores, with the core 1415 potentially implementing multiple artificial neurons that can be interconnected to implement an SNN.Each neuromorphic nucleus (e.g., 1415) can provide two loosely coupled asynchronous processes: an input dendrite process (e.g., 1480) that receives impulses from network 1410 and applies them to appropriate target dendrite compartments at suitable future times, and an output soma process (e.g., 1485) that receives the accumulated neurotransmitter levels of each dendrite compartment for the current time and further develops the membrane potential state of each dendrite and soma, generating outgoing spike messages at appropriate times (e.g., when a soma threshold potential has been observed). It should be noted that, from a biological perspective, the terms "dendrite" and "soma" used here only approximate the role of these functions and should not be interpreted too literally.
[0175] In at least one embodiment, the neuromorphic computing device 1405 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0176] Fig. Figure 15 is a block diagram of an embodiment of a multi-node network in which remote storage computing according to any embodiment can be implemented. System 1500 can represent a network of nodes that can be used, for example, to perform some or all of the operations described herein. System 1500 can represent a data center. System 1500 can represent a server farm. System 1500 can represent a data cloud or a processing cloud. System 1500 can represent a supercomputer. System 1500 can include dozens, hundreds, or thousands of nodes. The nodes of System 1500 can include processors such as, but not limited to, central processing units (CPUs), graphics processing units (GPUs), or any combination of processors described herein, such as, but not limited to, other processors in the Fig. 9-21. With respect to each of the processors in System 1500 and each of its components described above or elsewhere herein, one or more of the APIs or equivalents described herein may, for example, be compiled into instructions or equivalents that may be fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (for example, sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then withdrawn by a retract logic or equivalent. API(s) (and / or compiled instructions including API(s)) may be stored in any memory outside or inside a processor or node (for example, in the cache and / or memory).A result of API(s) can then be stored in memory inside or outside a processor or node, comprising registers, DRAM, flash, SRAM, cache, or other memory equivalents. The System 1500 can comprise over 9,000 nodes, each containing two Intel Xeon Max processors, six Intel Max-series GPUs, and a uniform memory such as, but not limited to, that used in Intel Corporation's Intel Aurora supercomputer in Santa Clara, California, or any other supercomputer sharing at least some of the components described herein.
[0177] One or more clients 1502 send requests to system 1500 over network 1504. Network 1504 represents one or more local area networks (LANs), wide area networks (WANs), or a combination thereof. Clients 1502 can be human or machine clients that generate requests for system 1500 to perform operations. System 1500 executes applications or data processing tasks requested by clients 1502.
[0178] The System 1500 can comprise one or more racks, which represent structural and interconnection resources for housing and connecting multiple compute nodes. The Rack 1510 can contain multiple nodes 1530. The Rack 1510 can host multiple blade components 1520(0) to 1520(N-1), where N is an integer greater than or equal to 2. Hosting can refer to the provision of power, structural or mechanical support, and interconnection. The Blades 1520(0) to 1520(N-1) can refer to compute resources on printed circuit boards (PCBs), with one PCB housing hardware components for one or more nodes 1530. The Blades 1520(0) to 1520(N-1) may or may not include an enclosure or other "box" that is separate from the Rack 1510. The Blades 1520(0) to 1520(N-1) can include a housing with an exposed connector for connection to the Rack 1510.The System 1500 may or may not include a Rack 1510, and each blade (for example, 1520(0)) may include an enclosure or shell that can be stacked or otherwise arranged in close proximity to other blades, enabling the interconnection of Nodes 1530. The System 1500 can comprise 10,624 compute blades, comprising 63,744 Intel Max Series GPUs and 21,248 Intel Xeon Max CPUs in 166 racks.
[0179] System 1500 can include a fabric 1570, which represents one or more connections for nodes 1530. The fabric 1570 can include multiple switches 1572, routers, or other hardware for forwarding signals between the nodes 1530. Additionally, the fabric 1570 can connect System 1500 to the network 1504 to enable access by clients 1502. Besides the forwarding devices, the fabric 1570 can also include cables, ports, or other hardware devices for connecting the nodes 1530 to each other. The fabric 1570 can have one or more associated protocols for managing the forwarding of signals through System 1500. One or more protocols depend, at least in part, on the hardware equipment used in System 1500.
[0180] As shown, Rack 1510 can include N blades (for example, 1520(0) to 1520(N-1)). In addition to Rack 1510, System 1500 can include Rack 1550. As shown, Rack 1550 can include M blades (for example, 1560(0) to 1560(M-1)). M is not necessarily equal to N; therefore, it is understood that different hardware components can be used and coupled to System 1500 via Structure 1570. Blades 1560(0) to 1560(M-1) can be identical to or similar to blades 1520(0) to 1520(N-1). Nodes 1530 can be any node type described herein and need not all be of the same node type. System 1500 is neither limited to homogeneity nor is it limited to non-homogeneity.
[0181] One node in Blade 1520(0) is shown in detail. However, other nodes in System 1500 may be the same or similar. At least some nodes 1530 may be compute nodes with a processor 1532 and memory 1540. A compute node refers to a node with processing resources (for example, one or more processors) that runs an operating system and can receive and process one or more tasks. At least some nodes 1530 may include storage server nodes with a server as processing resources 1532 and memory 1540. A storage server refers to a node with more storage resources than a compute node, and instead of having processors to perform tasks, a storage server includes processing resources to manage access to storage nodes within a storage server.
[0182] The node 1530 can include an interface controller 1534, which can represent logic for controlling the node 1530's access to the structure 1570. The logic can include hardware resources for connecting to physical interconnect hardware. The logic can include software or firmware logic for managing the interconnection. The interface controller 1534 can include a host interface, which can include a structure interface according to an embodiment described herein.
[0183] Node 1530 can include a memory subsystem 1540. Memory 1540 can include compute resources (comp) 1542, which represent one or more capabilities of memory 1540 for performing memory computations. System 1500 enables remote memory operations, such as, but not limited to, those described elsewhere herein. Thus, nodes 1530 can request memory computations from remote nodes, with data for the computation remaining locally on an executing node instead of being sent via structure 1570 or from memory to a structure interface. In response to the execution of the memory computation, the executing node can provide a result to a requesting node.
[0184] The 1532 processor can comprise one or more separate processors. Each separate processor can comprise a single processing unit, a multi-core processing unit, or a combination thereof. A processing unit can comprise a primary processor, such as, but not limited to, a CPU (central processing unit), a peripheral processor, such as, but not limited to, a GPU (graphics processing unit), or a combination thereof. The 1540 memory can comprise or include memory devices and a memory controller.
[0185] The term "storage device" can refer to different types of storage. Storage devices generally refer to volatile memory technologies. Volatile memory is memory whose state (and therefore the data stored on it) is indeterminate when the power supply is interrupted. Non-volatile memory refers to memory whose state is definite even when the power supply is interrupted. Dynamic volatile memory can update the data stored in a device to maintain its state. An example of dynamic volatile memory includes DRAM (Dynamic Random Access Memory) or a variant thereof, such as, but not limited to, synchronous DRAM (SDRAM).A memory subsystem described here can be compatible with a number of memory technologies, including DDR3 (Dual Data Rate Version 3, originally published by JEDEC (Joint Electronic Device Engineering Council) on 27.June 2007, currently in version 21), DDR4 (DDR Version 4, first specification published in September 2012 by JEDEC), DDR4E (DDR Version 4, extended, currently under discussion at JEDEC), LPDDR3 (Low Power DDR Version 3, JESD209-3B, August 2013 by JEDEC), LPDDR4 (Low Power Double Data Rate (LPDDR) Version 4, JESD209-4, originally published by JEDEC in August 2014), WIO2 (Wide I / A 2 (Widel02), JESD229-2, originally published by JEDEC in August 2014), HBM (HIGH BANDWIDTH MEMORY DRAM, JESD235, originally published by JEDEC in October 2013), DDR5 (DDR Version 5, currently under discussion at JEDEC), LPDDR5 (currently under discussion at JEDEC) HBM2 (HBM Version 2), currently under discussion at JEDEC) or other or combinations of memory technologies and technologies based on derivatives or extensions of such specifications.
[0186] In addition to or as an alternative to volatile memories, in one embodiment the reference to memory devices may refer to a non-volatile memory device whose state is determined even when the power supply is interrupted. In one embodiment, the non-volatile memory device is a block-addressable memory device, such as, but not limited to, NAND or NOR technologies. Thus, a memory device may also include future-generation non-volatile devices, such as, but not limited to, a three-dimensional cross-point memory (3DXP), other byte-addressable non-volatile memory devices, or memory devices that use a chalcogenide phase-change material (for example, chalcogenide glass).In one embodiment, a storage device may be or include a multi-threshold NAND flash memory, NOR flash memory, a single- or multi-stage phase-change memory (PCM) or a single-switch phase-change memory (PCMS), a resistive memory, a nanowire memory, a ferroelectric transistor random-access memory (FeTRAM), a magnetoresistive random-access memory (MRAM) incorporating memristor technology, or a spin-transfer-torque (STT) MRAM, or a combination of any of the above or any other memory.
[0187] In at least one embodiment, the System 1500 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0188] Fig. Figure 16 shows an accelerated processing unit 1600 according to at least one embodiment. The accelerated processing unit 1600 may comprise a processor based on the CDNA architecture of AMD Corporation in Santa Clara, California, or another processor that shares at least some of the components described herein. The accelerated processing unit 1600 may include one or more accelerator complex dies (XCDs) 1604 for performing operations otherwise described herein, such as, but not limited to, graphics processing and / or parallel processing, as well as instruction-level parallel computations, including support for a wide range of precisions (INT8, FP8, BF16, FP16, TF32, FP32, and FP64) and sparsity matrix data. XCDs may, in some cases, be referred to as graphics computational diodes (GCDs).The Accelerated Processing Unit 1600 can include one or more Complex Computing Units (CCDs) 1606 to perform operations described elsewhere herein, such as, but not limited to, those operations performed by host processors. CCDs may in some cases be referred to as core complexes or CCXs, such as, but not limited to, CCXs used in AMD Ryzen processors. XCDs and CCDs can share any type of cache or memory (for example, one or more memory units 1602) or have a cache or memory allocated to each XCD or CCD or group of XCDs or CCDs. For example, the on-package AMD Infinity Fabric connects XCDs and CCDs to the shared AMD Infinity Cache 1608 and, in some embodiments, to high-bandwidth memory (for example, HMB3).The Accelerated Processing Unit 1600 can include an AMD MI300a processor, which contains three CPU chiplets (or CCDs) and six accelerator chiplets (XCDs) on four input / output dies (IODs). These can be stacked on a single piece of silicon that connects them (for example, via AMD Infinity Fabric), with eight stacks of high-bandwidth DRAM surrounding a superchip. An AMD MI300x processor replaces the CCDs with two additional XCDs for a pure accelerator system.
[0189] The Accelerated Processing Unit 1600 can include one or more input / output (I / O) interfaces. For example, XCDs 1604 and CCDs 1606 can be arranged together on one or more input / output dies (IODs) 1610, which can include one or more I / O interfaces. IODs 1610 can include any number and type of I / O interfaces (for example, PCI, PCI-Extended (“PCI-X”), PCIe, Gigabit Ethernet (“GBE”), USB, etc.). Various types of peripheral devices can be connected to I / O interfaces 1670. I / O interfaces of IODs 1610 can also be used to connect one or more Accelerated Processing Units 1600, for example, in a server architecture.
[0190] The Accelerated Processing Unit 1600 can include one or more Memory Units 1602 for storing instructions and other information used to perform operations described elsewhere herein. The Memory Units 1602 can include any volatile memory, such as, but not limited to, the types of memory described elsewhere herein, and can include, for example, high-bandwidth memory (such as HMB3) or high-bandwidth DRAM. The memory (such as Memory Units 1602) associated with the Accelerated Processing Unit 1600 can include system memory, which can be used, for example, for instructions, statements, and constants, as well as inputs and outputs.The memory units 1602 can also include device memory that can be used as storage for, for example, instructions, directives, and constants, as well as inputs and outputs, as a return buffer, and for private data. Memory units 1602 can be connected to one or more IODs 1610. In at least one embodiment, the L1 cache 1620 initiates a memory hierarchy that includes a shared L2 cache 1628, for example, within XCDs. AMD Infinity Cache™ is a last-level cache (LLC) located on an active I / O die (IOD). CCDs 1606 and XCDs 1604 can have separate or shared memory.The AMD Infinity architecture and AMD Infinity Fabric™ technology enable a coherent, high-throughput unification of GPU and CPU chiplet technologies (e.g., XCDs, CCDs and / or CCXs) with memory (e.g., stacked HBM3 memory) in single devices and across multiple device platforms.
[0191] As in Fig. As shown in Figure 16, an XCD 1604 can include a common set of global resources 1630, which may include a hardware scheduler 1632 and an asynchronous compute engine (ACE) 1624 that sends tasks (for example, compute shader workgroups) to compute units (CUs or cores) 1634. ACEs 1624 (for example, four) can each be connected to CUs 1634 (for example, 40 CUs), and some of the CUs 1634 can be disabled for yield management. CUs 1634 can have a dedicated cache or a shared cache (for example, an L2 cache) 1628 that can be used to aggregate all memory traffic for a chip.CUs 1634 can include thread and parallel processor cores, including instruction retrieval and scheduling with Scheduler (S) 1612, Matrix Core Unit (MCU) 1616, and Shader Core (SC) 1618 (for example, execution units for scalar, vector, and matrix data types), as well as load / store pipelines with an L1 cache 1620 and Local Data Share (LDS) 1614. Local data sharing can include, for example, a scratch RAM with integrated arithmetic functions that enable data exchange between threads in a workgroup. An instruction cache 1640 (for example, for storing and providing instructions for performing operations described elsewhere herein) and a constant cache 1638 can be associated with one or more CUs and shared by two CUs. Matrix kernels 1616 can process a wide variety of data types, such as, but not limited to, INT8, FP8, FP16, BF16 and TF32 data types.The accelerated processing unit 1600 can include compute units 1634, which can be arranged in an array format, such as a data parallel processor (DPP) array. The ultra-thread dispatch processor 1642 can communicate with the compute units 1634, and the instruction processor 1644 can read instructions written by a host to memory-mapped registers in a system memory address space (not shown). The instruction processor 1644 can send hardware-generated interrupts to a host processor (such as a CCD) when an instruction completes. The memory controller 1636 can also have direct access to all of the device's memory and the host-specified portions of system memory.To satisfy read and write requests, the memory controller 1636 can execute functions of a direct memory access controller (DMA controller), which include calculating memory address offsets based on the format of the requested data in memory. For example, one or more of the APIs described here can be compiled into instructions that are stored in the instruction cache 1640 and then retrieved by the instruction fetch logic in the processor 1640, decoded by a processor decoder or similar, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or similar, executed by an execution logic or similar, reordered, and then withdrawn by a retraction logic or similar. API(s) (and / or compiled instructions comprising API(s)) can be stored in any memory outside or inside the processor 1600 (e.g., in the cache and / or memory).A result of API(s) can then be stored in a memory inside or outside the processor 1600, which includes registers, DRAM, Flash, SRAM, cache or other memory equivalents.
[0192] An application can consist of a program that runs on a host processor (for example, a CCD), as well as programs, called kernels, that run on one or more XCDs. Programs can be controlled by host instructions that set internal base addresses and other configuration registers, specify a data range in which the Accelerated Processing Unit 1600 can operate, invalidate and flush caches on the Accelerated Processing Unit 1600, and cause the Accelerated Processing Unit 1600 to begin executing a program. Kernels can be referred to as programs that are executed by the Accelerated Processing Unit 1600.A kernel can be executed independently on each work item or as groups of work items, which can be called a wavefront, and can execute a kernel on all work items in a group (for example, 64) in a single pass. Computing units 1634 can include a scalar arithmetic logic unit (ALU) that can process one value per wavefront (common to all work items), a vector ALU that can process unique values per work item, a local data exchange 1614 that allows work items within a workgroup to communicate and exchange data, a scalar memory (not shown) that can transfer data between scalar general-purpose registers (SGPRs) and memory via a cache, and a vector memory that can transfer data between vector general-purpose registers (VGPRs) and memory, including the sampling of texture maps.Kernel control flow can be controlled using scalar ALU instructions, which can include if / else statements, branches, and loops. Scalar ALU (SALU) and memory instructions can be applied to an entire wavefront and operate on one or more SGPRs. Vector memory and ALU instructions can be applied to all work elements in a wavefront simultaneously.
[0193] In at least one embodiment, the accelerated processing unit 1600 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0194] Fig. Figure 17 shows a 1700 processor, such as, but not limited to, a processor based on a Zen architecture (such as Zen 1, 2, 3, 4, 5, or others) from AMD Corporation in Santa Clara, California, or any other processor that shares at least some of the components described herein. The 1700 processor comprises one or more 1702(1)-1702(N) CPU chips, where N is any integer greater than 1. The 1702 CPU chip can comprise any number of 1716 processor cores (for example, for performing any of the operations described herein) and any number of caches (for example, for storing instructions and other information for performing any of the operations described herein) in any combination. For example, 1718 L2 cache units can be coupled to 1716 processor cores, which can share 1718 L2 cache units and / or be individually coupled to them.Processor cores 1716 can be individually coupled to the L3 cache 1722 and / or share the L3 cache, which can be a lowest-level cache (LLC) 1722 for accessing data and other information used by the processor cores 1716. One or more processor cores 1716 and one or more L2 cache units 1718 can be contained in a core complex (CCX) 1720, which can include a shared cache (for example, a 32 MB cache, such as the L3 cache 1722). The core complex 1720 can be manufactured on a chip (CCD or CPU chip) 1702. For example, up to 12 core complexes 1720, along with 8 CPU chips 1702, can be configured in a processor to provide up to 96 processor cores 1716 for the processor 1700. For example, a “Zen 4c” core complex 1720 can include up to eight cores 1716 and a shared 16 MB L3 cache 1722.Two of these core complexes 1720 can be combined on a single CPU chip 1702, providing 16 cores and a total of 32 MB of L3 cache 1722 per chip. Up to eight CPU chips 1702 can be combined with an I / O unit 1704 to provide CPUs with up to 128 processor cores 1716. Up to four of the "Zen 4c" chips described above can be combined to provide CPUs with up to 64 processor cores 1716.
[0195] The 1700 processor can incorporate a variety of input / output configurations, which are described in more detail below. The 1704 I / O unit can include one or more 1706 memory controllers that manage memory usage (for example, DDR5 memory) for the 1700 processor. The 1704 I / O unit can include one or more 1712 SATA hard disk controllers for managing memory and one or more 1714 Compute Express Link (CXL™) 1.1+ memory controllers, which provide CPU-to-device and CPU-to-memory connections and can be flexibly assigned to specific functions during server design. The 1704 I / O unit can also include a 1708 PCIe controller for connecting peripherals and other components connected to the 1700 processor. The I / O unit 1704 can include USB ports 1710 for connecting to other components that are separate from the processor 1700.CPU chips 1702 can support any number of connections, for example, one or two connections, to the I / O unit 1704. As shown, the I / O unit 1704 can include components, which are described further herein, and the I / O unit 1704 can be an I / O chip that houses several different components. The memory controller 1706, the PCIe controller 1708, the USB ports 1710, the SATA controller 1712, and / or the CXL controller 1714 can be integrated at any location within the processor 1700, either separately or in any groups or combinations thereof.
[0196] The 1700 processor can include Infinity Fabric 1724 interconnects (which resemble or are based on PCIe architectures) that provide connections between CPUs (for example, 1702(1)-1702(N) CPU chips), 1726 GPUs, 1732 inference engine(s), and other components in a multi-chip architecture, such as 1728 secure processors and a 1704 I / O unit. One or more AMD Infinity Fabric™ interconnects 1710 can be connected to 1702(1)-1702(N) CPU dies and serve as a connection between CPUs. One or more Infinity Fabric 1710 interconnects can connect each 1702 CPU chip to the 1710 I / O unit.
[0197] In at least one embodiment, the Processor 1700 may comprise central processing units (CPUs) and other associated hardware and software, as described above and further herein. The Processor 1700 may also comprise Graphics Processors 1726. The Graphics Processor 1726 may be used for image generation and processing, as well as for other computations and operations, as further described herein. The Graphics Processor 1726 may be based on AMD's RDNA 3 or 3.5 architecture in Santa Clara, California. The Graphics Processor 1726 may comprise Graphics Computing Units (GCDs) and Memory Cache Units (MCDs). GCDs may comprise any number of compute units (CUs) for graphics or other processing operations, such as operations performed by Arithmetic Logic Units (ALUs), as further described herein.The 1726 graphics processor can include an L2 cache that can be used by compute units. MCDs (not shown) can include any number of memory units and can include cache, such as L3 cache, as well as memory interfaces for coupling to memory, such as Memory 1742(1)-(N), where N is an integer. Components within the 1726 graphics processor can be interconnected using various approaches, such as Infinity Fabric 1724 connections outside or inside the 1726 graphics processor.
[0198] The inference engine 1732 can provide neural processing functions to the processor 1700 for computational processes used for neural networks, deep learning, and other artificial intelligence-related operations, which are further described herein. The processor 1700 can include one or more secure processors 1728 for managing the security of the processor 1700, a display controller 1730 for controlling displays, a system management unit 1734 for managing and operating some or all of the components on the processor 1700, multimedia engines 1736 for audio and video operations, a fusion controller hub 1738 for managing USB, SATA, and PCIe connections to the processor 1700, and a sensor fusion hub 1740 for managing sensors, such as accelerometers. The processor 1700 can also include a memory 1742(1)-(N), where N is any integer.The memory may include different memory types, such as LPDDR5 and / or DDR5 or others described elsewhere herein.
[0199] To perform the operations described herein, the 1700 processor can include an execution pipeline comprising a pipeline front end, which may include a cache (for example, an L1 cache) for storing instructions (not shown). The instruction flow can be modified by a branch predictor. Instructions can be decoded by a decoder, forwarded to a back end for execution, and renamed. Instruction fetch and decoding pipes can, for example, be forwarded to integer or floating-point execution operations, which can be scheduled by a scheduler and transferred to vector and / or general-purpose registers. Floating-point multiplication and / or addition operations can be processed, and arithmetic logic units (ALUs) can also be used to perform calculations, such as arithmetic and logical operations.The outputs of computing units can be coupled to a load / storage unit, which may be connected to a cache, such as an L1 cache and / or an L2 cache.
[0200] With respect to the Processor 1700 and each of its components described above or elsewhere herein, one or more of the APIs or equivalents described herein may, for example, be compiled into instructions or equivalents (for example, AVX-512 instructions based on a SIMD model), which are fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (for example, sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then withdrawn by an execution logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or inside the Processor 1700 (for example, in the cache and / or memory).A result of the API(s) can then be stored in a memory inside or outside the 1700 processor, which includes registers, DRAM, Flash, SRAM, cache or other memory equivalents.
[0201] In at least one embodiment, the 1700 processor may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0202] Fig. Figure 18 shows an example of a 1800 processing core that can implement the Arm architecture (for example, v9.0-A) or another processor that shares at least some of the components described here. The Neoverse™ V2 1800 core can be implemented within a DynamlQ Shared Unit (DSU) cluster via a DSU-110 interconnect 1854 for one or more interconnected cores, for example, for parallel processing. The Neoverse™ V2 core can be implemented as a single core in a DSU cluster configured for a direct connection, with or without L3 cache, snoop filtering, or snoop control unit (SCU) logic (not shown). The Neoverse™ V2 core can include a CPU Bridge 1852, which connects the Core 1800 to the DSU-110 interconnect, which can also connect the Core 1800 to an external storage system and the rest of a system-on-a-chip.The L1 instruction storage system 1802 can retrieve instructions from an instruction cache 1804 and deliver instructions (for example, one or more APIs described herein that can be compiled into instructions) to an instruction decoding unit 1810 to perform, for example, some or all of the operations described above or elsewhere herein. The L1 instruction storage system 1802 can include an L1 instruction cache 1804, for example with 64-byte cache lines; an L1 instruction translation lookaside buffer (TLB) 1806, for example with native support for 4 KB, 16 KB, 64 KB, and 2 MB page sizes; and a macro operation cache (MOP) 1808 (for example, an L0 MOP cache with 1536 entries and 4x asynchronous associativity), which can contain decoded and optimized instructions for higher performance. The instruction decoding unit 1810 can decode AArch64 instructions into an internal format.The Register Renaming Unit 1812 can perform register renaming to facilitate out-of-order execution and route decoded instructions to various output queues. The Instruction Output Unit 1814 can control when decoded instructions are allowed to be routed to execution pipelines and can include output queues for storing instructions waiting to be routed to execution pipelines. The Integer Execution Pipeline 1816 can be contained within an execution pipeline and may include an Integer Execution Unit 1818, which can perform arithmetic and logical data processing operations. The vector execution unit 1820 can be included in an execution pipeline and can perform extended SIMD and floating-point operations (FPU) 1822, SVE (Scalable Vector Extension) and SVE2 (Scalable Vector Extension 2) instructions 1824, and optionally cryptographic (Crypto) instructions 1826.Advanced SIMD can include a media and signal processing architecture that adds instructions primarily for audio, video, 3D graphics, image, and speech processing. A floating-point architecture provides support for single- and double-precision floating-point operations. The L1 Data Storage System 1830 can execute load and store instructions and handle memory coherence requirements. The L1 Data Storage System 1830 can include an L1 Data Cache 1832 and a fully associative L1 Data TLB 1834 with native support for page sizes of 4 KB, 16 KB, and 64 KB, and block sizes of 2 MB and 512 MB. The Memory Management Unit (MMU) 1828 can provide fine-grained memory system control through a series of mappings between virtual and physical addresses and memory attributes, which can be stored in translation tables that can be stored in the TLB 1834 when translating an address.The L2 memory system 1836 can include an L2 cache 1838 and be connected to the DSU-110 1854 via an asynchronous CPU bridge 1852. The Neoverse™ V2 core 1800 can support a range of debugging, testing, and tracing options, including a trace unit 1842, a trace buffer 1840, and an embedded logic analyzer (ELA) 1848. The Neoverse™ V2 core 1800 can implement the Statistical Profiling Extension (SPE) 1844 to provide a statistical view of the performance characteristics of executed instructions, which software developers can use to optimize their code for improved performance. The Performance Monitoring Unit (PMU) 1846 can provide performance monitors that can be configured to collect statistics on the operation of each core and memory system. The information can be used for debugging and code profiling.The CPU interface 1850 of the Generic Interrupt Controller (GIC), when integrated with an external distributor component, can serve as a resource for supporting and managing interrupts in a cluster system. In a cluster, there can be a CPU bridge 1852 between each Neoverse™ V2 core 1800 and DSU-110 1854. The CPU bridge 1852 can control buffering and synchronization between the core 1800 and the DSU-110 1854. The CPU bridge 1852 can be asynchronous to allow different frequency, power, and area implementation points for each core 1800. The CPU bridge 1852 can run synchronously without affecting other interfaces, such as, but not limited to, debugging and tracing, which can be asynchronous.
[0203] In at least one embodiment, the Core 1800 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0204] Fig. Figure 19 shows one or more chips comprising one or more tensor processing units (TPUs) 1900 according to at least one embodiment. The TPUs 1900 in Fig. 19 TPUs can include application-specific integrated circuits (ASICs), for example, to perform some or all of the operations described above or elsewhere herein, such as, but not limited to, accelerating machine learning workloads that perform matrix operations. TPUs 1900 can include ASICs from Alphabet Corporation in Mountain View, California. Cloud TPU includes a cloud service that provides TPUs as a scalable resource for processing tasks such as, but not limited to, machine learning workloads that can run on frameworks such as, but not limited to, TensorFlow, PyTorch, and JAX.
[0205] The Chip 1900 can comprise any number of TPUs, which can include Tensor Cores 1906. The Tensor Core 1906 can include one or more Core Sequencers 1908, a Vector Processing Unit (VPU) 1910, a Matrix Multiplication Unit (MXU) 1912(A)-1914(N), where N is any integer greater than 1, and a Transpose Permutation Unit 1916. The Core Sequencer 1908 can retrieve instructions (for example, VLIW (Very Long Instruction Word)) from the Core 1906's instruction memory (Imem), perform scalar operations using a scalar data memory (Smem) and scalar registers (Sregs) (not shown), and forward vector instructions to the Vector Processing Unit (VPU) (1910). For example, commands can trigger eight operations: two scalar operations, two vector ALU operations, vector load / store units, and a pair of slots that pass data to and receive data from matrix multiplication and transposition units.The VPU 1910 can perform vector operations using a large on-chip vector memory (Vmem) and vector registers (Vregs). The VPU 1910 can stream data to and from the MXU via decoupled FIFOs. The VPU 1910 can collect and distribute data to Vmem via data parallelism (2D matrix and vector functional units) and instruction parallelism (8 operations per instruction). A large two-dimensional matrix multiplication unit (MXU) 1912(A)-1912(N), for example, can use a systolic array to reduce area and power consumption, as well as large, software-controlled on-chip memories instead of caches. The transpose-reduction-permutation unit 1916 can perform matrix transpositions, reductions, and permutations of VPU 1910 lanes (e.g., 128x128). The high-bandwidth 1904 memory can be used for on-chip applications and can be coupled with 1902 host queues via PCIe, for example.One or more Chip 1900s can be interconnected for computation. For example, one or more Chip 1900s can be connected as a torus, such as a 2D torus. The Chip 1900 can also include any number (for example, four) of Inter-Core Interconnections (ICI) 1918, which specify that direct connections between chips are permitted to form a supercomputer.
[0206] With respect to all processors in Chip 1900 and all its components described above or elsewhere herein, one or more of the APIs or equivalents described herein may, for example, be compiled into instructions or equivalents that are fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then swapped out by a swap logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or within any processors in Chip 1900 (e.g., in the cache and / or memory).A result of the API(s) can then be stored in a memory inside or outside of any processor in the Chip 1900, which includes registers, DRAM, Flash, SRAM, cache or other memory equivalents.
[0207] In at least one embodiment, the Chip 1900 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0208] Fig. Figure 20 shows a vector processor according to at least one embodiment. The Vector Processor 2000 can support a RISC-V standard. The Vector Processor 2000 can include one or more Cores 2010 (for example, scalar units) with one or more Vector Processing Units (VPUs) 2042 (for example, vector units) that can perform, for example, some or all of the operations described above or elsewhere herein. The Core 2010 can include an Andes Custom Extension (ACE) 2016, which can be used, for example, via ACP 2038 for communicating user-defined instructions to the Processor 2000. The Core 2010 can include a 1-cycle multiplier and a 1-cycle instruction / data local memory (ILM / DLM) to increase parallelism by simultaneously fetching instructions and accessing data.The Memory Management Unit (MMU) 2024 can manage system memory and the cache, and is responsible for branch execution, instruction pair output, L1 instruction / data caches, and local memory. The Core 2010 can include a Physical Memory Protection Unit and a Programmable Physical Memory Attribute Unit (PMP / PPMA) 2022. The Core 2010 can include a Digital Signal Processor (DSP) 2028 and a Floating Point Unit (FPU) 2026, as well as a Load / Store Unit (LSU) 2032 for interface to the memory hierarchy (D$ 2034 and I$ 2030). The Core 2010 can include a Branch Prediction Unit 2018 and a Multiplier Unit 2020.
[0209] The Vector Processing Unit (VPU) 2042 can include one or more Vector Functional Units (FUs) 2046 (A)-2046(N) which can be chained together for parallel processing, independent memory paths for RISC-V Vector (RVV) load / store operations via ACE-RVV 2048 and Andes Streaming Port (ASP) 2044 load / store operations, and a Vector Load / Store Unit (VLSU) 2050.
[0210] The Vector Processor 2000 can include bus interfaces such as, but not limited to, an L2 cache memory port 2056 for cacheable access, an MMIO port 2054 for non-cacheable access, an Input-Output Coherence Port (IOCP) 2058 for a cache-free bus master, local memory access ports for ILM / DLM 2012 which can be coupled with SRAM 2006, and a High Bandwidth Vector Memory Access (HVM) 2036, a Common Peripheral Port (SPP) 2052 for external peripherals. Other memory ports include the LM slave port AXI 2002, the HVM subport AXI 2004, MEM (AXI) 2062, and AXI 2060. Trace I / F 2014, via Inst. Trace I / F 2008, can, for example, capture, encode, and transmit off-chip a recording of executed processor instructions, which software tools can use to reconstruct the exact execution sequence of a program.
[0211] With respect to all processors in Processor 2000 and all its components described above or elsewhere herein, one or more of the APIs or equivalents described herein may, for example, be compiled into instructions or equivalents that are fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler or equivalent, executed by an execution logic or equivalent, reordered, and then withdrawn by a retract logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or inside Processor 2000 (e.g., in the cache and / or memory).A result of the API(s) can then be stored in a memory inside or outside the Processor 2000, which includes registers, DRAM, Flash, SRAM, cache or other memory equivalents.
[0212] In at least one embodiment, the Vector Processor 2000 may comprise one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0213] Fig. Figure 21A shows a diagram of an exemplary microarchitecture of a tiled multi-core processor. The multi-core tiled processor in Fig. 21A may include a speech processing processor. As in Fig. As shown in Figure 21A, each "tile" of a processor architecture is a processing element interconnected via a network-on-chip (NoC) that can be used, for example, to perform some or all of the operations described above or elsewhere herein. For example, each tile may include an instruction output unit 2104, an integer (INT) 2106 and floating-point (FP) unit 2108, a load / store unit (LSU) 2112 for interface with the memory hierarchy (data cache (D$) 2110 and instruction cache (I$) 2114), and a network (NET) interface 2116 for communication with other tiles. Some tiles in the 2100 processor may include a memory controller 2102 for managing and controlling memory, as further described herein. The 2100 processor may have a functional slice architecture. The 2100 processor can be located on an application-specific integrated circuit (ASIC), and Fig. 21A can represent a layout of an ASIC. The 2100 processor can include a coprocessor designed to execute instructions for a predictive model. A predictive model is any model configured to make a prediction from input data. A predictive model can use a classifier to make a classification prediction. A predictive model can be a machine learning model, such as, but not limited to, a tensor flow model, and the 2100 processor is a tensor streaming processor.
[0214] The 2100 processor can use different microarchitectures, which are represented in each tile in Fig. The functional units shown in Figure 21B are divided. Instead, the functional styles 2124 of the processor 2100 can be grouped into a variety of functional processing units (hereinafter referred to as "slices") 2104, each corresponding to a specific functional type (for example, FP / INT 2118, NET 2120, MEM 2122). As shown in Fig. As shown in Figure 21B, for example, each slice can correspond to a column of functional tiles extending in a north-south direction. Furthermore, the 2100 processor can also include communication lanes to transfer data between tiles of different slices, each running horizontally in an east-west direction. Each communication lane can be connected to any of the 2104 slices of the 2100 processor.
[0215] The 2104 disks of the 2100 processor can each correspond to a different function and include arithmetic logic disks (for example, FP / INT2118), line-switching disks (for example, NET 2120), and memory disks (for example, MEM 2122). Arithmetic logic units can perform one or more arithmetic and / or logical operations on data received over communication channels to produce output data. Examples of arithmetic logic units include matrix multiplication units and vector multiplication units. Memory segments comprise memory cells that store data. Memory segments can provide data to other segments over communication channels. Memory segments can also receive data from other segments over communication channels. Lane-switching slices can configurably forward data from one communication lane to any other communication lane.For example, data from a first lane can be provided to a second lane via a lane-switching slice. In some embodiments, a lane-switching slice can be implemented as a crossbar switch. Each 2104 slice also includes its own instruction queue (not shown) that stores instructions and an instruction control unit (ICU) for controlling instruction execution. Instructions in a given instruction queue can only be executed by tiles in their associated function slice and cannot be executed by other slices of the 2100 processor.
[0216] By arranging the tiles of the 2100 processor in different functional slices 2104, the on-chip instruction and control processes of the 2100 processor can be decoupled from the data flow. For example, an arrow in Fig. 21B the flow of instructions within the processor architecture according to some embodiments. Another arrow in Fig. Figure 21B shows the data flow within the processor architecture according to at least one embodiment. As shown, instructions and control flow can flow in a first direction across the tiles of the 2100 processor (for example, from north to south along the length of the function segments, as shown by the first arrow), while data flows in a second direction across the tiles of the 2100 processor (for example, from east to west across the function segments, as shown by the second arrow), which is perpendicular to the first direction.
[0217] Different functional slices of the Processor 2100 can correspond to MEM 2122 (memory), VXM (vector execution module), MXM (matrix execution module), NIM (numeric interpretation module), and SXM (switching and permutation module). Each slice can comprise N tiles, all of which can be controlled by an identical instruction control unit (ICU) (not shown). Each slice can operate completely independently and can only be coordinated using barrier-like synchronization primitives or by a compiler exploiting tractable determinism. Each tile of the Processor 2100 can correspond to an execution unit organized as an xM SIMD tile. For example, each tile of the Processor 2100's on-chip memory can be organized to atomically store an L-element vector. Thus, MEM slices with N tiles can work together to store or process a large vector (for example, with a total of N×M elements).
[0218] Tiles within a slice can execute instructions in a staggered manner, with instructions within a slice being issued tile by tile over a period of N cycles. Functional slices can be physically arranged on the chip to enable efficient data flow for pipeline execution over hundreds of cycles for common patterns. Data flows can perform a single "reversal" (change of direction) corresponding to a single matrix operation before being written back to memory. In some embodiments, a given data flow can change direction multiple times (due to multiple matrix and vector operations) before the resulting data is written back to memory.
[0219] When using a Processor 2100 (for example, TSP) with a functional slice architecture that includes a functional slice, a TSP compiler (not shown) generates an explicit plan for how the Processor 2100 can execute a program (for example, a microprogram). The compiler can determine when each operation is performed, which functional slices perform the work, and which STREAM registers contain operands. The compiler can maintain a highly accurate (cycle-accurate) model of the Processor 2100's (for example, TSP's) hardware state, enabling a microprogram to coordinate data flow.
[0220] The Processor 2100 (e.g., TSP) can use a web-based compiler that takes a model as input (e.g., a machine learning model such as, but not limited to, a TensorFlow model) and outputs a proprietary instruction targeted to the Processor 2100 (e.g., TSP). The compiler is responsible for coordinating the control and data flow of a program and establishes any instruction-level parallelism by explicitly grouping instructions that can and should be executed concurrently so they can be sent together. The primary hardware structure includes an architecturally visible streaming register file (STREAMs), described in more detail below, which serves as a conduit for operands to flow from MEM slices (e.g., SRAM) to function slices and vice versa.
[0221] The MEM 2122 of the Processor 2100 can serve as: (1) memory for model parameters, microprograms, and the data on which they operate, and (2) network-on-chip (NoC) for communication of data operands from the MEM to functional slices and of computed results back to the MEM. In some embodiments, the on-chip memory can occupy approximately 75% of the Processor 2100's chip area. In some embodiments, the on-chip memory of MEM tiles may consist of SRAM rather than DRAM due to the bandwidth requirements of the Processor 2100. The on-chip memory capacity of the Processor 2100 can determine (i) the number of machine learning models that can be stored on the chip simultaneously, (ii) the size of a given model, and (iii) the partitioning of large models to accommodate multi-chip systems.In some embodiments, the MEM system of the 2100 processor can provide a multitude of memory disks organized into two distinct hemispheres (referred to as "MEM WEST" and "MEM EAST," respectively).
[0222] The memory segments of each hemisphere can be mirrored, so that the segments can be physically numbered {0, ... L} in the eastern hemisphere and {L, ... 0} in the western hemisphere, such that memory segment 0 for each hemisphere corresponds to a segment closest to the VXM segments between the hemispheres, with each hemisphere containing L segments. The direction of data transmission toward the center of a chip can be described as "inward," while data transmission toward the outer (easternmost or westernmost) edge of a chip can be described as "outward." Although the memory hemispheres of the 2100 processor can be referred to as East and West, it is understood that in other embodiments, other designations may be used to denote different memory hemispheres.
[0223] In some embodiments, a streaming register file called STREAMS transfers operands and results between the SRAM of MEM slices and the function slices of the 2100 processor. In some embodiments, a plurality of MEM slices (for example, between 2 and 10 adjacent MEM slices) can be physically organized as a set. Each set of slices can be located between a pair of STREAM register files, so that each slice can read from or write to STREAM registers in both directions. Positioning STREAM register files between sets of MEM slices reduces the number of cycles required to transfer data operands across one hemisphere (for example, by a factor corresponding to the number of slices per set). The number of slices per set can be configured based on a distance over which data can be transferred in a single clock cycle.
[0224] Regarding all processors in Fig. 21 and all components described above or elsewhere herein can, for example, compile one or more of the APIs or equivalents described herein into instructions or equivalents, which are fetched by an instruction fetch logic or equivalent, decoded by a processor decoder or equivalent, and scheduled for execution by a scheduler or equivalent (e.g., sequentially or out of sequence), executed by an execution logic or equivalent, reordered, and then withdrawn by a retract logic or equivalent. API(s) (and / or compiled instructions comprising API(s)) can be stored in any memory outside or inside the processor 2100 (e.g., in the cache and / or memory).A result of the API(s) can then be stored in memory inside or outside the 2100 processor, which may include registers, DRAM, Flash, SRAM, cache or other memory equivalents.
[0225] In at least one embodiment, the 2100 processor may include one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. Software designs
[0226] The following figures show, without limitation, examples of software designs for implementing at least one embodiment.
[0227] Fig. Figure 22 shows a software stack of a programming platform according to at least one embodiment. A programming platform may comprise a platform for utilizing hardware on a computer system to accelerate computational tasks. A programming platform may be accessible to software developers via libraries, compiler directives, and / or extensions to programming languages in at least one embodiment. A programming platform may be CUDA, Radeon Open Compute Platform (“ROCm”), OpenCL (OpenCL™ was developed by the Khronos Group), SYCL, or Intel oneAPL.
[0228] A Software Stack 2200 of a programming platform can provide an execution environment for an Application 2201. The Application 2201 can include any computer software that can be launched on the Software Stack 2200. The Application 2201 can include an artificial intelligence (AI) / machine learning (ML) application, a high-performance computing (HPC) application, a virtual desktop infrastructure (VDI), or a data center workload.
[0229] The application 2201 and the software stack 2200 run on the hardware 2208. The hardware 2208 can include one or more GPUs, CPUs, FPGAs, AI engines, and / or other types of computing devices that support a programming platform. The software stack 2200 can be vendor-specific and compatible only with devices from certain manufacturers, such as CUDA, ROCm, oneAPl, OpenCL, or other implementations. The hardware 2208 can include a host connected to one or more devices that can be accessed via an application programming interface (API) to perform computing tasks. A device within the hardware 2208 can include a GPU, an FPGA, an AI engine, or another computing device (but also a CPU) and its memory, in contrast to a host within the hardware 2208, which can include a CPU (but also a computing device) and its memory, at least in one embodiment.With respect to each of the 2208 hardware described above or elsewhere herein, one or more of the APIs described herein may, for example, be compiled into instructions that may be fetched by an instruction fetch logic, decoded by a processor decoder, scheduled for execution (e.g., sequentially or out of sequence) by a scheduler, executed by an execution logic, reordered, and then withdrawn by a retract logic. API(s) (and / or compiled instructions comprising API(s)) may be stored in any memory outside or inside the 2208 hardware (e.g., in the cache and / or memory). A result of the API(s) may then be stored in memory inside or outside the 2208 hardware, including registers, DRAM, flash, SRAM, cache, or other memory. One or more of the APIs described herein may receive a call.One or more of the APIs described here can communicate with a library or a section of a library to execute a function described by the call. One or more of the APIs described here can receive a call and communicate with a library or a section of a library to execute a function described by the call.
[0230] The software stack 2200 of a programming platform can include a set of libraries 2203, a runtime environment 2205, an optional driver / interface 2207, and a device kernel driver 2208. Each of the libraries 2203 can contain data and program code that can be used by computer programs and utilized in software development. Libraries 2203 can include pre-written code and subroutines, classes, values, type specifications, configuration data, documentation, help data, and / or message templates. Libraries 2203 can include functions optimized for execution on one or more devices. The libraries 2203 can include functions for performing mathematical, deep learning, and / or other types of operations on devices.The libraries 2203 can be linked to corresponding APIs 2202, which may include one or more APIs that expose functions implemented in the libraries 2203. A processor (e.g., CPU, GPU) can execute, call, or otherwise use one or more APIs to prioritize kernels. For example, a first kernel (e.g., a parent kernel) can start a second kernel (e.g., a child kernel), and this second kernel can be used by a processor to start additional kernels (e.g., child kernels) independently of the first kernel.A processor can execute an API or call an API from memory to support dynamic stream priority (for example, updating the priority while a stream is being used to perform operations). For example, when a processor executes the API, a programmer can copy the stream priority from one stream to one or more other streams.
[0231] Software Stack 2200 can include an API to support dynamic stream priority (for example, updating the priority while a stream is being used to perform operations), allowing a programmer to set a stream's priority at any time after its creation. Software Stack 2200 can also include an API to support dynamic stream priority (for example, updating the priority while the stream is being used to perform operations), which allows a programmer to determine a stream's current priority, where priority is one of a variety of attributes of a stream.The Software Stack 2200 can include an API to support dynamic stream priority (for example, updating the priority while the stream is being used to perform operations), allowing a programmer to retrieve the current priority of a stream as a single attribute. The Software Stack 2200 can also include an API to support dynamic stream priority (for example, updating the priority while the stream is being used to perform operations), allowing a programmer to start a kernel to perform operations on a stream with a specified priority that may differ from the stream priority.The Software Stack 2200 can include an API to specify whether an object (for example, a thread synchronization object such as, but not limited to, a barrier) is tracking whether all data movement operations for a group of threads running on a GPU have completed after a certain period of time and are in a certain state, where a certain state can be a state indicating that data has been moved and is ready for use, and is set using an expected parity value as input for the API.
[0232] The Software Stack 2200 can include one or more APIs for updating kernels. A processor can execute an API or call an API from memory to update an existing API to support context-free kernels, allowing a programmer to add a kernel node to a graph without a graphics context, thus enabling a graphics context to be dynamically associated with a kernel at runtime. The Software Stack 2200 can also include one or more APIs that allow a programmer to obtain a kernel identifier and a graphics context as separate parameters from a kernel node, enabling parameters to be obtained from both kernels and context-free kernels.The Software Stack 2200 can include one or more APIs to use parallel processors, such as one or more graphics processing units, to start task graphs (e.g., task graphs) and to execute one or more task graphs (e.g., including one or more programs).
[0233] The Software Stack 2200 can include one or more APIs to associate one or more commands with one or more memory order operations, such as, but not limited to, a fence or membar operation. Commands can be associated with one or more domains, so that a memory order operation is performed in conjunction with one or more specific domains without affecting commands in other domains. An API can indicate that a thread has arrived (for example, at a thread synchronization barrier) or has completed a work phase with respect to asynchronous data movement operations on a GPU. The Software Stack 2200 can include one or more APIs that allow programmers to manually specify an expected transaction count when a thread has completed a work phase.This can be used to update an object that tracks whether all data movement operations for a group of threads have been completed.
[0234] Application 2201 can be written as source code, which is compiled into executable code, as described below in conjunction with the Fig. 23 and Fig. 24 is explained in more detail. The executable code of application 2201 can be executed, at least partially, in an environment provided by software stack 2200. During the execution of application 2201, code may be encountered that must be executed on a device rather than a host. In such a case, runtime 2205 can be invoked to load and start the necessary code on the device. Runtime 2205 can encompass any technically feasible runtime system capable of supporting the execution of application 2201.
[0235] The 2205 runtime environment can be implemented as one or more runtime libraries connected to corresponding APIs, represented as API(s) 2204. One or more of these runtime libraries may include, among other things, memory management, execution control, device management, error handling, and / or synchronization functions. Memory management functions may include functions for allocating, releasing, and copying device memory, as well as for transferring data between host memory and device memory. Execution control functions may include functions for starting a function (sometimes called a "kernel" if it is a global function that can be called by a host) on a device and for setting attribute values in a buffer managed by a runtime library for a particular function to be executed on a device.
[0236] Runtime libraries and their corresponding APIs can be implemented in any technically feasible way. One (or any number of) APIs can provide a set of low-level functions for fine-grained control of a device, while another (or any number of) APIs can provide a set of such functions at a higher level. A high-level language runtime API can build upon a low-level API. One or more language runtime APIs can be language-specific APIs built upon a language-independent language runtime API.
[0237] An optional driver or interface 2207 can be implemented, for example, for CUDA and ROCm implementations, which are described in more detail below. The optional driver / interface 2207 can be connected to optional driver or interface APIs, such as, but not limited to, CUDA and / or ROCm APIs.
[0238] One or more processors disclosed in "Processing Systems" can execute, access, or otherwise use Software Stack 2200. For example, System-on-a-Chip 900, Parallel Processor 1000, Graphics Multiprocessor 1034, Processor 1100, Processor 1200, Accelerator 1300, Neuromorphic Processor 1405, Supercomputer 1500, Acceleration Processing Unit 1600, Processor 1700, Processor 1800, Tensor Processing Unit 1900, Processor 2000, and Speech Processing Unit 2100 can execute, use, call, or otherwise implement (for example, by accessing memory) one or more APIs contained in Software Stack 2200.
[0239] The Device Kernel Driver 2208 can be configured to facilitate communication with an underlying device. The Device Kernel Driver 2208 can provide low-level functions upon which APIs, such as but not limited to API(s) 2204, and / or other software are based. The Device Kernel Driver 2208 can be configured to compile Intermediate Representation (“IR”) code into binary code at runtime. For CUDA or other implementations such as but not limited to ROCm, oneAPI, or OpenCL, the Device Kernel Driver 2208 can compile Parallel Thread Execution (“PTX”) IR. It can also compile non-hardware-specific code at runtime into binary code for a specific target device (with caching of the compiled binary code), a process sometimes referred to as “finalizing” the code.This allows the finalized code to be executed on a target device that may not have existed when the source code was originally compiled into PTX code. Alternatively, the device code can be compiled offline into binary code without the need for the device kernel driver 2208 to compile the IR code at runtime.
[0240] Processors described elsewhere herein, such as, but not limited to, processors in the Fig. 9-21, may include one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software, for example, the Software Stack 2200, to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0241] According to at least one embodiment, the software stack 2200 can consist of Fig. 22 in a CUDA implementation. A CUDA software stack 2200, on which an application 2201 can be started, can include CUDA libraries 2203, a CUDA runtime 2205, a CUDA driver 2207, and a device kernel driver 2208. The CUDA software stack 2200 can be run on hardware (for example, a graphics multiprocessor 1034, which can include a GPU that supports CUDA and was developed by NVIDIA Corporation in Santa Clara, California).
[0242] The application 2201, the CUDA runtime 2205, and the device kernel driver 2208 can perform functions described above and elsewhere in this document. The CUDA driver 2207 can include a library (libcuda.so) that can implement a CUDA driver API 2206. Similar to a CUDA runtime API 2204 implemented by a CUDA runtime library (cudart), the CUDA driver API 2206 can provide, among other things, functions for memory management, execution control, device management, error handling, synchronization, and / or graphics interoperability. The CUDA driver API 2206 may differ from the CUDA runtime API 2204 in that the CUDA runtime API 2204 simplifies the management of device code by providing implicit initialization, context management (analogous to a process) and module management (analogous to dynamically loaded libraries).In contrast to the high-level CUDA runtime API 2204, the CUDA driver API 2206 can be a low-level API that allows for finer control of a device, particularly regarding contexts and module loading. The CUDA driver API 2206 can provide context management features that the CUDA runtime API 2204 might not. The CUDA driver API 2206 can also be language-independent and, in addition to the CUDA runtime API 2204, may support languages such as OpenCL. Furthermore, development libraries that comprise the CUDA runtime environment 2205 can be considered separate from driver components, including the user-mode CUDA driver 2207 and the kernel-mode device driver 2208 (sometimes referred to as the "display").
[0243] CUDA libraries 2203 can include mathematical libraries, deep learning libraries, parallel algorithm libraries, and / or signal / image / video processing libraries that can be used by parallel computing applications such as, but not limited to, application 2201. CUDA libraries 2203 can include mathematical libraries such as, but not limited to, a cuBLAS library, which is an implementation of Basic Linear Algebra Subprograms (“BLAS”) for performing linear algebra operations; a cuFFT library for computing fast Fourier transforms (“FFTs”); and a cuRAND library for generating random numbers. CUDA libraries 2203 can include deep learning libraries such as, but not limited to, a cuDNN library with primitives for deep neural networks and a TensorRT platform for high-performance deep learning inference.
[0244] In at least one embodiment, processors described elsewhere herein, such as, but not limited to, processors in the Fig. 9-21, comprise one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits can be configured by software, for example, the Software Stack 2200, to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0245] According to at least one embodiment, the software stack 2200 can consist of Fig. 22 in an ROCm implementation. An ROCm software stack 2200, on which an application 2201 can be started, includes a language runtime 2203, a system runtime environment 2205, a Thunk 2207, and an ROCm kernel driver 2208. The ROCm software stack 2200 runs on hardware 2209, which may include a GPU that supports ROCm and was developed by AMD Corporation in Santa Clara, California.
[0246] Application 2201 can perform similar functions to those described above in conjunction with Fig. 22 described. In addition, Srpach runtime 2203 and system runtime environment 2205 can perform similar functions to those described above in conjunction with Fig. The runtime environment 22205 described in section 2203 and the system runtime environment 2205 can differ in that the system runtime environment 2205 is a language-independent runtime environment that implements a ROCr system runtime API 2204 and uses a Heterogeneous System Architecture (“HSA”) runtime API. The HSA runtime API can include a lean, user-mode API that provides interfaces for accessing and interacting with an AMD GPU, including functions for memory management, execution control via the architectural distribution of kernels, error handling, system and agent information, and initializing and shutting down the runtime environment. In contrast to the system runtime 2205, the language runtime API 2203 can be an implementation of a language-specific runtime API 2202 that sits on top of the ROCr system runtime API 2204 as a layer.The language runtime API can include, among other things, a Heterogeneous Compute Interface for Portability (“HIP”) language runtime API, a Heterogeneous Compute Compiler (“HCC”) language runtime API, or an OpenCL API. In particular, the HIP language is an extension of the C++ programming language with functionally similar versions of CUDA mechanisms, and a HIP language runtime API can include functions similar to those of the above in conjunction with... Fig. 22 described CUDA runtime API may be similar to, but not limited to, functions for memory management, execution control, device management, error handling and synchronization.
[0247] Thunk (ROCt) 2207 can be an interface 2206 that can be used to interact with the underlying ROCm driver 2208. The ROCm driver 2208 can be a ROCk driver, which is a combination of an AMDGPU driver and an HSA kernel driver (amdkfd). The AMDGPU driver can be a device kernel driver for GPUs developed by AMD, which provides similar functionality to the one above in conjunction with Fig. The device kernel driver 2209, as described in section 22, is executed. The HSA kernel driver can be a driver that allows different types of processors to share system resources more effectively via hardware functions.
[0248] Various libraries (not shown) may be included in the ROCm software stack 2200 above the Srpach runtime 2203 and provide functions that complement the above in conjunction with Fig. The CUDA libraries described in section 2203 are similar. Various libraries may include mathematical, deep learning, and / or other libraries, such as, but not limited to, a hipBLAS library that implements functions similar to those of CUDA cuBLAS, a rocFFT library for calculating FFTs that is similar to CUDA cuFFT, among others.
[0249] Processors described elsewhere herein, such as, but not limited to, processors in the Fig. 9-21, may include one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits may be configured by software, for example, the Software Stack 2200, to use one or more SNR values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0250] According to at least one embodiment, the software stack 2200 can consist of Fig. 22 in an OpenCL implementation. An OpenCL software stack 2200, on which an application 2201 can be started, may include an OpenCL framework 2203, an OpenCL runtime 2205, and a driver 2208. The OpenCL software stack 2200 can run on hardware 2209 that is not vendor-specific. Because OpenCL is supported by devices from different manufacturers, specific OpenCL drivers may be required to work with the hardware of these manufacturers.
[0251] The application 2201, the OpenCL runtime 2205, the device kernel driver 2208, and the hardware 2209 can perform similar functions to other implementations of the application 2201, the runtime environment 2205, the device kernel driver 2208, and the hardware 2209 mentioned above in conjunction with Fig. 22 were explained. Application 2201 can further include an OpenCL kernel (not shown) with code to be executed on a device.
[0252] OpenCL can define a "platform" that allows a host to control devices connected to it. An OpenCL framework can provide a platform layer API and a runtime API, represented as Platform API 2202 and Runtime API 2204, respectively. Runtime API 2204 can use contexts to manage kernel execution on devices. Each device to be identified can be associated with a corresponding context, which Runtime API 2204 can use to manage command queues, program objects, and kernel objects, as well as to share memory objects for that device, among other things. Platform API 2202 can provide functionality that allows device contexts to be used to select and initialize devices, submit tasks to devices via command queues, and enable data transfer to and from devices.In addition, the OpenCL framework can provide various built-in functions (not shown), including mathematical functions, relational functions, and image functions.
[0253] A compiler (not shown) can also be included in the OpenCL framework 2203. Source code can be compiled offline before an application is executed or online during application execution. In contrast to CUDA and ROCm, OpenCL applications can be compiled online by a compiler that is representative of any number of compilers that can be used to compile source code and / or IR code, such as, but not limited to, Standard Portable Intermediate Representation (“SPIR-V”) code, into binary code.
[0254] Alternatively, OpenCL applications can be compiled offline before running such applications.
[0255] In at least one embodiment, the processors described elsewhere herein, such as, but not limited to, the processors in the Fig. 9-21, comprise one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits can be configured by software, for example, the Software Stack 2200, to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0256] According to at least one embodiment, software can be supported by a programming platform configured to support various programming models, middleware, and / or libraries and frameworks upon which an application can rely. The application can be an AI / ML application implemented, for example, using a deep learning framework such as MXNet, PyTorch, or TensorFlow, which can rely on libraries such as cuDNN, NVIDIA Collective Communications Library (“NCCL”), and / or NVIDIA Developer Data Loading Library (“DALI”) CUDA libraries to enable accelerated computation on the underlying hardware.
[0257] The programming platform can be one of the above in conjunction with Fig. The programming platform can be one of the CUDA, ROCm, or OpenCL platforms described in section 22. It can support multiple programming models, which may be abstractions of an underlying computing system that allow for the expression of algorithms and data structures. Programming models can expose features of the underlying hardware to improve performance. These models can include CUDA, HIP, OpenCL, C++ Accelerated Massive Parallelism (“C++AMP”), Open Multi-Processing (“OpenMP”), Open Accelerators (“OpenACC”), and / or Vulcan Compute.
[0258] Libraries and / or middleware can provide implementations of programming model abstractions. Such libraries can include data and program code that can be used by computer programs and leveraged in software development. Middleware can also include software that provides services to applications beyond those available from the programming platform. Libraries and / or middleware can include cuBLAS, cuFFT, cuRAND, and other CUDA libraries, or rocBLAS, rocFFT, rocRAND, and other ROCm libraries. Furthermore, libraries and / or middleware can include NCCL and ROCm Communication Collectives Library ("RCCL") libraries, which provide communication routines for GPUs, an MIOpen library for deep learning acceleration, and / or an Eigen library for linear algebra, matrix and vector operations, geometric transformations, numerical solvers, and related algorithms.
[0259] Application frameworks can depend on libraries and / or middleware. Any application framework can be a software framework used to implement a standard application software structure. Returning to the AI / ML example discussed above, an AI / ML application can be implemented using a framework such as, but not limited to, Caffe, Caffe2, TensorFlow, Keras, PyTorch, or MxNet deep learning frameworks.
[0260] In at least one embodiment, the processors described elsewhere herein, such as, but not limited to, the processors in the Fig. 9-21, comprise one or more circuits to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein. One or more circuits can be configured by software, such as the programming platforms described herein, to use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates and / or otherwise perform any of the operations described above or elsewhere herein.
[0261] Fig. Figure 23 shows how to compile code for execution on one of the programming platforms described above. Fig. 22 according to at least one embodiment. A compiler 2301 is configured to receive source code 2300, compile source code 2300, and output an executable file 2310. The compiler 2301 can be configured to convert the source code 2300 into host-executable code 2307 and device-executable code 2308. The source code 2300 can be compiled either offline before the execution of an application or online during the execution of an application. The source code 2300 can include code in any programming language supported by the compiler 2301, such as, but not limited to, C++, C, Fortran, etc. The source code 2300 can be contained in a single source file that contains a mixture of host code and device code, specifying the positions of the device code within it. A single source file can be a .cu file containing CUDA code, or a .hip file.The source code 2300 may consist of a .cpp file containing HIP code, or a file in another format containing both host code and device code. Alternatively, instead of a single source file, the source code 2300 may consist of multiple source code files in which host code and device code may be separate. The compiler 2301 includes or incorporates one or more libraries for recognizing a sequence of API calls to execute a single fused API, where a single fused API is a combined API for two or more APIs. In at least one embodiment, the compiler 2301 may be an NVIDIA CUDA compiler (“NVCC”) for compiling CUDA code into .cu files, an HCC compiler for compiling HIP code into .hip.cpp files, or other compilers.
[0262] The 2301 compiler can be configured to compile the 2300 source code into a host-executable code 2307 for execution on a host and a device-executable code 2308 for execution on the device. The 2301 compiler performs operations that include parsing the 2300 source code into an abstract system tree (AST), performing optimizations, and generating executable code. If the 2300 source code comprises a single source file, the 2301 compiler can separate the device code from the host code in such a single source file, compile the device code and the host code into device-executable code 2308 and host-executable code 2307 respectively, and join the device-executable code 2308 and the host-executable code 2307 together in a single file.
[0263] The compiler 2301 can comprise a compiler frontend 2302, a host compiler 2305, a device compiler 2306, and a linker 2309. The compiler frontend 2302 can be configured to separate the device code 2304 from the host code 2303 in the source code 2300. The device code 2304 can be compiled by the device compiler 2306 into a device-executable code 2308, which, as described, can comprise binary code or IR code in at least one embodiment. Separately, the host code 2303 can be compiled by the host compiler 2305 into a host-executable code 2307.For NVCC compilers, such as, but not limited to, those for oneAPL, ROCm, and OpenCL, the host compiler 2305 can be a general-purpose C / C++ compiler that outputs native object code, while the device compiler 2306 can be a Low Level Virtual Machine (LLVM)-based compiler that branches an LLVM compiler infrastructure and outputs PTX code or binary code. For HCC, both the host compiler 2305 and the device compiler 2306 can be LLVM-based compilers that output target binary code.
[0264] After compiling the source code 2300 into host-executable code 2307 and device-executable code 2308, the linker 2309 can link the host-executable code 2307 and the device-executable code 2308 together in an executable file 2310. Native object code for a host and PTX or binary code for a device can be linked in an ELF (Executable and Linkable Format) file, a container format used to store object code. Host-executable code 2307 and device-executable code 2308 can be in any suitable format, for example, but not limited to, binary code and / or IR code. In the case of CUDA, the host-executable code 2307 can, in at least one embodiment, include native object code, and the device-executable code 2308 can include code in PTX intermediate representation.In the case of ROCm, both the host-executable code 2307 and the device-executable code 2308 can, in at least one embodiment, include target binary code. Other implementations, such as, but not limited to, oneAPL and OpenCL, are conceivable and can be implemented similarly to the CUDA and ROCm implementations mentioned above.
[0265] The source code 2300 can be translated before the source code is compiled. The source code is passed through a translation tool (not shown) that translates the source code 2300 into translated source code. A compiler 2301 can be used to compile the translated source code into host-executable code 2307 and device-executable code 2308 in a process similar to the compilation of the source code 2300 by the compiler 2301 into host-executable code 2307 and device-executable code 2308, as described above in conjunction with Fig. 23 explained.
[0266] A translation performed with a translation tool can be used to port the 2300 source code for execution in a different environment than originally intended. The translation tool may include a HIP translator, which is used to "hipify" CUDA code intended for a CUDA platform into HIP code that can be compiled and executed on a ROCm platform using a compiler. The translation of the 2300 source code may involve parsing the 2300 source code and converting calls to APIs provided by one programming model (e.g., CUDA) into corresponding calls to APIs provided by another programming model (e.g., HIP), as described below in conjunction with Fig. This will be explained in more detail in section 24. To return to the example of HIP conversion of CUDA code: Calls to the CUDA runtime API, the CUDA driver API, and / or CUDA libraries can be converted into corresponding HIP API calls. The automated translations performed by translation tool 2301 can sometimes be incomplete, so additional manual work is required to fully port the source code 2300.
[0267] One or more of the techniques described here may employ other methods for converting one code type to another to enable interchangeability between different device architectures. In at least one embodiment, an application for one platform (for example, a CUDA application) may be compiled into code for implementation on another platform (for example, an AMD processor, an Intel processor, or another processor). For example, the source code 2300 may comprise source code for one platform (for example, CUDA). The compiler 2301 may compile the source code 2300 into an executable file 2310 that can be used by another platform (for example, AMD or Intel). Programming toolkits may compile applications for one platform (for example, CUDA) for another platform (for example, AMD or Intel) using a native compiler (for example, natively).For example, a GPGPU programming toolkit can enable the native compilation of CUDA applications for AMD GPUs. Programs (such as CUDA programs) or their build system do not need to show signs of modification or translation into another language before being compiled into code for a different platform. A compiler can accept the same command-line options and programming dialect (such as the CUDA dialect) as another compiler (such as nvcc for CUDA) and serves as a replacement for installing a toolkit (such as the NVIDIA CUDA Toolkit), allowing existing build tools and scripts (such as cmake) to function without further modification. In at least one embodiment, an nvcc-compatible compiler can be used to compile nvcc-dialect CUDA for AMD GPUs, including PTX assembly implementations of CUDA runtime and driver APIs for AMD GPUs.Libraries (such as open-source wrapper libraries) can provide APIs, such as CUDA-X APIs, by delegating to the corresponding ROCm libraries. One example implementation is SCALE from Spectral Compute in London, England. Instead of providing a new method for writing GPGPU software, SCALE allows direct instruction of programs written in the widely used CUDA language for AMD GPUs. Additional implementations can include a Clang compiler, which provides a language frontend and tooling infrastructure for languages in the C family of languages (C, C++, Objective-C / C++, OpenCL, CUDA, and RenderScript).In at least one embodiment, the compilers described herein, such as, but not limited to, compiler 2301, compiler 2305 and / or compiler 2306, can compile one or more circuits for compiling code (for example, CUDA, HIP, OpenCL, oneAPI or others) to use one or more signal-to-noise ratio (SNR) values, to select one or more neural networks, to generate one or more channel estimates and / or to perform one of the operations described above or elsewhere herein.
[0268] Fig. Figure 24 shows a system 2400 configured according to at least one embodiment for compiling and executing CUDA source code 2410 using different types of processing units. The system 2400 comprises CUDA source code 2410, a CUDA compiler 2450, host-executable code 2470(1), host-executable code 2470(2), CUDA device-executable code 2484, a CPU 2490, a CUDA-enabled GPU 2494, a GPU 2492, a CUDA-to-HIP translation tool 2420, HIP source code 2430, an HIP compiler driver 2440, an HCC 2460, and HCC device-executable code 2482.
[0269] CUDA source code 2410 can be a collection of human-readable code written in a CUDA programming language. A CUDA programming language can be an extension of the C++ programming language that includes mechanisms for defining device code and distinguishing between device code and host code. Device code can contain source code that, after compilation, is executable in parallel on a device. A device can be a processor optimized for parallel instruction processing, such as, but not limited to, a CUDA-enabled GPU 2490, a GPU 2492, or another GPGPU, etc. Host code is source code that, after compilation, is executable on a host. A host is a processor optimized for sequential instruction processing, such as, but not limited to, a CPU 2490.
[0270] The CUDA source code 2410 can include any number (including zero) of global functions 2412, any number (including zero) of device functions 2414, any number (including zero) of host functions 2416, and any number (including zero) of host / device functions 2418. Global functions 2412, device functions 2414, host functions 2416, and host / device functions 2418 can be mixed in the CUDA source code 2410. Each of the global functions 2412 can be executable on a device and callable by a host. One or more of the global functions 2412 can therefore serve as entry points for a device. Each of the global functions 2412 can be a kernel. In a technique known as dynamic parallelism, one or more of the global functions 2412 can define a kernel that is executable on a device and callable by such a device.A kernel can be executed N (where N is any positive integer) times in parallel by N different threads on a device during execution.
[0271] Each of the device functions 2414 can be executed on a device and can only be called from such a device. Each of the host functions 2416 can be executed on a host and can only be called from such a host. Each of the host / device functions 2416 can define both a host version of a function, which can be executed on a host and can only be called from such a host, and a device version of the function, which can be executed on a device and can only be called from such a device.
[0272] CUDA source code 2410 can also include any number of calls to any number of functions that can be defined via a CUDA runtime API 2402. The CUDA runtime API 2402 can include any number of functions that run on a host to allocate and release device memory, transfer data between host and device memory, manage multi-device systems, and so on. CUDA source code 2410 can also include any number of calls to any number of functions that can be specified in any number of other CUDA APIs. A CUDA API can be any API designed for use by CUDA code. CUDA APIs can include the CUDA runtime API 2402, a CUDA driver API, APIs for any number of CUDA libraries, and so on, including all APIs described elsewhere herein.Compared to the CUDA runtime API 2402, a CUDA driver API can be a lower-level API, but it allows for finer control of a device. Examples of CUDA libraries include cuBLAS, cuFFT, cuRAND, cuDNN, etc.
[0273] The CUDA compiler 2450 can compile the input CUDA code (for example, the CUDA source code 2410) to produce host-executable code 2470(1) and CUDA device-executable code 2484. The CUDA compiler 2450 can be, but is not limited to, NVCC. The host-executable code 2470(1) can be a compiled version of the host code contained in the input source code and executable on the CPU 2490. The CPU 2490 can be any processor optimized for sequential instruction processing.
[0274] The device executable code 2484 can be a compiled version of the device code contained in the input source code, executable on a CUDA-enabled GPU 2494. The device executable code 2484 can contain binary code. The device executable code 2484 can include IR code, such as, but not limited to, PTX code, which is further compiled at runtime by a device driver into binary code for a specific target device (for example, a CUDA-enabled GPU 2494). The CUDA-enabled GPU 2494 can include any processor optimized for parallel instruction processing that supports CUDA. The CUDA-enabled GPU 2494 may have been developed by NVIDIA Corporation in Santa Clara, California.
[0275] The CUDA-to-HIP translation tool 2420 can be configured to translate CUDA source code 2410 into functionally similar HIP source code 2430. HIP source code 2430 can comprise a collection of human-readable code written in a HIP programming language. HIP code can contain human-readable code written in a HIP programming language. A HIP programming language can include an extension of the C++ programming language that contains functionally similar versions of CUDA mechanisms to define device code and distinguish between device code and host code. A HIP programming language can comprise a subset of the functionality of a CUDA programming language.For example, a HIP programming language includes mechanisms for defining global functions 2412, but such a HIP programming language may not support dynamic concurrency, so global functions 2412 defined in HIP code may only be called from one host.
[0276] The HIP source code 2430 can include any number (including zero) of global functions 2412, any number (including zero) of device functions 2414, any number (including zero) of host functions 2416, and any number (including zero) of host / device functions 2418. The HIP source code 2430 can also include any number of calls to any number of functions that can be specified in a HIP runtime API 2432. The HIP runtime API 2432 can include functionally similar versions of a subset of functions contained in the CUDA runtime API 2402. The HIP source code 2430 can also include any number of calls to any number of functions that can be specified in any number of other HIP APIs. A HIP API can be any API designed for use by HIP code and / or ROCm.HIP APIs can include the HIP runtime API 2432, a HIP driver API, APIs for any number of HIP libraries, APIs for any number of ROCm libraries, etc.
[0277] The CUDA-to-HIP translation tool 2420 can convert any kernel call in CUDA code from CUDA syntax to HIP syntax and convert any number of other CUDA calls in CUDA code to any number of other functionally similar HIP calls. A CUDA call can include a call to a function specified in a CUDA API, and a HIP call can include a call to a function specified in a HIP API. The CUDA-to-HIP translation tool 2420 can convert any number of calls to functions specified in the CUDA runtime API 2402 to any number of calls to functions specified in the HIP runtime API 2432.
[0278] The CUDA-to-HIP translation tool 2420 may include a tool called hipify-perl, which performs a text-based translation process. Alternatively, the CUDA-to-HIP translation tool 2420 may include a tool called hipify-clang, which performs a more complex and robust translation process compared to hipify-perl. In hipify-clang, CUDA code is parsed with clang (a compiler frontend), and the resulting symbols are then translated. The conversion of CUDA code to HIP code may include manual modifications in addition to those performed by the CUDA-to-HIP translation tool 2420.
[0279] The HIP compiler driver 2440 can include a front end that determines a target device 2446 and then configures a compiler compatible with the target device 2446 to compile the HIP source code 2430. The target device 2446 can include a processor optimized for parallel instruction processing. The HIP compiler driver 2440 can determine the target device 2446 in any technically feasible way.
[0280] If the target device 2446 is CUDA-compatible (for example, CUDA-enabled GPU 2494), the HIP compiler driver 2440 can generate a HIP / NVCC compilation instruction 2442. The HIP / NVCC compilation instruction 2442 can configure the CUDA compiler 2450 to compile the HIP source code 2430 using a HIP-to-CUDA translation header and a CUDA runtime. In response to the HIP / NVCC compilation instruction 2442, the CUDA compiler 2450 can generate host-executable code 2470(1) and CUDA device-executable code 2484.
[0281] If the target device 2446 is not CUDA-compatible, the HIP compiler driver 2440 can generate a HIP / HCC compilation command 2444. The HIP / HCC compilation command 2444 can configure HCC 2460 to compile the HIP source code 2430 using an HCC header and a HIP / HCC runtime. In response to the HIP / HCC compilation command 2444, HCC 2460 can generate host-executable code 2470(2) and device-executable code 2482. The HCC device executable code 2482 can be a compiled version of the device code contained in the HIP source code 2430, which is executable on the GPU 2492. The GPU 2492 can be any processor optimized for parallel instruction processing, not CUDA-compatible, and HCC-compatible. The GPU 2492 may have been developed by AMD Corporation of Santa Clara, California. The GPU 2492 may include a non-CUDA-enabled GPU 2492.
[0282] For explanatory purposes only, in Fig. Figure 24 illustrates three different processes that can be implemented in at least one embodiment to compile the CUDA source code 2410 for execution on the CPU 2490 and various devices. A direct CUDA process can compile the CUDA source code 2410 for execution on the CPU 2490 and the CUDA-enabled GPU 2494 without translating the CUDA source code 2410 into HIP source code 2430. An indirect CUDA process can translate the CUDA source code 2410 into HIP source code 2430 and then compile the HIP source code 2430 for execution on the CPU 2490 and the CUDA-enabled GPU 2494. A CUDA / HCC process can translate the CUDA source code 2410 into HIP source code 2430 and then compile the HIP source code 2430 for execution on the CPU 2490 and the GPU 2492 using a compiler.
[0283] A direct CUDA flow that can be implemented is depicted by dashed lines and a series of bubbles with annotations A1-A3. As shown in bubble A1, the CUDA compiler 2450 can receive the CUDA source code 2410 and a CUDA compilation instruction 2448, which can configure the CUDA compiler 2450 to compile the CUDA source code 2410. The CUDA source code 2410, which can be used in a direct CUDA flow, can be written in a CUDA programming language based on a programming language other than C++ (for example, C, Fortran, Python, Java, etc.). In response to the CUDA compilation command 2448, the CUDA compiler 2450 can generate executable code 2470(1) on the host and executable code 2484 on the CUDA device (shown in the bubble with the annotation A2).As illustrated by the bubble with annotation A3, executable code 2470(1) on the host and executable code 2484 on the CUDA device can be executed on the CPU 2490 and the CUDA-enabled GPU 2494, respectively.
[0284] The CUDA device executable code 2484 can include binary code. The device executable code 2484 can include PTX code and can be further compiled into binary code at runtime for a specific target device.
[0285] An indirect CUDA flow that can be implemented is illustrated by dotted lines and a series of bubbles labeled B1-B6. As shown in the bubble labeled B1, the CUDA-to-HIP translation tool 2420 can receive the CUDA source code 2410. As shown in the bubble labeled B2, the CUDA-to-HIP translation tool 2420 can translate the CUDA source code 2410 into HIP source code 2430. As shown in the bubble labeled B3, the HIP compiler driver 2440 can receive the HIP source code 2430 and determine that the target device 2446 is CUDA-enabled.
[0286] As shown in the bubble labeled B4, the HIP compiler driver 2440 can generate a HIP / NVCC compilation instruction 2442 and transfer both the HIP / NVCC compilation instruction 2442 and the HIP source code 2430 to the CUDA compiler 2450. The HIP / NVCC compilation instruction 2442 can configure the CUDA compiler 2450 to compile the HIP source code 2430 using a HIP-to-CUDA translation header and a CUDA runtime. The HIP-to-CUDA translation header can translate any number of mechanisms (for example, functions) specified in any number of HIP APIs into any number of mechanisms specified in any number of CUDA APIs.The CUDA compiler 2450 can use the HIP-to-CUDA translation header in conjunction with a CUDA runtime library that conforms to the CUDA runtime API 2402 to generate host executable code 2470(1) and CUDA device executable code 2484. In response to the HIP / NVCC compilation instruction 2442, the CUDA compiler 2450 can generate host executable code 2470(1) and CUDA device executable code 2484 (illustrated by the bubble with annotation B5). As illustrated by bubble B6, host executable code 2470(1) and CUDA device executable code 2484 can be executed on the CPU 2490 and the CUDA-enabled GPU 2494, respectively. The CUDA device's executable code 2484 can include binary code. The device's executable code 2484 can include PTX code and can be further compiled at runtime into binary code for a specific target device.
[0287] A CUDA / HCC workflow that can be implemented is depicted by solid lines and a series of bubbles labeled C1-C6. As shown in the bubble labeled C1, the CUDA-to-HIP translation tool 2420 can receive the CUDA source code 2410. As shown in the bubble labeled C2, the CUDA-to-HIP translation tool 2420 can translate the CUDA source code 2410 into HIP source code 2430. As shown in the bubble labeled C3, the HIP compiler driver 2440 can receive the HIP source code 2430 and determine that the target device 2446 is not CUDA-enabled.
[0288] The HIP compiler driver 2440 can generate the HIP / HCC compile command 2444 and transfer both the HIP / HCC compile command 2444 and the HIP source code 2430 to HCC 2460 (illustrated by the bubbled annotation C4). The HIP / HCC compile command 2444 can configure HCC 2460 to compile the HIP source code 2430 using an HCC header and a HIP / HCC runtime. The HIP / HCC runtime library can correspond to the HIP runtime API 2432. The HCC header can include any number and type of interoperability mechanisms for HIP and HCC. In response to the HIP / HCC compilation instruction 2444, HCC 2460 can generate on the host executable code 2470(2) and on the device executable code 2482 (shown with the bubble labeled C5).As shown in bubble C6, executable code 2470(2) can be run on the host and executable code 2482 can be run on the CPU 2490 and GPU 2492, respectively, on the HCC device.
[0289] After the CUDA source code 2410 has been translated into HIP source code 2430, the HIP compiler driver 2440 can then be used to generate executable code for the CUDA-enabled GPU 2494 or the GPU 2492 without re-executing the CUDA-to-HIP translation tool 2420. The CUDA-to-HIP translation tool 2420 can translate the CUDA source code 2410 into HIP source code 2430, which is then stored in memory. The HIP compiler driver 2440 can then configure HCC 2460 to generate the host-executable code 2470(2) and the HCC device-executable code 2482 based on the HIP source code 2430. In at least one embodiment, the HIP compiler driver 2440 then configures the CUDA compiler 2450 to generate, based on the stored HIP source code 2430, a host-executable code 2470(1) and a CUDA device-executable code 2484.
[0290] An example kernel can be generated according to at least one embodiment by the CUDA-to-HIP translation tool 2420 from Fig. 24. The CUDA source code 2410 divides an overall problem, for which a specific kernel is designed, into relatively coarse subproblems that can be solved independently using thread blocks. Each thread block contains any number of threads. Each subproblem can be divided into relatively fine parts that can be solved cooperatively in parallel by threads within a thread block. Threads within a thread block can cooperate by exchanging data via shared memory and synchronizing execution to coordinate memory accesses.
[0291] The CUDA source code 2410 can organize thread blocks associated with a given kernel into a one-dimensional, two-dimensional, or three-dimensional grid of thread blocks. Each thread block can contain any number of threads, and a grid can contain any number of thread blocks.
[0292] A kernel can be a function in the device code defined with a "__global__" declaration specifier. The dimension of a raster that executes a kernel for a specific kernel call and associated streams can be specified using a CUDA kernel start syntax. The CUDA kernel start syntax is specified as follows: "KernelName<<<GridSize, BlockSize, GemeinsameSpeicherGröße, Stream> >>(KernelArguments);". An execution configuration syntax can include a "<<<...>>>" construct inserted between a kernel name ("KernelName") and a list of kernel arguments enclosed in parentheses ("KernelArguments"). The CUDA kernel startup syntax can include a CUDA startup function syntax instead of an execution configuration syntax.
[0293] `GridSize` can be of type `dim3` and specify the dimension and size of a grid. The `dim3` type can be a CUDA-defined structure comprising unsigned integers `x`, `y`, and `z`. If `z` is not specified, it can be set to one by default. If `y` is not specified, it can be set to one by default. The number of thread blocks in a grid can be the product of `GridSize.x`, `GridSize.y`, and `GridSize.z`. `BlockSize` can be of type `dim3` and specify the dimension and size of each thread block. The number of threads per thread block can be the product of `BlockSize.x`, `BlockSize.y`, and `BlockSize.z`. Each thread executing a kernel can be assigned a unique thread ID, which can be accessed within the kernel via a built-in variable (for example, `threadldx`).
[0294] Regarding the CUDA kernel start syntax, "SharedMemorySize" can be an optional argument that specifies, in addition to the statically allocated memory, a number of bytes in shared memory that is dynamically allocated per thread block for a given kernel call. Regarding the CUDA kernel start syntax, "SharedMemorySize" can be set to zero by default. Regarding the CUDA kernel start syntax, "Stream" can be an optional argument that specifies an associated stream and is set to zero by default to specify a default stream. A stream can be a sequence of commands (which may be issued by different host threads) that are executed sequentially. Different streams can execute commands related to each other in any order or concurrently.
[0295] The CUDA source code 2410 can include a kernel definition for a sample kernel "MatAdd" and a main function. The main function can be host code that runs on a host and contains a kernel call that causes the MatAdd kernel to execute on a device. The MatAdd kernel can add two matrices A and B of size NxN, where N is a positive integer, and store the result in a matrix C. The main function can define a variable "threadsPerBlock" as 16x16 and a variable "numBlocks" as N / 16xN / 16. The main function can then make the kernel call "MatAdd<<<numBlocks, threadsPerBlock> >>(A, B, C);” Specify. According to the CUDA kernel start syntax, the kernel MatAdd can be executed using a grid of thread blocks with a dimension of N / 16 by N / 16, where each thread block has a dimension of 16 by 16.Each thread block can contain 256 threads, a raster can be created with enough blocks to have one thread per matrix element, and each thread in such a raster can execute the kernel MatAdd to perform pairwise addition.
[0296] When translating CUDA source code 2410 into HIP source code 2430, the CUDA-to-HIP translation tool 2420 can translate each kernel call in the CUDA source code 2410 from CUDA kernel start syntax to HIP kernel start syntax and convert any number of other CUDA calls in the source code 2410 into any number of other functionally similar HIP calls. The HIP kernel start syntax can be expressed as "hipLaunchKernelGGL(KernelName,RasterSize,BlockSize,CommonMemorySize,Stream,KernelArguments);" KernelName, RasterSize, BlockSize, SharedMemorySize, Stream, and KernelArguments can have the same meaning in the HIP kernel start syntax as in the CUDA kernel start syntax (as described earlier herein). The SharedMemorySize and Stream arguments may be required in the HIP kernel start syntax and optional in the CUDA kernel start syntax.
[0297] A section of the HIP source code 2430 can be identical to a section of the illustrated CUDA source code 2410, except for a kernel call that executes the MatAdd kernel on a device. The MatAdd kernel can be defined in the HIP source code 2430 with the same declaration specifier "__global__" as the MatAdd kernel in the CUDA source code 2410. A kernel call in the HIP source code 2430 might be "hipLaunchKernelGGL(MatAdd, numBlocks, threadsPerBlock, 0, 0, A, B, C);", while a corresponding kernel call in the CUDA source code 2410 might be "MatAdd<<<numBlocks, threadsPerBlock> >>(A, B, C);” reads.
[0298] Other implementations are conceivable and can be carried out similarly to the CUDA and HIP implementations mentioned above, for example, oneAPI, OpenCL, and other programming platforms. Code can be translated in either direction. For example, CUDA can be translated to HIP, and CUDA can be translated to OpenCL. SnuCL-Tr and CUCL can be used to translate OpenCL to CUDA and CUDA to OpenCL, respectively. Compiled code or intermediate representations (for example, CUDA-PTX code) can also be translated to run on other processor platforms (for example, AMD or Intel). For example, PTX code can be translated using a translation tool like ZLUDA to run on Intel or AMD processors.
[0299] One or more of the techniques described here can use a oneAPI programming model. A oneAPI programming model can refer to a programming model for interacting with various compute accelerator architectures. OneAPI can refer to an application programming interface (API) designed for interacting with various compute accelerator architectures. A oneAPI programming model can use a DPC++ programming language. A DPC++ programming language can refer to a high-level language for the productivity of data-parallel programming. A DPC++ programming language can be at least partially based on the C and / or C++ programming languages. A oneAPI programming model can be a programming model such as, but not limited to, those developed by Intel Corporation in Santa Clara, California.
[0300] OneAPI and / or the oneAPI programming model can be used to interact with various accelerator, GPU, processor, and / or variations thereof architectures. OneAPI can include a number of libraries that implement different functions. OneAPI can include at least one oneAPI DPC++ library, oneAPI math kernel library, oneAPI data analysis library, oneAPI deep neural network library, oneAPI collective communication library, oneAPI threading component library, oneAPI video processing library, and / or variations thereof.
[0301] A oneAPI DPC++ library, also known as oneDPL, can be a library that implements algorithms and functions to accelerate DPC++ kernel programming. OneDPL can implement one or more Standard Template Library (STL) functions. OneDPL can implement one or more parallel STL functions. OneDPL can provide a range of library classes and functions, such as parallel algorithms, iterators, function object classes, range-based APIs, and / or variations thereof. OneDPL can implement one or more classes and / or functions from a C++ standard library. OneDPL can implement one or more random number generator functions.
[0302] A oneAPl math kernel library, also known as oneMKL, can be a library that implements various optimized and parallelized routines for different mathematical functions and / or operations. OneMKL can implement one or more basic linear algebra subroutines (BLAS) and / or linear algebra packages (LAPACK) with dense linear algebra routines. OneMKL can implement one or more sparse BLAS linear algebra routines. OneMKL can implement one or more random ...
Claims
Processor, comprising: one or more circuits for using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates. Processor according to claim 1, wherein each of the one or more neural networks is to be trained at least partially based on a different range of SNR values. Processor according to claim 1, wherein each of the one or more neural networks is to be trained using one or more physical resource block (PRB) assignments. Processor according to claim 1, wherein the one or more SNR values correspond to one or more wireless signals received from one or more base stations, and wherein the one or more wireless signals further correspond to one or more orthogonal signals from one or more user device devices. Processor according to claim 1, wherein the one or more circuits use the one or more SNR values and a set of one or more physical resource blocks (PRBs) allocated to a user device to select the one or more neural networks. Processor according to claim 1, wherein one or more SNR values are compared with one or more properties of a neural network to select the one or more neural networks. Processor according to claim 1, wherein one or more neural networks are trained to perform channel estimation of wireless signals. System comprising: one or more processors that use one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates. System according to claim 8, wherein each of the one or more neural networks is to be trained at least partially based on a different range of SNR values. System according to claim 8, wherein each of the one or more neural networks is to be trained using one or more physical resource block (PRB) assignments. System according to claim 8, wherein the one or more SNR values correspond to one or more wireless signals received from one or more base stations, and wherein the one or more wireless signals further correspond to one or more orthogonal signals from one or more user device devices. System according to claim 8, wherein the one or more processors use the one or more SNR values and a set of one or more physical resource blocks (PRBs) allocated to a user device to select the one or more neural networks. System according to claim 8, wherein one or more SNR values are compared with one or more properties of a neural network to select the one or more neural networks. System according to claim 8, wherein the one or more neural networks are trained to perform channel estimation of wireless signals. Methods, comprising: Using one or more signal-to-noise ratio (SNR) values to select one or more neural networks to generate one or more channel estimates. Method according to claim 15, wherein each of the one or more neural networks is to be trained at least partially based on a different range of SNR values. The method of claim 15, wherein each of the one or more neural networks is to be trained using one or more physical resource block (PRB) assignments. Method according to claim 15, wherein the one or more SNR values correspond to one or more wireless signals received from one or more base stations, and wherein the one or more wireless signals further correspond to one or more orthogonal signals from one or more user device devices. The method of claim 15, further comprising the use of one or more SNR values and a set of one or more physical resource blocks (PRBs) allocated to a user device to select the one or more neural networks. Method according to claim 15, wherein the one or more SNR values are compared with one or more properties of a neural network in order to select the one or more neural networks.
Citation Information
Patent Citations
Machine Learning-Based Channel Estimation
US20220376957A1
Signaling for additional training of neural networks for multiple channel conditions
US20230021835A1