Data disenchantment using various neural diffusion networks

Neural diffusion networks trained through diffusion distillation techniques address inefficiencies in neural networks by enabling efficient denoising of different data types, reducing computational resources and enhancing performance in generating high-quality data.

DE102025138451A1Pending Publication Date: 2026-03-26NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing neural networks face inefficiencies in memory, time, and computing resource usage due to the architecture and size of parameters, which impact their performance in various tasks.

Method used

Implementing neural diffusion networks for denoising data by training multi-stage and single-stage networks using diffusion distillation techniques, allowing for efficient denoising of different data types with reduced computational resources.

Benefits of technology

The solution enables faster and more efficient generation of high-quality denoised data across various data types, utilizing fewer computational resources and improving the performance of neural networks in tasks such as image, video, and audio generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Devices, systems, and techniques for denoising inference or training data of neural networks are described. In at least one embodiment, the denoising of inference or training data of neural networks can be performed based at least partially on identifying different types of inference or training data within the inference or training data of neural networks, which are to be denoised separately using a corresponding number of neural diffusion networks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] At least one embodiment relates to processing resources used to execute and enable artificial intelligence for various tasks. For example, at least one embodiment relates to processors or computing systems that train or infer using neural networks to denoise data. BACKGROUND

[0002] Artificial intelligence techniques are used to implement various tasks. For example, the architecture and size of parameters in neural networks can significantly impact the use of memory, time, or computing resources for performing different tasks. The amount of memory, time, or computing resources used to perform various tasks can be improved through training or inference using neural networks. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates an example of a logical block diagram implementing noise reduction data using various neural diffusion networks, according to at least one embodiment; Fig. Figure 2 illustrates an example of a logical block diagram using multi-stage and single-stage neural diffusion networks for denoising data using different neural diffusion networks, according to at least one embodiment; Fig. Figure 3 illustrates an example of a logical block diagram implementing diffusion distillation training of single-stage neural networks according to at least one embodiment; Fig. Figure 4 illustrates an example of a logical block diagram implementing an inference system that implements various single-stage neural diffusion networks for different data types, according to at least one embodiment; Fig. Figure 5 illustrates an example of a flowchart implementing noise reduction data using various neural diffusion networks, according to at least one embodiment; Fig. Figure 6 illustrates an example of a flowchart implementing diffusion distillation training of single-stage neural networks according to at least one embodiment; Fig. Figure 7 illustrates an example of a flowchart implementing inference using various single-stage neural diffusion networks for different file types, according to at least one embodiment; Fig. 8A illustrates a logic according to at least one embodiment; Fig. 8B illustrates a logic according to at least one embodiment; Fig. Figure 9 illustrates the training and use of a neural network, according to at least one embodiment; Fig. Figure 10 illustrates an exemplary data center system, according to at least one embodiment; Fig. Figure 11A illustrates an example of an autonomous vehicle, according to at least one embodiment; Fig. 11B illustrates an example of camera locations and fields of view for the autonomous vehicle from Fig. 11A, according to at least one embodiment; Fig. 11C is a block diagram showing an exemplary system architecture for the exemplary autonomous vehicle from Fig. 11A illustrates, according to at least one embodiment; Fig. 11D is a diagram that depicts a system for communication between one or more cloud-based servers and the autonomous vehicle. Fig. 11A illustrates, according to at least one embodiment; Fig. 12 is a block diagram illustrating a computer system according to at least one embodiment; Fig. 13 is a block diagram illustrating a computer system according to at least one embodiment; Fig. Figure 14 illustrates a computer system according to at least one embodiment; Fig. Figure 15 illustrates a computer system according to at least one embodiment; Fig. 16A illustrates a computer system according to at least one embodiment; Fig. 16B illustrates a computer system according to at least one embodiment; Fig. 16C illustrates a computer system according to at least one embodiment; Fig. Figure 16D illustrates a computer system according to at least one embodiment; Fig. 16E and Fig. Figure 16F illustrates a shared programming model, according to at least one embodiment; Fig. 17 illustrates exemplary integrated circuits and associated graphics processors, according to at least one embodiment; Fig. 18A-18B illustrate exemplary integrated circuits and associated graphics processors, according to at least one embodiment; Fig. Figures 19A-19B illustrate an additional exemplary graphics processor logic, according to at least one embodiment; Fig. 20 illustrates a computer system according to at least one embodiment; Fig. 21A illustrates a parallel processor, according to at least one embodiment; Fig. 21B illustrates a partition unit according to at least one embodiment; Fig. 21C illustrates a processing cluster according to at least one embodiment; Fig. Figure 21D illustrates a graphics multiprocessor, according to at least one embodiment; Fig. Figure 22 illustrates a multi-graphics processing unit (GPU) system, according to at least one embodiment; Fig. 23 illustrates a graphics processor according to at least one embodiment; Fig. Figure 24 is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment; Fig. Figure 25 illustrates a deep learning application processor, according to at least one embodiment; Fig. Figure 26 is a block diagram illustrating an exemplary neuromorphic processor according to at least one embodiment; Fig. 27 illustrates at least parts of a graphics processor, according to one or more embodiments; Fig. 28 illustrates at least parts of a graphics processor, according to one or more embodiments; Fig. 29 illustrates at least parts of a graphics processor, according to one or more embodiments; Fig. Figure 30 is a block diagram of a graphics processing machine of a graphics processor, according to at least one embodiment; Fig. 31 is a block diagram of at least parts of a graphics processor core, according to at least one embodiment; Fig. Figures 32A-32B illustrate a thread execution logic that includes an array of processing elements of a graphics processor core, according to at least one embodiment. Fig. Figure 33 illustrates a parallel processing unit (“PPU”) according to at least one embodiment; Fig. Figure 34 illustrates a general processing cluster (“GPC”) according to at least one embodiment; Fig. Figure 35 illustrates a working memory partition unit of a parallel processing unit (“PPU”), according to at least one embodiment; Fig. Figure 36 illustrates a streaming multiprocessor, according to at least one embodiment; Fig. Figure 37 is an exemplary data flow diagram for an advanced computing pipeline, according to at least one embodiment; Fig. Figure 38 is a system diagram for an exemplary system for training, adapting, instantiating and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment; Fig. 39 contains an exemplary illustration of an extended computing pipeline 3810A for processing imaging data, according to at least one embodiment; Fig. 40A contains an exemplary data flow diagram of a virtual instrument supporting an ultrasound device, according to at least one embodiment; Fig. 40B contains an exemplary data flow diagram of a virtual instrument supporting a CT scanner, according to at least one embodiment; Fig. 41A illustrates a data flow diagram for a process for training a machine learning model, according to at least one embodiment; Fig. Figure 41B is an exemplary illustration of a client-server architecture for improving annotation tools with pre-trained annotation models, according to at least one embodiment; and Fig. 42 illustrates components of a system for accessing a large language model, according to at least one embodiment. DETAILED DESCRIPTION

[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not fall within the scope of protection are described herein.

[0004] Disclosed are devices, systems, and techniques that cause inference or training data of neural networks to be denoised. In at least one embodiment, the denoising of inference or training data of neural networks can be performed based at least partially on identifying different types of inference or training data within the inference or training data of neural networks, which are to be denoised separately using a corresponding number of neural diffusion networks.

[0005] The revelation extends to all novel aspects or features described and / or illustrated here.

[0006] Further features of the disclosure are characterized by the independent and dependent claims.

[0007] Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.

[0008] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features here should be interpreted accordingly.

[0009] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.

[0010] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.

[0011] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.

[0012] The disclosure also provides a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out the procedures described herein and / or for embodying the device or system features described herein.

[0013] The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.

[0014] The revelation also provides a signal that transmits one or more of the aforementioned computer programs.

[0015] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.

[0016] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings.

[0017] Fig. Figure 1 illustrates an example of a logical block diagram implementing noise reduction using various neural diffusion networks, according to at least one embodiment. In at least one embodiment, neural diffusion networks implement probabilistic diffusion models or score-based generative models. In at least one embodiment, neural diffusion networks can be trained to perform a reverse diffusion process to remove noise from data conditioned according to an input. In at least one embodiment, neural diffusion networks can be used, for example, to implement generative artificial intelligence by generating new data (e.g., new video, image, or audio data) as an inference in the face of an input (e.g., text, speech, or image prompt).In at least one embodiment, neural diffusion networks can generate new image or video data when a text prompt is provided describing the new image or video data to be created.

[0018] In at least one embodiment, a dataset, such as dataset 110, can be used to train one or more multi-stage neural diffusion networks 130. In at least one embodiment, the one or more multi-stage neural diffusion networks 130 can perform a reverse diffusion process over several steps, repeatedly estimating noise in the data at each step and removing the estimated noise from the data, so that a final denoised version of the data can be provided as an inference output generated by the one or more multi-stage neural networks 130 (e.g., image, video, or sound). In at least one embodiment, for example, the one or more multi-stage neural networks can distribute data p at different time steps t. real , which is disturbed by independent Gaussian noise: p t,real (x t ) = ∫preal (x)q t (x t |x)dx where qt(xt|x)~N(αtx,σt2I) with a given α t , σ t In at least one embodiment, the one or more multi-stage neural diffusion networks learn the score of noisy data, which is considered damaged data. sreal:=∇xtlog pt,real(xt)=−(xt−αtx) / σt2 can be represented by a denoised x: µ(x t ,t) ≈ x is predicted.

[0019] In at least one embodiment, the training of multi-stage neural diffusion networks 120 can be implemented, which trains one or more multi-stage neural diffusion networks 130 using different data types of the dataset 110, such as data types 112a, 112b, 112c, 112d, 112e, and 112f. In at least one embodiment, the training dataset 110 can consist of data pairs, input data (e.g., a prompt or other input data, such as text, speech, or image data), and a ground-truth output (e.g., image, video, audio, or other data) that represents, sounds similar to, or otherwise represents the input data (e.g., "cat" as text input paired with "cat video"). In at least one embodiment, different data types of the training dataset can represent 110 different categories of data within the same data modality (e.g., image data, video data, audio data).In at least one embodiment, for example, different types of training data can represent different types of objects depicted in images (e.g., cats, dogs, trees, flowers, people, cars, trucks, etc.). In at least one embodiment, different types of training data set 110 can represent different data modalities of the same category (e.g., image, video, sound of cats). In at least one embodiment, different data types of the training data set 110 can represent different types of input for the same category and / or modality (e.g., text, speech, or image input to generate sound data of cats).

[0020] In at least one embodiment, the training of multi-stage neural diffusion networks can implement a forward-run diffusion technique that adds noise over multiple steps to ground-truth data that is conditioned or otherwise matches an input, enabling the multi-stage neural network to learn to remove the noise over multiple steps. In at least one embodiment, a score-based generative technique can be used in a forward-run diffusion technique to add levels of noise to ground-truth data as part of training multi-stage neural diffusion networks. In at least one embodiment, a probabilistic diffusion model that probabilistically adds noise over multiple steps by predicting at each step what the data was like before the noise was added was included as part of a forward-run diffusion technique.In at least one embodiment, a stochastic differential equation can be used to add levels of noise to ground-truth data as part of training multi-stage neural diffusion networks.

[0021] In at least one embodiment, the diffusion distillation training of single-stage neural networks 140 can implement the type identification 142 to train the one or more different data-type-specific single-stage neural diffusion networks using different data types of the dataset 110, such as the one or more data-type-specific single-stage neural diffusion networks 150a, the one or more data-type-specific single-stage neural diffusion networks 150b and the one or more data-type-specific single-stage neural diffusion networks 150n, using a multi-stage neural diffusion network 130 as a teacher or source for distillation.In at least one embodiment, the type identification 142 for training one or more different data-type-specific single-stage neural diffusion networks using different data types of the dataset 110 enables data-type-specific neural diffusion networks, such as data-type-specific neural networks 150a, 150b, and 150n, to generate denoised data (e.g., image, video, or sound / audio data) in a single step. In at least one embodiment, data-type-specific neural diffusion networks can have a different architecture than a multi-stage neural diffusion network.In at least one embodiment, data-type-specific neural diffusion networks can, for example, be smaller or designed to operate faster and / or use fewer computational resources if similar (or better) quality denoised data can be generated from the same input as a multi-stage neural diffusion network 130 in significantly less time and / or using significantly fewer computational resources (e.g., memory, GPU, CPU, and / or I / O). In at least one embodiment, performing diffusion distillation training on single-stage neural networks 140 can train several different neural diffusion networks that can perform different generative tasks for different data types in less time and / or using significantly fewer computational resources.

[0022] In at least one embodiment, the type identification 142 can be implemented as part of a one-stage diffusion distillation training to identify different data types of a dataset, such as data types 112a-112f of dataset 110. In at least one embodiment, the type identification 142 can be implemented as a drop-in framework (e.g., a library, a software package, or other software / hardware implementation) that can be combined with a conditional one-stage diffusion distillation procedure, which enhances the quality of data-type-specific neural diffusion networks without affecting the inference speed of data-type-specific neural diffusion networks. In at least one embodiment, the type identification 142 can be implemented to identify different data types of a dataset, such as data types 112a-112f of dataset 110.In at least one embodiment, data types of a data record can be selected or specified according to a user configuration (e.g., one or more parameters specified to a programmatic interface (e.g., API) and / or a command interface (e.g., command line or graphical interface)). In at least one embodiment, a user configuration can select from a set of predefined or standard data type configuration strategies or mechanisms and / or specify instructions, formulas, or other parameters to directly identify different data types of a data record.In at least one embodiment, a user configuration can, for example, specify a number of data types (and corresponding data-type-specific neural diffusion networks to be trained) that can determine a number of data types available for identification during type identification 142. In at least one embodiment, type identification 142 can identify data types according to various criteria that may (or may not) be specified as parameters (e.g., as part of a user configuration). In at least one embodiment, data types can, for example, be determined based on whether data types are disjoint, thereby ensuring that data types are unique at inference time and preventing redundant training.In at least one embodiment, data types can be determined based on whether a number of data elements falling into a data type are of the same (or similar) size, thereby ensuring that different data-type-specific neural diffusion networks function similarly when trained with the same resources. In at least one embodiment, data types can be determined based on whether conditions or properties of elements within a type are more similar than in other groups, thereby ensuring that data-type-specific neural diffusion networks are trained to learn different types of data distribution and achieve training efficiency. In at least one embodiment, a technique for data type identification using class labels (e.g.,A data type identification technique can be implemented using neural networks or another classifier of a machine learning model to label training data pairs as belonging to different classes, in order to have semantically similar data types that are similarly measured. In at least one embodiment, a data type identification technique can be implemented using a pretrained embedding of inputs (e.g., class or prompt embeddings determined from a multi-stage neural diffusion network or other input encoders to generate embeddings) and subsequent clustering of the embeddings to determine a data type as one or more clusters of embeddings.

[0023] Fig. Figure 2 illustrates an example of a logical block diagram for the use of multi-stage and single-stage neural diffusion networks for denoising data using various neural diffusion networks, according to at least one embodiment. At least one embodiment shows a multi-stage diffusion technique for image data, with exemplary denoising steps 200a, 200b, 200c, and finally 200n. As discussed above, in at least one embodiment, a multi-stage diffusion technique using multi-stage neural diffusion networks can accept and generate results for a number of different data types, such as different types of objects (e.g., inanimate objects like the lamp shown at 200n, or living objects (not illustrated)).In at least one embodiment, various single-stage neural diffusion networks trained using single-stage diffusion distillation are illustrated for different types of image data, as shown between 210a and 210b and 220a and 220b, such that one or more multi-stage neural diffusion networks capable of producing the lamp in 200n can be produced in a single step with similar quality to that shown in 220b using a single-stage neural diffusion network. In at least one embodiment, the distillation techniques described above in relation to the diffusion distillation training of single-stage neural networks 140 can be performed to generate a number of steps for noise-reduced data (e.g.,(from hundreds of steps to less than ten steps) by using type identification to generate a different respective small number of neural networks with step diffusion to significantly reduce the number of steps.

[0024] Fig. Figure 3 illustrates an example of a logical block diagram implementing diffusion-distillation training of the single-stage neural network according to at least one embodiment. In at least one embodiment, the type identification 310 for a training data set 302 can be implemented. As above with respect to Fig. As discussed in Figure 1, the type identification 310 can be implemented in different ways to partition or recognize different types of data within the training dataset 302. For example, in at least one embodiment, the type identification 310 can predict or determine class labels to label training data pairs as belonging to different classes, in order to have semantically similar data types that are similarly measured. In at least one embodiment, a type identification 310 can use an encoder or other components that generate embeddings of inputs (e.g., class or prompt embeddings determined from multi-stage neural diffusion networks or other input encoders to generate embeddings) to cluster the embeddings in order to determine a data type as one or more clusters of embeddings.In at least one embodiment, the type identification 310 can be used as a data filter function. Ψ(D,Yk)=Dk:=D|Yk be implemented, the data types as different subsets D k from the training dataset Y k can be described.

[0025] In at least one embodiment, one or more different types of distillation can be used to determine weights for a one-stage neural diffusion network for a given data type. In at least one embodiment, one or more of distribution matching 340, teacher score matching 330, and / or opponent distribution matching 350 can be implemented to generate data-type-specific one-stage neural diffusion networks 322.

[0026] In at least one embodiment, different data types 312 of the training data set 302 can be provided to the diffusion-distillation training of single-stage neural networks 320 together with a neural teacher network 314. In at least one embodiment, the neural teacher network 314 can be a multi-stage neural diffusion network.

[0027] In at least one embodiment, the diffusion-distillation training of single-stage neural networks 320 can implement distribution matching 340. In at least one embodiment, a pre-trained teacher denoising diffusion model (e.g., neural teacher network 302), µ Lehrer , a training dataset D (e.g., data set 302), which may contain real data, fake data, and generated conditions, and a distillation technique Φ. In at least one embodiment, a distillation technique may divide the teacher into single-stage K-generators. {Gk(z;y∈Yk)}k=1Ku¨ber Gk=Φ(μTeacher;Dk=Ψ(D,Yk)),k=1,…,K distill. In at least one embodiment, each distilled student G k be data type-specific and handle a subset of input conditions that data types Yk⊂Y correspond to those based on a subset of training data Dk⊂D can be trained and through Yk can be determined via a filter function Ψ. In at least one embodiment, a obtained generator G can take a random latent z and an input condition (e.g., in the entire condition set). Y) Train the system using output data (e.g., an image, video, or audio). In at least one embodiment, a distribution matching technique can be used as... Gk(1)=ΦDM(μTeacher,ΨDM(DDM,Yk)),k=1,…,K described. In at least one embodiment, distribution matching can be implemented using a closed regression loss or a two-timescale update rule (TTUR). In at least one embodiment, for example, the distribution matching can train a single-stage distilled neural diffusion network (e.g., a student model) to match a generated distribution of a teacher diffusion model (e.g., a multi-stage neural diffusion network) by minimizing inverse KL divergence between teacher and student output distributions, which is diffused over the ambient space at different noise levels for better support, represented as: EtDKL(pt,falsified||pt,real)=Ext(log(pt,falsified(xt)pt,real))

[0028] In at least one embodiment, the training can use a gradient of the inverse KL divergence, which (with fully extended expectation and a user-defined weight w) t ), represented as: ∇θEtDKL≃Ez,t,xt[wt,αt(sfake(xt,t)−real(xt,t)∇θGθ(z)) where z~N(0,I),t~Uniforme[Tmin,Tmax]and xt~q(xt|x), the noise-amplified version of x = G θ (z), which is generated by a single-stage student model. In at least one embodiment, a teacher denoising model can approximate a score from real data, and a "fake" denoising model can approximate a score from generated fake data, represented as: six(xt,t)≈−xt−αtμteacher(xt,t)σt2,sfake(xt,t)≈−xt−αtμffake(xt,t)σt2 where the “fake” denoising model is simultaneously used with the canonical denoising target with weighting λ tThe training process can be represented as follows: Lent noiseϕ=Et[λt‖μfakeϕ(xt,t)−x‖22].

[0029] In at least one embodiment, various strategies can be used to improve generation performance. For example, in at least one embodiment, the KL loss can be completed with a regression loss to promote mode superposition. Lreg=E(z,y)~Dpairedl(Gθ(z),y) where D gepaart which may involve a dataset of latent data pairs generated by an offline teacher model, and which are l This could involve learned perceptual image patch similarity (LPIPS). In at least one embodiment, the distribution matching can apply a TTUR to a generator and a "fake" score model, which enables more stable convergence.

[0030] In at least one embodiment, the diffusion-distillation training of single-stage neural networks 320 can implement adversarial distribution matching 350. In at least one embodiment, adversarial distribution matching (ADM) can be described, for example, as follows: Gk(2)=ΦADM(μLehr,ΨADM(DADM,Yk);Gk(1)),k=1,…,K. In at least one embodiment, an adversarial loss for ADM can be implemented by adding a minimal discriminator head to a bottleneck layer of neural networks, which implements a "fake" denoising model, µ. gefärscht, implemented after distribution matching has been performed. In at least one embodiment, a filter function for distribution matching and opponent distribution matching can be different. For distribution matching, for example, in at least one embodiment, a filter function can have abundant input conditions on a desired type for variational score distribution (VSD) loss, but use the entire matched dataset for regression loss. In at least one embodiment, for opponent distribution matching, a data filter function can be implemented such that both opponent and VSD losses concentrate on a corresponding subset of a training dataset.

[0031] In at least one embodiment, the diffusion distillation training of single-stage neural networks 320 can implement teacher-score matching 330. In at least one embodiment, the teacher-score matching 330 can take into account differences in architecture between a neural teacher network 314 and student models that become single-stage data-type-specific neural networks 322. In at least one embodiment, the teacher-score matching 330 can, for example, handle scenarios in which the architecture of a neural student network is different (e.g., smaller) than that of the neural teacher network 314.In at least one embodiment, the teacher score matching 330 can be performed as a pre-training phase to train a student neural network as if it were a multi-stage neural diffusion network, using a teacher score matching loss function to initialize student neural network weights. In at least one embodiment, a teacher score matching (TSM) loss can be described as follows: LTSM=Et[λt‖μTSMφ(xt,t)−μTeacher(xt,t)‖22] where the smaller student (with weights φ) can be trained to match a teacher's score with real data (e.g., real images) at different noise levels. In at least one embodiment, the teacher score matching can initialize 330 weights for subsequent distillation stages (e.g., distribution matching 340 and / or 350).

[0032] Fig. Figure 4 illustrates an example of a logical block diagram implementing an inference system that implements different single-stage neural diffusion networks for different file types, according to at least one embodiment. In at least one embodiment, an inference system 410 can be implemented as a service as part of a cloud computing service provider (e.g., hosting different single-stage neural diffusion networks on different hosts and forwarding requests to suitable hosts to perform inference). In at least one embodiment, the inference system 410 can be implemented as a standalone system (e.g., on a single host computer system that loads the one or more single-stage neural diffusion networks of different data types to serve different inference requests). In at least one embodiment, the inference request 402 can take an input (e.g.,a condition y (as discussed above). In at least one embodiment, an input can be a prompt (e.g., text, audio, such as speech and / or image data). In at least one embodiment, the type identification 404 can identify a data type of data to be denoised using different single-stage neural diffusion networks of data types, such as one or more single-stage neural diffusion networks of data type A 440a, one or more single-stage neural diffusion networks of data type N 40b, and one or more single-stage neural diffusion networks of data type N 440n. As discussed above, different techniques of the type identification 404 can be implemented in at least one embodiment.In at least one embodiment, the type identification 404 can, for example, generate an embedding of an input from the inference request 402 and use a vector similarity search using the embedding in a data type clustering, wherein a nearest cluster indicates which data type should be used in the one-stage neural diffusion network for a detected data type. In at least one embodiment, a classification model can be used to classify an input, such as a specific category and / or modality, for the inference request 402. In at least one embodiment, a filter function (Ψ), as described above, can be used to determine the nature of the data type (e.g., a subset of a training dataset).

[0033] In at least one embodiment, the inference system 410 can implement request routing 420. In at least one embodiment, the request routing 420 can use a detected data type for the inference request 402 to select one of the data types in single-stage neural diffusion networks. For example, an identifier for a detected data type can identify a model identifier that corresponds to one of the data types in single-stage neural diffusion networks in order to generate, send, and / or otherwise cause the single-stage neural diffusion network of the selected data type to generate denoised data.

[0034] In at least one embodiment, the inference system 410 can provide one or more different single-stage neural diffusion networks for different data types, such as data type A in one or more single-stage neural diffusion networks 440a, data type B in one or more single-stage neural diffusion networks 44b, and data type N in one or more single-stage neural diffusion networks 440n. In at least one embodiment, the inference system 410 can provide different types of denoised data using selected data types in single-stage neural diffusion networks, such as denoised data of type A 442, denoised data of type B 444, and denoised data of type N 446.

[0035] Fig. Figure 5 illustrates an example of a flowchart implementing denoising data using various neural diffusion networks according to at least one embodiment. As stated in Figure 510, inference or training data for neural networks can be obtained in at least one embodiment. In at least one embodiment, inference data for neural networks can be random noise data conditioned according to an input (e.g., a prompt) (e.g., linked to an embedding that is generated, modified, or otherwise changed for an input, or specified according to an input). In at least one embodiment, training data for neural networks for training multiple single-stage neural diffusion networks can be received from or obtained from a multi-stage neural diffusion network.

[0036] As stated in document 520, different types of inference or training data for neural networks can be identified in at least one embodiment. In at least one embodiment, an embedding of an input from an inference request can be determined (e.g., using one or more coding layers of a multi-level diffusion neural network or a text, image, or other encoder pre-trained for a corresponding modality of input data), and a vector similarity search can be performed using the embedding in a data type clustering, with a next cluster indicating which data type a single-level diffusion neural network should use for a detected data type. In at least one embodiment, a classification model can be used to categorize an input, such as...to classify a specific category and / or modality by identifying different types of inference or training data for neural networks. In at least one embodiment, the identification of different types of inference or training data for neural networks can be performed using a filter function (Ψ), as described above.

[0037] As stated in 530, it can be arranged that the inference or training data of neural networks are denoised, at least in part, based on the identified different types of inference or training data, which are to be denoised separately using a corresponding number of neural diffusion networks, in at least one embodiment. As above with respect to Fig. 3 and below with regard to Fig. As discussed in detail in section 6, in at least one embodiment, training data of neural networks can be denoised as part of various distillation techniques that perform one or more training steps to denoise data using a loss that compares a denoised element of training data to a ground-truth label. As discussed above with reference to Fig. 4 and below with reference to Fig. As discussed in detail in section 7, inference data from neural networks can be denoised as part of performing an inference request in order to generate denoised data (e.g., to perform a generative task to generate image data, audio / sound data, and / or video data) and to use a selected one of several single-stage neural diffusion networks for different data types.

[0038] Fig. Figure 6 illustrates an example of a flowchart implementing diffusion-distillation training of single-stage neural networks according to at least one embodiment. As specified in Figure 610, a training dataset can be received in at least one embodiment. In at least one embodiment, distillation techniques can be used, for example, to distill a multi-stage neural network using a training dataset received as part of a software library, framework, and / or hardware support for identifying and accessing a training dataset containing several different types of data.

[0039] As stated in 620, different types of data in the training dataset can be identified in at least one embodiment. In at least one embodiment, embeddings of different elements in a training dataset can be determined (e.g., using one or more coding layers of a multi-level diffusion neural network or a text, image, or other encoder pre-trained on a corresponding modality of input data) and clustered according to a vector similarity technique, such as cosine similarity or other technique clustering vectors according to the distance between vectors in vector space. In at least one embodiment, a cluster of embeddings can specify a data type. In at least one embodiment, a classification model can be used to identify elements of training data, such as...to classify a specific category and / or modality in order to identify data types of inference or training data for neural networks. In at least one embodiment, the identification of different types of training data can be performed using a filter function (Ψ), as described above.

[0040] As stated in 630, in at least one embodiment, a number of single-stage neural diffusion networks corresponding to different types of data can be trained based on a multi-stage neural diffusion network. In at least one embodiment, one or more distillation techniques can be performed to train a number of single-stage neural diffusion networks corresponding to different types of data based on a multi-stage neural diffusion network. As stated in 640, weights of the single-stage neural diffusion networks with respect to the multi-stage neural diffusion network and the different types of data can, for example, be initialized in at least one embodiment as if the multi-stage neural diffusion networks were using a score-matching technique, such as TSM, discussed above.In at least one embodiment, a TSM loss for one or more training steps can be used to compare denoised data generated by a single-stage neural diffusion network (e.g., a student neural network) with denoised data generated using a multi-stage neural diffusion network (e.g., a teacher model). In at least one embodiment, values ​​of weights from a single-stage neural diffusion network can be checked, stored, or otherwise saved after one or more training steps (which can be used as weight values ​​for applying one or more other distillation techniques).

[0041] As stated in reference 650, the weights of the single-stage neural diffusion networks can be updated using distribution matching with respect to the multi-stage neural diffusion network and the different types of data in at least one embodiment. In at least one embodiment, a distribution matching technique can be used to train neural student networks using single-stage diffusion that correspond to different data types identified in a training dataset. In at least one embodiment, a distribution matching technique can use one or more training steps to train a neural student network using single-stage diffusion to match a generated distribution of a multi-stage neural diffusion network using a data type identified from a training dataset.In at least one embodiment, neural student networks can have weights initialized using one-step diffusion, as discussed above in 640. In at least one embodiment, neural student networks can be randomly initialized with values ​​using one-step diffusion.

[0042] As stated in 660, the weights of the single-stage neural diffusion networks can be updated using adversarial distribution matching with respect to the multi-stage neural diffusion network and the different types of data in at least one embodiment. As discussed above, in at least one embodiment, an adversarial matching technique can update weights of single-stage neural student networks that include an adversarial loss for one or more training steps using a prediction for a fake-score model prediction. In at least one embodiment, a final output layer of a single-stage neural student network can map a given vector to a scalar predicted probability and a checkpoint of the generator (e.g.,load a student neural diffusion network from a distribution matching stage, but reinitialize a fake score and a Generative Adversarial Network (GAN) classification model from teacher weights of a multi-stage neural diffusion network.

[0043] Fig. Figure 7 illustrates an example of a flowchart implementing inference using various single-stage neural diffusion networks for different file types, according to at least one embodiment. As stated in Figure 710, in at least one embodiment, an inference request can be received by a neural diffusion network. In at least one embodiment, an inference request can include a prompt or other input that can be used to condition starting noise data that can be denoised to generate denoised data using a neural diffusion network. In at least one embodiment, an inference request can include information that indicates or specifies a data type for generating denoised data (e.g., a classification label corresponding to a data type).

[0044] As stated in 720, in at least one embodiment, a diffusion neural network can be selected from a number of diffusion neural networks based on a type of inference input data in the inference request. In at least one embodiment, a type of data for inference can be identified. In at least one embodiment, as discussed above, an embedding of an input from an inference request can be determined (e.g., using one or more coding layers of a multi-level diffusion neural network or a text, image, or other encoder pre-trained for a corresponding modality of input data), and a vector similarity search can be used using the embedding in a data type cluster, wherein a next cluster specifies which data type a single-level diffusion neural network should use for a detected data type.In at least one embodiment, a classification model can be used to classify an input, such as a specific category and / or modality, by identifying different types of inference or training data for neural networks.

[0045] As stated in 730, the selected neural diffusion network can be caused to generate denoised data using the inference input data in at least one embodiment. In at least one embodiment, a selected neural diffusion network can be loaded into host system hardware for execution, and a request or instruction can be forwarded to a specific host for a selected neural diffusion network. In at least one embodiment, more than one neural diffusion network can be selected (e.g., by identifying more than one data type in an inference request). In at least one embodiment, denoised data can contain various different modalities (e.g., image data, audio / sound data, and / or video data). In at least one embodiment, denoised data can be processed as part of the execution of an inference request (e.g.,Providing video data to a real-time video system (which provides / displays real-time video) to a destination or downstream process. In at least one embodiment, denoised data can be provided to a customer who has submitted an inference request in response. LOGIC

[0046] Fig. Figure 8A illustrates the logic 815, which, as described elsewhere herein, can be used in one or more devices to perform operations such as those discussed herein according to at least one embodiment. In at least one embodiment, the logic 815 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the logic 815 is an inference and / or training logic. Details regarding the logic 815 are given below in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, logic refers to any combination of software logic, hardware logic and / or firmware logic to provide the functionality or operations described herein, wherein the logic may be collectively or individually embodied as a circuit forming part of a larger system, for example, an integrated circuit (IC), a system-on-a-chip (SoC), or one or more processors (e.g., CPU, GPU).

[0047] In at least one embodiment, the logic 815 can, without limitation, include a code and / or data memory 801 for storing forward and / or output weights and / or input / output data and / or other parameters for configuring neurons or layers of a neural network that is trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the logic 815 can include or be coupled to the code and / or data memory 801 to store graph code or other software for controlling the timing and / or sequence in which weight and / or other parameter information for configuring the logic is loaded, including integer and / or floating-point units (collectively, arithmetic logic units, or ALUs).In at least one embodiment, a code, such as a graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, the code and / or data storage 801 stores weight parameters and / or input / output data of each layer of a neural network that was trained or used in conjunction with one or more embodiments during the forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, a portion of the code and / or data storage 801 may be contained in another on-chip or off-chip data storage, including an L1, L2, or L3 cache or system memory of a processor.

[0048] In at least one embodiment, a section of the code and / or data storage 801 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data storage 801 can be a cache memory, dynamically randomly addressable memory (DRAM), static randomly addressable memory (SRAM), non-volatile storage (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether code and / or code and / or data storage 801 is, for example, internal or external to a processor or includes DRAM, SRAM, flash memory or another type of storage, may depend on on-chip versus off-chip available storage, latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.

[0049] In at least one embodiment, the logic 815 can, without limitation, include a code and / or data storage 805 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network that is trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, the code and / or data storage 805 stores weight parameters and / or input / output data of each layer of a neural network that was trained or used in conjunction with one or more embodiments during the backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, the logic 815 may include or be coupled to the code and / or data memory 805 to store graph code or other software for controlling the timing and / or sequence in which weight and / or other parameter information for configuring the logic is to be loaded, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).

[0050] In at least one embodiment, the code, such as the graph code, causes weight or other parameter information to be loaded into processor ALUs based on a neural network architecture to which this code corresponds. In at least one embodiment, a portion of the code and / or data storage 805 may be contained in another on-chip or off-chip data storage, including an L1, L2, or L3 cache or system memory of a processor. In at least one embodiment, a portion of the code and / or data storage 805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 805 is, for example, internal or external to a processor, or whether it comprises DRAM, SRAM, flash memory, or another type of memory, may depend on on-chip versus off-chip available storage, latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.

[0051] In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 can be separate storage structures. In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 can be a combined storage structure. In at least one embodiment, the code and / or data storage 801 and the code and / or data storage 805 can be partially combined and partially separate. In at least one embodiment, a portion of the code and / or data storage 801 and the code and data storage 805 can be contained in another on-chip or off-chip data storage, including an L1, L2, or L3 cache or system memory of a processor.

[0052] In at least one embodiment, the logic 815 can, without limitation, include one or more arithmetic logic units (“ALUs”) 810, comprising integer and / or floating-point units, to perform logical and / or mathematical operations that are at least partially based on or indicated by a training and / or inference code (e.g., graph code), wherein a result thereof can generate activations (e.g., output values ​​of layers or neurons within a neural network) that are stored in an activation memory 820, which are functions of input / output and / or weight parameter data that are stored in the code and / or data memory 801 and / or code and / or data memory 805.In at least one embodiment, activations stored in the activation memory 820 are generated according to linear algebraic and / or matrix-based mathematics, which are performed by ALU(s) 810 in response to the execution of instructions or other code, wherein weight values ​​stored in the code and / or data storage 805 and / or data storage 801 are used as operands together with other values, such as bias values, gradient information, momentum values ​​or other parameters or hyperparameters, some or all of which may be stored in the code and / or data storage 805 or code and / or data storage 801 or some other storage on or off the chip.

[0053] In at least one embodiment, one or more ALUs 810 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 810 may be external to a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, the ALUs 810 may be contained within the execution units of a processor or otherwise within a series of ALUs that can be accessed by the execution units of a processor, either within the same processor or distributed across different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, the code and / or data storage 801, the code and / or data storage 805, and the activation storage 820 can share a processor or other hardware logic device or circuit, while in another embodiment they can be located in different processors or other hardware logic devices or circuits, or in a combination of identical and different processors or other hardware logic devices or circuits. In at least one embodiment, a portion of the activation storage 820 can be included in another on-chip or off-chip data storage, including an L1, L2, or L3 cache or system memory of a processor.Furthermore, the inference and / or training code can be stored with other code that a processor or other hardware logic or circuitry can access and be retrieved and / or processed using a processor's pick-up, decode, scheduling, execution, elimination, and / or other logic circuitry.

[0054] In at least one embodiment, the activation memory 820 can be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or another type of storage. In at least one embodiment, the activation memory 820 can be located wholly or partially within or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation memory 820 is, for example, internal to or external to a processor, or whether it comprises DRAM, SRAM, flash memory, or another type of storage, can depend on on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.

[0055] In at least one embodiment, the logic 815, which is in Fig. Figure 8A illustrates that the logic can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, the logic 815, which is illustrated in Fig. Figure 8A illustrates how they can be used in conjunction with hardware of the Central Processing Unit (CPU), Graphics Processing Unit (GPU), or other hardware, such as Field Programmable Gate Arrays (FPGAs).

[0056] Fig. Figure 8B illustrates logic 815 according to at least one embodiment. In at least one embodiment, logic 815 is inference and / or training logic. In at least one embodiment, logic 815 can, without limitation, include hardware logic in which computing resources are dedicated or otherwise used exclusively in connection with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, logic 815, which is described in Fig. 8B is illustrated, in conjunction with an application-specific integrated circuit (ASIC), such as a Google TensorFlow® processing unit, a Graphcore™ inference processing unit (IPU), or an Intel Corp. Nervana® processor (e.g., “Lake Crest”). In at least one embodiment, the logic 815, which is illustrated in Fig. Figure 8B illustrates the use of logic 815 in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as field-programmable gate arrays (FPGAs). In at least one embodiment, the logic 815 includes, without limitation, a code and / or data memory 801 and a code and / or data memory 805, which can be used to store a code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment, which is illustrated in Figure 8B, the logic 815 includes, without limitation, a code and / or data memory 801 and a code and / or data memory 805, which can be used to store a code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Fig. As illustrated in Figure 8B, each of the code and / or data storage 801 and code and / or data storage 805 is assigned to a dedicated computing resource, such as computing hardware 802 and computing hardware 806, respectively. In at least one embodiment, each of the computing hardware 802 and computing hardware 806 comprises one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in the code and / or data storage 801 and code and / or data storage 805, respectively, the result of which is stored in the activation storage 820.

[0057] In at least one embodiment, each of the code and / or data storage 801 and 805 and the corresponding computing hardware 802 and 806, respectively, corresponds to different layers of a neural network, such that a resulting activation from one storage / computing pair 801 / 802 of the code and / or data storage 801 and the computing hardware 802 is provided as an input to the next storage / computing pair 805 / 806 of the code and / or data storage 805 and the computing hardware 806, in order to reflect a conceptual organization of a neural network. In at least one embodiment, each of the storage / computing pairs 801 / 802 and 805 / 806 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computing pairs (not shown) can be included after or in parallel to the storage / computing pairs 801 / 802 and 805 / 806 in the logic 815. TRAINING AND USE OF A NEURAL NETWORK

[0058] Fig. Figure 9 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, the untrained neural network 906 is trained using a training dataset 902. In at least one embodiment, the training framework 904 is a PyTorch framework, while in other embodiments, the training framework 904 is a TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or another training framework. In at least one embodiment, the training framework 904 trains an untrained neural network 906 and enables it to be trained using the processing resources described herein to generate a trained neural network 908. In at least one embodiment, weights can be selected randomly or by pretraining using a deep belief network.In at least one embodiment, the training can be carried out in a supervised, partially supervised or unsupervised manner.

[0059] In at least one embodiment, the untrained neural network 906 is trained using supervised learning, wherein the training dataset 902 contains an input paired with a desired output for that input, or wherein the training dataset 902 contains an input with a known output, and an output of the neural network 906 is manually evaluated. In at least one embodiment, the untrained neural network 906 is trained in a supervised manner and processes inputs from the training dataset 902 and compares the resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are then propagated back by the untrained neural network 906. In at least one embodiment, the training framework 904 adjusts weights that control the untrained neural network 906.In at least one embodiment, the training framework 904 includes tools to monitor how well the untrained neural network 906 converges to a model, such as the trained neural network 908, that is capable of generating correct answers, as in result 914, based on input data, such as a new dataset 912. In at least one embodiment, the training framework 904 repeatedly trains the untrained neural network 906 while adjusting weights to refine an output of the untrained neural network 906 using a loss function and an adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 904 trains the untrained neural network 906 until the untrained neural network 906 achieves a desired accuracy.In at least one embodiment, the trained neural network 908 can then be provided to implement any number of machine learning operations.

[0060] In at least one embodiment, the untrained neural network 906 is trained using unsupervised learning, wherein the untrained neural network 906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised training dataset 902 includes input data without associated output data or ground-truth data. In at least one embodiment, the untrained neural network 906 can learn groupings within the training dataset 902 and determine how individual inputs are related to the untrained dataset 902. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in the trained neural network 908 that is capable of performing operations useful for reducing the dimensionality of the new dataset 912.In at least one embodiment, unsupervised training can also be used to perform anomaly detection, enabling the identification of data points in the new data set 912 that deviate from normal patterns of the new data set 912.

[0061] In at least one embodiment, semi-supervised learning can be used, which is a technique in which the training dataset 902 contains a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 904 can be used to perform incremental learning, such as through transferred learning techniques. In at least one embodiment, the incremental learning enables the trained neural network 908 to adapt to the new dataset 912 without forgetting the knowledge that was introduced into the trained neural network 908 during the initial training.

[0062] In at least one embodiment, the training framework 904 is a framework that is processed in conjunction with a software development toolkit, such as an OpenVINO (Open Visual Inference and Neural Network Optimization) toolkit. In at least one embodiment, an OpenVINO toolkit is a toolkit such as the one developed by Intel Corporation of Santa Clara, California. In at least one embodiment, OpenVINO includes or uses Logic 815 to perform operations described herein. In at least one embodiment, a system-on-a-chip (SoC), an integrated circuit, or a processor uses OpenVINO to perform operations described herein.

[0063] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications, particularly neural network applications, for various tasks and operations, such as human vision simulation, speech recognition, natural language processing, recommendation systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks, such as convolutional neural networks (CNNs), recurrent and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries, such as OpenCV, OpenCL, and / or variations thereof.

[0064] In at least one embodiment, OpenVINO supports neural network models for various tasks and operations, such as classification, segmentation, object recognition, face recognition, speech recognition, pose estimation (e.g., people and / or objects), monocular depth estimation, image inpainting, style transfer, action recognition, coloring, and / or variations thereof.

[0065] In at least one embodiment, OpenVINO comprises one or more software tools and / or modules for model optimization, also referred to as model optimizers. In at least one embodiment, a model optimizer is a command-line tool that facilitates transitions between training and deployment of neural network models. In at least one embodiment, a model optimizer optimizes neural network models for execution on various devices and / or processing units, such as a GPU, CPU, PPU, GPGPU, and / or variations thereof. In at least one embodiment, a model optimizer creates an internal representation of a model and optimizes the model to generate an intermediate representation. In at least one embodiment, a model optimizer reduces the number of layers of a model. In at least one embodiment, a model optimizer removes layers of a model that are used for training.In at least one embodiment, a model optimizer performs various neural network operations, such as modifying inputs to a model (e.g., changing the size of inputs to a model), modifying the size of inputs to a model (e.g., modifying a batch size of a model), modifying a model structure (e.g., modifying layers of a model), normalization, standardization, quantization (e.g., converting weights of a model from a first representation, such as floating point, to a second representation, such as an integer), and / or variations thereof.

[0066] In at least one embodiment, OpenVINO comprises one or more inference software libraries, also referred to as an inference engine. In at least one embodiment, an inference engine is a C++ library or any suitable programming language library. In at least one embodiment, an inference engine is used to infer input data. In at least one embodiment, an inference engine implements various classes to infer input data and generate one or more results. In at least one embodiment, an inference engine implements one or more API functions to process an intermediate representation, specify input and / or output formats, and / or execute a model on one or more devices.

[0067] In at least one embodiment, OpenVINO provides various capabilities for the heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computing refers to one or more computing processes and / or systems that use one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software functions for executing a program on one or more devices. In at least one embodiment, OpenVINO provides various software functions for executing a program and / or parts of a program on different devices. In at least one embodiment, OpenVINO provides various software functions for, for example, executing a first section of code on a CPU and a second section of code on a GPU and / or an FPGA.In at least one embodiment, OpenVINO provides various software functions to run one or more layers of a neural network on one or more devices (e.g., a first set of layers on a first device, such as a GPU, and a second set of layers on a second device, such as a CPU).

[0068] In at least one embodiment, OpenVINO includes various functions similar to those associated with a CUDA programming model, such as different neural network model operations associated with frameworks like TensorFlow, PyTorch, and / or variations thereof. In at least one embodiment, one or more CUDA programming model operations are performed using OpenVINO. In at least one embodiment, various systems, methods, and / or techniques described herein are implemented using OpenVINO. DATA CENTER

[0069] Fig. Figure 10 illustrates an exemplary data center 1000 in which at least one embodiment can be used. In at least one embodiment, the data center 1000 includes an infrastructure layer 1010, a framework layer 1020, a software layer 1030, and an application layer 1040.

[0070] In at least one embodiment, as in Fig. As shown in Figure 10, the infrastructure layer 1010 of the data center can include a resource orchestrator 1012, clustered compute resources 1014 and node compute resources (“node RRs”) 1016(1)-1016(N), where “N” is a positive integer (which may be a different integer “N” than the one used in other figures). In at least one embodiment, the node RRs 1016(1)-1016(N) can include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units, etc.), storage devices 1018(1)-1018(N) (e.g., dynamic solid-state storage, solid-state storage, or floppy disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules, and / or cooling modules, etc. In at least one embodiment, one or more node RRs can be located among the node RRs.s 1016(1)-1016(N) be a server that has one or more of the computing resources mentioned above.

[0071] In at least one embodiment, the grouped computing resources can include 1014 separate groupings of node RRs located in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). In at least one embodiment, separate groupings of node RRs within grouped computing resources can include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs, including the CPUs or processors, can be grouped within one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, the one or more racks can also include any number of power supply modules, cooling modules and network switches in any combination.

[0072] In at least one embodiment, the resource orchestrator 1012 can configure or otherwise control one or more node RRs 1016(1)-1016(N) and / or grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1012 can include an entity for managing the software design infrastructure (“SDI”) for the data center 1000. In at least one embodiment, the resource orchestrator 1012 can include hardware, software, or a combination thereof.

[0073] In at least one embodiment, as in Fig. As shown in Figure 10, the framework layer 1020 includes a job scheduler 1022, a configuration manager 1024, a resource manager 1026, and a distributed file system 1028. In at least one embodiment, the framework layer 1020 can include a framework that supports the software 1032 of software layer 1030 and / or one or more applications 1042 of application layer 1040. In at least one embodiment, the software 1032 or the one or more applications 1042 can each include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1020 can be a type of free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark"), which can utilize a distributed file system 1028 for processing large amounts of data (e.g., "Big Data"), but are not limited to this.In at least one embodiment, the job scheduler 1022 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 1000. In at least one embodiment, the configuration manager 1024 can be able to configure different layers, such as the software layer 1030 and the framework layer 1020, including Spark and the distributed file system 1028, to support the processing of large amounts of data. In at least one embodiment, the resource manager 1026 can be able to manage clustered or grouped computing resources allocated or assigned to support the distributed file system 1028 and the job scheduler 1022. In at least one embodiment, the clustered or grouped computing resources can include the grouped computing resources 1014 on the infrastructure layer 1010 of the data center.In at least one embodiment, the resource manager 1026 can coordinate with the resource orchestrator 1012 to manage these allocated or assigned computing resources.

[0074] In at least one embodiment, the software contained in software layer 1030 may include software 1032 that is used by at least sections of the node RRs 1016(1)-1016(N), the grouped compute resources 1014, and / or the distributed file system 1028 of framework layer 1020. In at least one embodiment, one or more types of software may include, but are not limited to, web page search software, email virus scanning software, database software, and streaming video content software.

[0075] In at least one embodiment, the applications 1042 contained in the application layer 1040 may include one or more types of applications used by at least sections of the node RRs 1016(1)-1016(N), the grouped compute resources 1014, and / or the distributed file system 1028 of the framework layer 1020. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computation applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0076] In at least one embodiment, the configuration manager 1024, the resource manager 1026, and / or the resource orchestrator 1012 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible way. In at least one embodiment, self-modifying actions can relieve a data center operator of data center 1000 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.

[0077] In at least one embodiment, the data center 1000 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture, using the software and computing resources described above with reference to the data center 1000.In at least one embodiment, trained machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 1000 by using weight parameters calculated by one or more training techniques such as those described herein.

[0078] In at least one embodiment, the data center can use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services.

[0079] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the data center 1000 can be used to infer or predict operations that are based at least partially on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0080] In at least one embodiment, an embodiment of at least one of the Fig. 8A, Fig. 8B, Fig. 9 and / or 10 contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks, according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented. AUTONOMOUS VEHICLE

[0081] Fig. Figure 11A illustrates an example of an autonomous vehicle 1100, according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1100 (alternatively referred to herein as "vehicle 1100") can be, without limitation, a passenger vehicle, such as a car, truck, bus, and / or any other type of vehicle that carries one or more passengers. In at least one embodiment, the vehicle 1100 can be a tractor unit used for transporting cargo. In at least one embodiment, the vehicle 1100 can be an aircraft, a robotic vehicle, or any other type of vehicle.

[0082] Autonomous vehicles can be described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). In at least one embodiment, the vehicle 1100 can exhibit functionality corresponding to one or more of Levels 1 through 5 of the autonomous driving levels.In at least one embodiment, the vehicle 1100 may, for example, be capable of conditional automation (level 3), high automation (level 4) and / or full automation (level 5), depending on the embodiment.

[0083] In at least one embodiment, the vehicle 1100 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components without restriction. In at least one embodiment, the vehicle 1100 can include a drive system 1150, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of drive without restriction. In at least one embodiment, the drive system 1150 can be connected to a drivetrain of the vehicle 1100, which can include a transmission without restriction to enable the propulsion of the vehicle 1100. In at least one embodiment, the drive system 1150 can be controlled in response to the reception of signals from one or more throttle valves / accelerator devices 1152.

[0084] In at least one embodiment, a steering system 1154, which may include a steering wheel, is used to steer the vehicle 1100 (e.g., along a desired path or route) when the drive system 1150 is in operation (e.g., when the vehicle 1100 is in motion). In at least one embodiment, the steering system 1154 can receive signals from one or more steering actuators 1156. In at least one embodiment, a steering wheel can be optional for full automation (Level 5). In at least one embodiment, the brake sensor system 1146 can be used to actuate vehicle brakes in response to receiving signals from one or more brake actuators 1148 and / or brake sensors.

[0085] In at least one embodiment, the one or more controllers 1136, which represent one or more systems-on-chips (“SoCs”) (in Fig. 11A not shown) and / or may include, without limitation, one or more graphics processing units (“GPUs”), providing signals (e.g., representative of instructions) to one or more components and / or systems of the vehicle 1100. For example, in at least one embodiment, one or more controllers 1136 may send signals to operate the vehicle brakes via one or more brake actuators 1148, to operate the steering system 1154 via one or more steering actuators 1156, or to operate the propulsion system 1150 via one or more throttle valves / accelerator devices 1152. In at least one embodiment, one or more controllers 1136 may include one or more built-in (e.g., integrated) computing devices that process sensor signals and issue operating commands (e.g.,output signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1100. In at least one embodiment, one or more controllers 1136 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may perform two or more of the above-mentioned functionalities, two or more controllers may perform a single functionality, and / or any combination thereof.

[0086] In at least one embodiment, one or more controllers 1136 provide signals for controlling one or more components and / or systems of the vehicle 1100 in response to sensor data received from one or more sensors (e.g. sensor inputs). In at least one embodiment, the sensor data can be received, for example, without limitation, by one or more of the following: Global Navigation Satellite Systems (GNSS) sensors 1158 (e.g., global positioning system sensors), radar sensors 1160, ultrasonic sensors 1162, LiDAR sensors 1164, inertial measurement unit (IMU) sensors 1166 (e.g., accelerometers, gyroscopes, a magnetic compass or magnetic compasses, magnetometers, etc.), microphones 1196, stereo cameras 1168, wide-angle cameras 1170 (e.g., fisheye cameras), infrared cameras 1172, environmental cameras 1174 (e.g.,360-degree cameras), long-range cameras (not in . Fig. 11A shown), medium-range cameras (not shown in Fig. 11A shown), speed sensors 1144 (e.g. for measuring the speed of the vehicle 1100), vibration sensors 1142, steering sensors 1140, brake sensors (e.g. as part of the brake sensor system 1146) and / or other sensor types.

[0087] In at least one embodiment, one or more of the controllers 1136 can receive inputs (e.g., in the form of input data) from an instrument cluster 1132 of vehicle 1100 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 1134, an acoustic alarm, a loudspeaker, and / or via other components of vehicle 1100. In at least one embodiment, the outputs can include information such as vehicle speed, rotational speed, time, map data (e.g., a high-definition map (in Fig. 11A not shown), location data (e.g., the location of vehicle 1100, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by one or more controllers 1136, etc. In at least one embodiment, for example, information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.) can be displayed on the HMI display 1134.

[0088] In at least one embodiment, the vehicle 1100 further includes a network interface 1124, which can use one or more wireless antennas 1126 and / or one or more modems for communication over one or more networks. In at least one embodiment, the network interface 1124 can, for example, be suitable for communication over Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000) networks, etc. In at least one embodiment, one or more wireless antennas 1126 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.).) enable the use of one or more local networks, such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or low power wide area networks (“LPWANs”), such as LoRaWAN, SigFox protocols, etc.

[0089] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the vehicle 1100 can be used to infer or predict operations that are at least partially based on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0090] In at least one embodiment, an embodiment of at least one of Fig. 11A include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0091] Fig. Figure 11B illustrates an example of camera locations and fields of view for the autonomous vehicle 1100. Fig. 11A, according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are exemplary and are not intended to be limiting. For example, at least one embodiment may include additional and / or alternative cameras and / or cameras may be located at different points on the vehicle 1100.

[0092] In at least one embodiment, camera types may include, but are not limited to, digital cameras designed for use with components and / or systems of the vehicle 1100. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. In at least one embodiment, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the cameras may use roller shutters, global shutters, another type of shutter, or a combination thereof.In at least one embodiment, the color filter array can include a red-clear-clear-clear color filter array (RCCC), a red-clear-clear-blue color filter array (RCCB), a red-blue-green clear color filter array (RBGC), a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, can be used to increase light sensitivity.

[0093] In at least one embodiment, one or more of the cameras can be used to perform advanced driver assistance systems (ADAS) (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign recognition, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0094] In at least one embodiment, one or more cameras can be mounted in a bracket, e.g., in a specially designed (three-dimensional ("3D") printed) bracket, to eliminate stray light and reflections from inside the vehicle 1100 (e.g., reflections of the dashboard in the windshield) that could impair the camera's image acquisition capabilities. With regard to the mounting of exterior mirrors, in at least one embodiment, exterior mirror assemblies can be individually 3D printed so that a camera mounting plate is adapted to the shape of an exterior mirror. In at least one embodiment, one or more cameras can be integrated into exterior mirrors. In at least one embodiment, for side cameras, one or more cameras can also be integrated into four pillars at each corner of a cabin.

[0095] In at least one embodiment, cameras with a field of view that includes sections of the environment in front of the vehicle 1100 (e.g., forward-facing cameras) can be used for environmental viewing to help identify forward paths and obstacles, and to provide, with the aid of one or more of the controllers 1136 and / or control SoCs, information that is crucial for creating an occupancy grid and / or determining preferred vehicle paths. In at least one embodiment, forward-facing cameras can be used to perform many of the similar ADAS functions as LiDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance.In at least one embodiment, forward-facing cameras can also be used for ADAS functions and systems that include lane departure warnings (LDW), autonomous cruise control (ACC) and / or other functions such as traffic sign recognition without restriction.

[0096] In at least one embodiment, a plurality of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform containing a CMOS (complementary metal oxide semiconductor) color imager. In at least one embodiment, a wide-angle camera 1170 can be used to detect objects entering the field of view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although only one wide-angle camera 1170 is used in Fig. As illustrated in Figure 11B, in other embodiments the vehicle 1100 can have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of long-range cameras 1198 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, one or more long-range cameras 1198 can also be used for object detection and classification, as well as for basic object tracking.

[0097] In at least one embodiment, any number of stereo cameras 1168 can also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 1168 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. In at least one embodiment, such a unit can be used to create a 3D map of the vehicle 1100's environment, including a distance estimate for all points in an image.In at least one embodiment, one or more alternative stereo cameras 1168 can, without limitation, include one or more compact stereo vision sensors, which can, without limitation, include two camera lenses (one each on the left and right) and an image processing chip. These sensors can measure the distance between the vehicle 1100 and the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 1168 can be used in addition to or as an alternative to those described herein.

[0098] In at least one embodiment, cameras with a field of view that includes sections of the environment to the side of the vehicle 1100 (e.g., side cameras) can be used for the surround view by providing information that is used to create and update an occupancy grid and to generate side-impact collision warnings. For example, in at least one embodiment, one or more surround cameras 1174 (e.g., four surround cameras, as in Fig. (Illustrated in Figure 11B) are positioned on the vehicle 1100. In at least one embodiment, one or more surround-view cameras 1174 can include any number and combination of wide-angle cameras, fisheye cameras, 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye cameras can be positioned on the front, rear, and sides of the vehicle 1100. In at least one embodiment, the vehicle 1100 can use three surround-view cameras 1174 (e.g., left, right, and rear) and one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0099] In at least one embodiment, cameras with a field of view that includes sections of the environment behind the vehicle 1100 (e.g., reversing cameras) can be used for parking assistance, surround view, rear-impact warnings, and the creation and updating of an occupancy grid. In at least one embodiment, a plurality of cameras can be used, including, but not limited to, cameras that are also suitable as one or more forward-facing cameras (e.g., one or more long-range cameras 1198 and / or one or more medium-range cameras 1176, one or more stereo cameras 1168, one or more infrared cameras 1172, etc.) as described herein.

[0100] In at least one embodiment, an embodiment of at least one of Fig. 11B include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0101] Fig. 11C is a block diagram showing an exemplary system architecture for the exemplary autonomous vehicle 1100. Fig. Figure 11A illustrates at least one embodiment. In at least one embodiment, each of the components, features, and systems of vehicle 1100 is shown. Fig. Figure 11C is illustrated as being connected via bus 1102. In at least one embodiment, bus 1102 can, without restriction, include a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, a CAN bus can be a network within the vehicle 1100, used to support the control of various features and functions of the vehicle 1100, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 1102 can be configured to have dozens or even hundreds of nodes, each of which has its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1102 can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators.In at least one embodiment, the bus 1102 can be a CAN bus that is ASIL B compliant.

[0102] In at least one embodiment, FlexRay and / or Ethernet protocols can be used in addition to or as an alternative to CAN. In at least one embodiment, there can be any number of buses constituting Bus 1102, which can include, without restriction, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using a different protocol. In at least one embodiment, two or more buses can be used to perform different functions and / or for redundancy. For example, a first bus can be used for collision avoidance functionality and a second bus for actuation control.In at least one embodiment, each bus 1102 can communicate with one of the components of the vehicle 1100, and two or more buses of bus 1102 can communicate with the corresponding components. In at least one embodiment, each of any number of one or more system-on-chips (“SoCs”) 1104 (such as SoC 1104(A) and SoC 1104(B)), each of the controllers 1136, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from sensors of the vehicle 1100) and can be connected to a common bus, such a CAN bus.

[0103] In at least one embodiment, vehicle 1100 may include one or more controllers 1136 as described herein with reference to Fig. 11A. In at least one embodiment, one or more controllers 1136 can be used for a variety of functions. In at least one embodiment, one or more controllers 1136 can be coupled with one of the various other components and systems of the vehicle 1100 and can be used for controlling the vehicle 1100, for the artificial intelligence of the vehicle 1100, for infotainment for the vehicle 1100, and / or the like.

[0104] In at least one embodiment, the vehicle 1100 can include any number of SoCs 1104. In at least one embodiment, each of the SoCs 1104 can, without limitation, include central processing units (“CPUs”) 1106, graphics processing units (“GPUs”) 1108, one or more processors 1110, one or more caches 1112, one or more accelerators 1114, one or more data storage devices 1116, and / or other components and features not illustrated. In at least one embodiment, one or more SoCs 1104 can be used to control the vehicle 1100 in a variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 1104 can be combined in a system (e.g., the system of the vehicle 1100) with a high-definition (“HD”) card 1122, which is accessed via a network interface 1124 by one or more servers (in Fig. (11C not shown) may receive map refreshes and / or updates.

[0105] In at least one embodiment, one or more CPUs 1106 can comprise a CPU cluster or CPU complex (hereinafter alternatively referred to as "CCPLEX"). In at least one embodiment, one or more CPUs 1106 can comprise multiple cores and / or Level 2 ("L2") caches. For example, in at least one embodiment, one or more CPUs 1106 can comprise eight cores in a coherent multiprocessor configuration. In at least one embodiment, one or more CPUs 1106 can comprise four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2-megabyte (MB) L2 cache). In at least one embodiment, one or more CPUs 1106 (e.g., CCPLEX) can be configured to support the simultaneous operation of clusters, so that any combination of clusters of one or more CPUs 1106 can be active at any given time.

[0106] In at least one embodiment, one or more CPUs can implement 1106 power management functions that, without limitation, include one or more of the following features: Individual hardware blocks can be automatically clock-controlled when idle to save dynamic power; each core clock can be controlled when such core is not actively executing instructions due to the execution of Wait for Interrupt ("WFI") / Wait for Event ("WFE") instructions; each core can be independently power-controlled; each core cluster can be independently clock-controlled when all cores are clock-controlled or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled.In at least one embodiment, one or more CPUs 1106 can further implement an improved power state management algorithm in which permissible power states and expected wake-up times are defined, and the hardware / microcode determines the best power state to input for the core, cluster, and CCPLEX. In at least one embodiment, processing cores can support simplified sequences for inputting the power state into the software, thereby offloading the work to the microcode.

[0107] In at least one embodiment, one or more GPUs 1108 can include an integrated GPU (hereinafter referred to as the "iGPU"). In at least one embodiment, one or more GPUs 1108 can be programmable and efficient for parallel workloads. In at least one embodiment, one or more GPUs 1108 can use an extended Tensor instruction set. In at least one embodiment, one or more GPUs 1108 can include one or more streaming microprocessors, each streaming microprocessor being able to include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96 KB of memory), and two or more streaming microprocessors being able to share an L2 cache (e.g., an L2 cache with 512 KB of memory). In at least one embodiment, one or more GPUs 1108 can include at least eight streaming microprocessors.In at least one embodiment, one or more GPUs 1108 can use one or more application programming interfaces (APIs) for computations. In at least one embodiment, one or more GPUs 1108 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).

[0108] In at least one embodiment, one or more GPUs 1108 can be energy-optimized for best performance in automotive and embedded applications. In at least one embodiment, one or more GPUs 1108 could, for example, be fabricated on a Fin field-effect transistor (“FinFET”). In at least one embodiment, each streaming microprocessor can contain a number of mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 FP64 cores could be divided into four processing blocks. In at least one embodiment, each processing block could be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, a level 0 ("L0") instruction cache, a scheduler (e.g., a warp scheduler) or sequencer, a dispatch unit, and / or a 64 KB register bank.In at least one embodiment, the streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computations and addressing operations. In at least one embodiment, the microprocessors can include independent thread scheduling to enable fine-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessors can include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.

[0109] In at least one embodiment, one or more of the GPUs 1108 can include high-bandwidth memory (“HBM”) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In at least one embodiment, synchronous graphics random-access memory (“SGRAM”), such as double-data-rate type five synchronous graphics random-access memory (“GDDR5”), can be used in addition to or as an alternative to the HBM memory.

[0110] In at least one embodiment, one or more GPUs 1108 can incorporate unified memory technology. In at least one embodiment, support for address translation services (ATS) can be used so that one or more GPUs 1108 can directly access the page tables of one or more CPUs 1106. In at least one embodiment, if a GPU of the memory management unit (MMU) of one or more GPUs 1108 fails, an address translation request can be sent to one or more CPUs 1106. In at least one embodiment, two CPUs of one or more CPUs 1106 can, in response, search their page tables for the virtual physical mapping for the address and transmit the translation back to one or more GPUs 1108.In at least one embodiment, the unified memory technology can enable a single unified virtual address space for the main memory of both one or more CPUs 1106 and one or more GPUs 1108, thereby simplifying the programming of one or more GPUs 1108 and the porting of applications to one or more GPUs 1108.

[0111] In at least one embodiment, one or more GPUs 1108 can include any number of access counters that can track the frequency of access by one or more GPUs 1108 to the memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved to the physical memory of a processor that accesses pages most frequently, thereby improving the efficiency of memory areas shared by processors.

[0112] In at least one embodiment, one or more SoCs 1104 can include any number of caches 1112, including those described herein. In at least one embodiment, for example, one or more caches 1112 could include a Level 3 ("L3") cache available to both one or more CPUs 1106 and one or more GPUs 1108 (e.g., the one associated with one or more CPUs 1106 and one or more GPUs 1108). In at least one embodiment, one or more caches 1112 can include a write-back cache capable of tracking the states of the rows, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, an L3 cache can include 4 MB of memory or more, depending on the embodiment, although smaller cache sizes can also be used.

[0113] In at least one embodiment, one or more SoCs 1104 can include one or more accelerators 1114 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoCs 1104 can, for example, include a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large on-chip memory. In at least one embodiment, a large on-chip memory (e.g., 4 MB SRAM) can enable a hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, a hardware acceleration cluster can be used as an adjunct to one or more GPUs 1108 and offload some of the tasks from one or more GPUs 1108 (e.g., to free up more cycles of one or more GPUs 1108 for performing other tasks).In at least one embodiment, one or more accelerators 1114 could be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc.) that are stable enough to be suitable for acceleration. In at least one embodiment, a CNN could include a region-based or regional convolutional neural network (RCNN) and fast RCNNs (e.g., as used for object detection) or another type of CNN.

[0114] In at least one embodiment, one or more accelerators 1114 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). In at least one embodiment, one or more DLAs can, without limitation, include one or more tensor processing units (“TPUs”) configured to provide an additional ten trillion operations per second for deep learning applications and inference. In at least one embodiment, one or more TPUs can be accelerators configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can further be optimized for a specific set of neural network types and floating-point operations, as well as for inference.In at least one embodiment, the design of one or more DLAs can provide more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU. In at least one embodiment, one or more TPUs can perform multiple functions, including a convolution function for a single instance that supports, for example, INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, in particular CNNs, on processed or unprocessed data for a variety of functions, including, for example, but not limited to: a CNN for identifying and detecting objects using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for detecting and identifying emergency vehicles using data from microphones; a CNN for facial recognition and identifying vehicle owners using data from camera sensors; and / or a CNN for safety and / or security-related events.

[0115] In at least one embodiment, one or more DLAs can perform each function of one or more GPUs 1108, and by using an inference accelerator, a developer can, for example, target either one or more DLAs or one or more GPUs 1108 for each function. For example, in at least one embodiment, a developer can concentrate the processing of CNNs and floating-point operations on one or more DLAs and leave other functions to one or more GPUs 1108 and / or one or more accelerators 1114.

[0116] In at least one embodiment, one or more accelerators 1114 may include a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, one or more PVAs may be designed and configured to accelerate computer vision algorithms for an advanced driver assistance system (“ADAS”) 1138, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, the PVA may provide a balance between performance and flexibility.For example, in at least one embodiment, it can include any PVA, for example, and without limitation, any number of cores of Reduced Instruction Set Computers (RISC), Direct Memory Access (DMA), and / or any number of vector processors.

[0117] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any camera described herein), one or more image signal processors, etc. In at least one embodiment, each RISC core can include any amount of memory. In at least one embodiment, RISC cores can use any number of protocols, depending on the embodiment. In at least one embodiment, RISC cores can run a real-time operating system (RTOS). In at least one embodiment, RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. In at least one embodiment, RISC cores could, for example, include an instruction cache and / or tightly coupled RAM.

[0118] In at least one embodiment, DMA can enable PVA components to access system memory independently of one or more CPUs. In at least one embodiment, DMA can support any number of features that serve to optimize a PVA, including, but not limited to, supporting multidimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0119] In at least one embodiment, the vector processors can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, a vector processing subsystem can operate as the primary processing machine of a PVA and can include a vector processing unit (VPU), an instruction cache, and / or a working memory (e.g., VMEM).In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction, multiple data (SIMD) and very long instruction word (VLIW) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve data throughput and speed.

[0120] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to dedicated memory. Therefore, in at least one embodiment, each vector processor can be configured to operate independently of other vector processors. In at least one embodiment, the vector processors included in a given PVA can be configured to utilize data parallelism. For example, in at least one embodiment, the plurality of vector processors included in a single PVA can execute a common computer vision algorithm, but on different regions of an image.In at least one embodiment, vector processors contained in a particular PVA can simultaneously execute different computer vision algorithms on an image, or even different algorithms on successive images or sections of an image. In at least one embodiment, an arbitrary number of PVAs can be included in the hardware acceleration cluster, and an arbitrary number of vector processors can be included in each of the PVAs. In at least one embodiment, a PVA can include additional memory for an error-correcting code (ECC) to enhance the overall system safety.

[0121] In at least one embodiment, one or more accelerators 1114 can include an on-chip computer vision network and static random-access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for one or more accelerators 1114. In at least one embodiment, the on-chip memory can include at least 4 MB of SRAM, comprising, for example, and without limitation, eight field-configurable memory blocks accessible to both a PVA and a DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used.In at least one embodiment, a PVA and a DLA can access main memory via a backbone that enables high-speed access to the main memory for both the PVA and the DLA. In at least one embodiment, a backbone can include an on-chip computer vision network that connects a PVA and a DLA to main memory (e.g., using an APB).

[0122] In at least one embodiment, an on-chip computer vision network can include an interface that determines, prior to the transmission of control signals / addresses / data, that both a PVA and a DLA provide ready-to-use and valid signals. In at least one embodiment, an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. In at least one embodiment, an interface can conform to the standards of the International Organization for Standardization (ISO) 26262 or the International Electrotechnical Commission (IEC) 61508, although other standards and protocols may be used.

[0123] In at least one embodiment, one or more of the SoCs 1104 can include a real-time ray tracing hardware accelerator. In at least one embodiment, a real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) for generating real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or purposes.

[0124] In at least one embodiment, one or more accelerators 1114 can have a wide range of uses for autonomous driving. In at least one embodiment, a PVA can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of the PVA are well suited to algorithmic areas that require predictable processing with low power consumption and low latency. In other words, a PVA is well suited for semi-dense or dense regular computations, even with small datasets, that require predictable runtimes with low latency and low power consumption. In at least one embodiment, such as in vehicle 1100, PVAs could be designed to execute classical computer vision algorithms, as they can be efficient at object detection and operate with integer mathematics.

[0125] For example, according to at least one embodiment of the technology, a PVA is used to perform computer stereovision. In at least one embodiment, a semi-global matching algorithm can be used, although this is not intended as a limitation. In at least one embodiment, applications for Level 3-5 autonomous driving use spontaneous motion estimation or spontaneous stereo matching (e.g., structure of motion, pedestrian detection, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereovision functions on input from two monocular cameras.

[0126] In at least one embodiment, a PVA can be used to perform a dense optical flow. For example, in at least one embodiment, a PVA could process raw radar data (e.g., using a four-dimensional fast Fourier transform) to provide processed radar data. In at least one embodiment, a PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to provide processed time-of-flight data.

[0127] In at least one embodiment, a DLA can be used to operate any type of network to improve control and driving safety, including, for example, but not limited to, a neural network that outputs a confidence measure for each object detection. In at least one embodiment, confidence can be represented or interpreted as a probability or as providing a relative "weight" for each detection compared to other detections. In at least one embodiment, a confidence measure enables a system to make further decisions about which detections should be considered true positives rather than false positives. In at least one embodiment, a system can set a confidence threshold and consider only detections exceeding the threshold as true positives.In at least one embodiment where an automatic emergency braking (AEB) system is used, false positive detections would cause a vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered triggers for AEB. In at least one embodiment, a DLA can execute a neural network to regress the confidence value. In at least one embodiment, a neural network can take as input at least a subset of parameters, such as, among others, the dimensions of the bounding box, a (e.g.,ground plane estimate obtained from another subsystem, output from one or more inertial measurement unit (IMU) sensors 1166 that correlate with the orientation of the vehicle 1100, distance, 3D location estimates of the object obtained from the neural network and / or other sensors (e.g. one or more LIDAR sensors 1164 or one or more RADAR sensors 1160).

[0128] In at least one embodiment, one or more of the SoCs 1104 can include one or more data memories 1116. In at least one embodiment, one or more data memories 1116 can be on-chip memory of one or more SoCs 1104 in which neural networks can be stored that are to be executed on one or more GPUs 1108 and / or a DLA. In at least one embodiment, one or more data memories 1116 can be large enough to store multiple instances of neural networks for redundancy and security. In at least one embodiment, one or more data memories 1116 can include one or more L2 or L3 caches.

[0129] In at least one embodiment, one or more SoCs 1104 can include any number of processors 1110 (e.g., embedded processors). In at least one embodiment, one or more processors 1110 can include a boot and power management processor, which can be a dedicated processor and subsystem for handling boot power and management functions and the associated security enforcement. In at least one embodiment, a boot and power management processor can be part of a boot sequence of one or more SoCs 1104 and can provide runtime power management services.In at least one embodiment, a boot and power management processor can provide clock and voltage programming, support for system transitions to a low-power state, thermal and temperature sensor management of one or more SoCs 1104, and / or energy state management of one or more SoCs 1104. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and one or more SoCs 1104 can use ring oscillators to sense the temperatures of one or more CPUs 1106, one or more GPUs 1108, and / or one or more accelerators 1114.In at least one embodiment, if it is determined that the temperatures exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoCs 1104 into a lower power state and / or put the vehicle 1100 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 1100 to a safe stop).

[0130] In at least one embodiment, one or more processors 1110 may further include an array of embedded processors that can serve as an audio processing machine, which may be an audio subsystem enabling full hardware support for multi-channel audio over multiple interfaces and including a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing machine is a dedicated processor core with a digital signal processor and dedicated RAM.

[0131] In at least one embodiment, one or more processors 1110 may further include an always-on processor engine, which can provide the necessary hardware functions to support low-power sensor management and the waking of use cases. In at least one embodiment, an always-on processor engine may, without limitation, include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0132] In at least one embodiment, one or more processors 1110 can further include a security cluster engine, which without limitation includes a dedicated processor subsystem for the security management of automotive applications. In at least one embodiment, a security cluster engine can without limitation include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a security mode, two or more cores in at least one embodiment can operate in a lockstep mode and function as a single core with comparison logic to detect any differences between their operations.In at least one embodiment, one or more processors 1110 may further include a real-time camera engine, which may, without limitation, include a dedicated processor subsystem for managing the real-time camera. In at least one embodiment, one or more processors 1110 may further include a high dynamic range signal processor, which may, without limitation, include an image signal processor that is a hardware machine that is part of a camera processing pipeline.

[0133] In at least one embodiment, one or more processors 1110 can include a video image compositor, which can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate a final image for a player window. In at least one embodiment, a video image compositor can perform lens distortion correction on one or more wide-angle cameras 1170, one or more ambient cameras 1174, and / or on one or more sensors of the surveillance camera in the cabin. In at least one embodiment, one or more sensors of the surveillance camera in the cabin are preferably monitored by a neural network running on another instance of the SoC 1104 and configured to detect events in the cabin and respond accordingly.In at least one embodiment, a system in the cabin can, without restriction, perform lip-reading to activate the mobile phone service and make a call, dictate emails, change a destination, activate or change an infotainment system and vehicle settings, or enable voice-controlled internet browsing. In at least one embodiment, certain functions are available to a driver when a vehicle is operating in autonomous mode and are otherwise deactivated.

[0134] In at least one embodiment, a video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in a video, the noise reduction weights the spatial information accordingly and reduces the weight of information provided by adjacent frames. In at least one embodiment, when a frame or a portion of a frame does not contain motion, the temporal noise reduction performed by the video image compositor can use information from a previous frame to reduce noise in the current frame.

[0135] In at least one embodiment, a video image compositor can also be configured to perform stereo equalization of the input stereo lens images. In at least one embodiment, a video image compositor can also be used for user interface design when an operating system desktop is in use and one or more GPUs 1108 do not need to constantly render new surfaces. In at least one embodiment, when one or more GPUs 1108 are powered on and actively performing 3D rendering, a video image compositor can be used to offload one or more GPUs 1108, thereby improving performance and responsiveness.

[0136] In at least one embodiment, one or more SoCs 1104 may further include a serial camera interface with a mobile industry processor interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for a camera and associated pixel input functions. In at least one embodiment, one or more SoCs 1104 may further include one or more software-controlled input / output controllers that can be used to receive I / O signals that are not assigned to a specific role.

[0137] In at least one embodiment, one or more SoCs 1104 can further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio encoders / decoders (“codecs”), power management, and / or other devices. In at least one embodiment, one or more SoCs 1104 can be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., one or more LiDAR sensors 1164, one or more radar sensors 1160, etc., which may be connected via Ethernet channels), data from bus 1102 (e.g., vehicle speed 1100, steering wheel position, etc.), data from one or more GNSS sensors 1158 (e.g., connected via an Ethernet bus or a CAN bus), etc.In at least one embodiment, one or more SoCs of one or more SoCs 1104 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to offload one or more CPUs 1106 from routine data management tasks.

[0138] In at least one embodiment, one or more SoCs 1104 can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture and effectively and efficiently employing computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack along with deep learning tools. In at least one embodiment, one or more SoCs 1104 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, in at least one embodiment, one or more accelerators 1114, in combination with one or more CPUs 1106, one or more GPUs 1108, and one or more data stores 1116, can form a fast, efficient platform for autonomous vehicles of levels 3-5.

[0139] In at least one embodiment, for example, computer vision algorithms can be executed on CPUs that can be configured using a high-level programming language, such as C, to run a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, as used in in-vehicle ADAS applications and in purpose-built Level 3-5 autonomous vehicles.

[0140] The embodiments described here allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to enable autonomous driving functionality at levels 3-5. In at least one embodiment, for example, a CNN running on the DLA or a discrete GPU (e.g., one or more GPUs 1120) can include text and word recognition, enabling the reading and understanding of traffic signs, including signs for which a neural network has not been specifically trained. Furthermore, in at least one embodiment, a DLA can include a neural network capable of identifying and interpreting a sign, providing a semantic understanding, and passing this semantic understanding to path planning modules running on a CPU complex.

[0141] In at least one embodiment, multiple neural networks can run simultaneously, as when driving at levels 3, 4, or 5. In at least one embodiment, for example, a warning sign with the inscription "Caution: Flashing lights indicate black ice" can be interpreted independently or jointly by multiple neural networks, along with an electric light. In at least one embodiment, such a warning sign itself can be identified as a traffic sign by a first neural network (e.g., a trained neural network), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which informs the vehicle's path planning software (preferably executed on a CPU complex) that black ice is present when flashing lights are detected.In at least one embodiment, a flashing light can be identified across multiple images by a third neural network, which informs a vehicle's path planning software about the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks can run simultaneously, e.g., within a DLA and / or on one or more GPUs 1108.

[0142] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1100. In at least one embodiment, an always-on sensor processing machine can be used to unlock a vehicle when an owner approaches a driver's door and turn on the lights, and to deactivate such a vehicle in security mode when an owner leaves it. In this way, one or more SoCs 1104 provide security against theft and / or carjacking.

[0143] In at least one embodiment, a CNN for detecting and identifying emergency vehicles can use data from microphones 1196 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1104 use a CNN to classify environmental and urban noise as well as visual data. In at least one embodiment, a CNN running on a DLA is trained to detect the relative approach speed of an emergency vehicle (e.g., by using a Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles specific to a local area in which a vehicle is operating, as identified by one or more GNSS sensors 1158.In at least one embodiment, a CNN will attempt to detect European sirens when operating in Europe, and when operating in North America, a CNN will attempt to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slowing down a vehicle, pulling over to the side of the road, parking a vehicle, and / or idling a vehicle, using one or more ultrasonic sensors 1162, until the emergency vehicles pass.

[0144] In at least one embodiment, a vehicle 1100 can include one or more CPUs 1118 (e.g., one or more discrete CPUs or one or more dCPUs) which can be coupled to one or more SoCs 1104 via a high-speed interface (e.g., PCIe). In at least one embodiment, one or more CPUs 1118 can, for example, include an x86 processor. One or more CPUs 1118 can, for example, be used to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and one or more SoCs 1104 and / or monitoring the status and state of one or more controllers 1136 and / or an infotainment system-on-a-chip (“infotainment SoC”) 1130.In at least one embodiment, one or more SoCs 1104 include one or more intermediate connections, and an intermediate connection may include a Peripheral Component Interconnect Express (PCIe).

[0145] In at least one embodiment, a vehicle 1100 can include one or more GPUs 1120 (e.g., one or more discrete GPUs or one or more dGPUs) which can be coupled to one or more SoCs 1104 via a high-speed intermediate connection (e.g., NVIDIA's NVLINK channel). In at least one embodiment, one or more GPUs 1120 can provide additional artificial intelligence functions, e.g., by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least partially on inputs (e.g., sensor data) from sensors of the vehicle 1100.

[0146] In at least one embodiment, vehicle 1100 can further include the network interface 1124, which can include, without limitation, one or more wireless antennas 1126 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, network interface 1124 can be used to enable a wireless connection to cloud services via the internet (e.g., with one or more servers and / or other network devices), to other vehicles, and / or to computing devices (e.g., client devices of passengers). In at least one embodiment, a direct connection can be established between vehicle 1100 and another vehicle, and / or an indirect connection can be established (e.g., via networks and the internet) to communicate with other vehicles.In at least one embodiment, direct connections can be established using a vehicle-to-vehicle communication link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 1100 with information about vehicles in the vicinity of vehicle 1100 (e.g., vehicles in front of, beside, and / or behind vehicle 1100). In at least one embodiment, this functionality can be part of a cooperative adaptive cruise control function of vehicle 1100.

[0147] In at least one embodiment, a network interface 1124 can include a SoC that provides modulation and demodulation functions and enables one or more controllers 1136 to communicate via wireless networks. In at least one embodiment, the network interface 1124 can include a high-frequency front end for upconversion from baseband to high frequency and downconversion from high frequency to baseband. In at least one embodiment, frequency conversions can be performed in any technically feasible manner. The frequency conversions could, for example, be performed using known methods and / or superheterodyne methods. In at least one embodiment, the high-frequency front-end functionality can be provided by a separate chip.In at least one embodiment, network interfaces can include wireless functionality for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN and / or other wireless protocols.

[0148] In at least one embodiment, the vehicle 1100 may further include one or more data storage devices 1128, which may be located outside the chip (e.g., outside one or more SoCs 1104). In at least one embodiment, one or more data storage devices 1128 may, without limitation, include one or more memory elements, including RAM, SRAM, dynamic random access memory (DRAM), video random access memory (VRAM), flash memory, hard disks, and / or other components and / or devices capable of storing at least one bit of data.

[0149] In at least one embodiment, the vehicle 1100 may further include one or more GNSS sensors 1158 (e.g., GPS and / or GPS-enabled sensors) to support mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1158 may be used, including, for example, and without limitation, a GPS unit that uses a USB port with an Ethernet-to-serial bridge (e.g., RS-232).

[0150] In at least one embodiment, the vehicle 1100 may further include one or more radar sensors 1160. In at least one embodiment, one or more radar sensors 1160 of the vehicle 1100 may be used for long-range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, the functional safety levels of the radar may be ASIL B. In at least one embodiment, one or more radar sensors 1160 may use a CAN bus and / or bus 1102 (e.g., for transmitting the data generated by one or more radar sensors 1160) for control and access to object tracking data, with access to Ethernet channels to access raw data in some examples. In at least one embodiment, a wide variety of radar sensor types may be used.One or more RADAR sensors 1160 can be suitable, for example and without limitation, for front, rear, and side radar applications. In at least one embodiment, one or more of the RADAR sensors 1160 are pulse-Doppler radar sensors.

[0151] In at least one embodiment, one or more RADAR sensors 1160 can incorporate various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, side coverage with short-range, etc. In at least one embodiment, long-range RADAR can be used for the adaptive cruise control function. In at least one embodiment, long-range RADAR systems can provide a wide field of view, achieved by two or more independent scans, such as within a range of 250 m (meters). In at least one embodiment, one or more RADAR sensors 1160 can assist in distinguishing between stationary and moving objects and can be used by the ADAS system 1138 for emergency braking assistance and forward collision warning.In at least one embodiment, one or more sensors 1160 in a long-range, unrestricted radar system comprise a monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface. In at least one embodiment with six antennas, the central four antennas can generate a focused beam pattern intended to detect the surroundings of the vehicle 1100 at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, two other antennas can extend the field of view so that vehicles entering or leaving the lane of the vehicle 1100 can be detected quickly.

[0152] In at least one embodiment, medium-range RADAR systems can, for example, include a range of up to 160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, short-range RADAR systems can, without limitation, include any number of RADAR sensors 1160 designed for installation at both ends of a rear bumper. When installed at both ends of a rear bumper, in at least one embodiment such a RADAR sensor system can generate two beams that continuously monitor blind spots in a reversing direction and alongside a vehicle. In at least one embodiment, short-range RADAR systems can be used in an ADAS system 1138 for blind spot detection and / or as a lane change assistant.

[0153] In at least one embodiment, the vehicle 1100 can further include one or more ultrasonic sensors 1162. In at least one embodiment, one or more ultrasonic sensors 1162, which can be positioned on the front, rear, and / or sides of the vehicle 1100, can be used for parking assistance and / or for creating and updating an occupancy grid. In at least one embodiment, a plurality of ultrasonic sensors 1162 can be used, and different ultrasonic sensors 1162 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, one or more ultrasonic sensors 1162 can operate with functional safety levels of ASIL B.

[0154] In at least one embodiment, the vehicle 1100 can further include one or more LIDAR sensors 1164. In at least one embodiment, one or more LIDAR sensors 1164 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 1164 can be operated at functional safety level ASIL B. In at least one embodiment, the vehicle 1100 can include multiple LIDAR sensors 1164 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).

[0155] In at least one embodiment, one or more LIDAR sensors 1164 can provide a list of objects and their distances for a 360-degree field of view. In at least one embodiment, one or more commercially available LIDAR sensors 1164 can have a reported range of approximately 100 m, with an accuracy of 2 cm to 3 cm, and support for a 100 Mbit / s Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such an embodiment, one or more LIDAR sensors 1164 can include a small device that can be embedded in a front, rear, side, and / or corner location of the vehicle 1100.In at least one embodiment, one or more LIDAR sensors 1164 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even for objects with low reflectivity. In at least one embodiment, one or more front-mounted LIDAR sensors 1164 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0156] In at least one embodiment, LIDAR technologies, such as 3D flash LIDAR, can also be used. In at least one embodiment, the 3D flash LIDAR uses a laser pulse as a transmission source to illuminate the area around the vehicle 1100 up to approximately 200 m. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records the laser pulse time of flight and the reflected light at each pixel, corresponding to an area from the vehicle 1100 to objects. In at least one embodiment, the flash LIDAR can generate highly accurate and distortion-free images of the environment with each laser pulse. In at least one embodiment, four flash LIDAR sensors can be used, one on each side of the vehicle 1100.In at least one embodiment, 3D flash lidar systems without limitation include a solid-state 3D focal plane array lidar camera that contains no moving parts other than a fan (e.g., a non-scanning lidar device). In at least one embodiment, a flash lidar device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture reflected laser light as a 3D distance point cloud and co-registered intensity data.

[0157] In at least one embodiment, the vehicle 1100 may further include one or more IMU sensors 1166. In at least one embodiment, one or more IMU sensors 1166 may be located in the center of a rear axle of the vehicle 1100. In at least one embodiment, one or more IMU sensors 1166 may include, for example, and without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, such as in six-axis applications, one or more IMU sensors 1166 may include accelerometers and gyroscopes. In at least one embodiment, such as in nine-axis applications, one or more IMU sensors 1166 may include, without limitation, accelerometers, gyroscopes, and magnetometers.

[0158] In at least one embodiment, one or more IMU sensors 1166 can be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, speed, and orientation. In at least one embodiment, one or more IMU sensors 1166 enable a vehicle 1100 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating speed changes from a GPS with one or more IMU sensors 1166. In at least one embodiment, one or more IMU sensors 1166 and one or more GNSS sensors 1158 can be combined in a single integrated unit.

[0159] In at least one embodiment, the vehicle 1100 can include one or more microphones 1196 arranged in and / or around the vehicle 1100. In at least one embodiment, among other things, one or more microphones 1196 can be used for detecting and identifying emergency vehicles.

[0160] In at least one embodiment, the vehicle 1100 can further include any number of camera types, including one or more stereo cameras 1168, one or more wide-angle cameras 1170, one or more infrared cameras 1172, one or more surround-view cameras 1174, one or more long-range cameras 1198, medium-range cameras 1176, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around the entire periphery of the vehicle 1100. In at least one embodiment, the types of cameras used depend on the vehicle 1100. In at least one embodiment, any combination of camera types can be used to provide the necessary coverage around the vehicle 1100. In at least one embodiment, the number of cameras used can vary depending on the embodiment.In at least one embodiment, vehicle 1100 could, for example, include six cameras, seven cameras, ten cameras, twelve cameras, or any other number of cameras. In at least one embodiment, the cameras could, by way of example and without limitation, support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communications. In at least one embodiment, each camera could, as above with reference to… Fig. 11A and Fig. 11B is described in more detail.

[0161] In at least one embodiment, the vehicle 1100 may further include one or more vibration sensors 1142. In at least one embodiment, one or more vibration sensors 1142 may measure vibrations of components of the vehicle 1100, such as one or more axles. For example, in at least one embodiment, changes in vibrations may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 1142 are used, differences between vibrations may be used to determine friction or slip on the road surface (e.g., when there is a difference in vibration between a driven axle and a freely rotating axle).

[0162] In at least one embodiment, the vehicle 1100 can include the ADAS system 1138. In at least one embodiment, the ADAS system 1138 can include a SoC in some examples without restriction.In at least one embodiment, the ADAS system 1138 may include a number and combinations of an autonomous / adaptive / automatic cruise control (“ACC”) system, a cooperative adaptive cruise control (“CACC”) system, a forward crash warning (“FCW”) system, an automatic emergency braking (“AEB”) system, a lane departure warning (“LDW”) system, a lane keep assist (“LKA”) system, a blind spot warning (“BSW”) system, a rear cross-traffic warning (“RCTW”) system, a collision warning (“CW”) system, a lane centering (“LC”) system and / or other systems, features and / or functionalities.

[0163] In at least one embodiment, the ACC system can use one or more radar sensors 1160, one or more lidar sensors 1164, and / or any number of cameras. In at least one embodiment, the ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to another vehicle immediately in front of the vehicle 1100 and automatically adjusts the speed of the vehicle 1100 to maintain a safe distance from vehicles ahead. In at least one embodiment, a lateral ACC system performs distance control and advises the vehicle 1100 to change lanes if necessary. In at least one embodiment, a lateral ACC system interacts with other ADAS applications, such as LC and CW.

[0164] In at least one embodiment, a CACC system uses information from other vehicles, which can be received via a network interface 1124 and / or one or more wireless antennas 1126 from other vehicles via a wireless connection, or indirectly, via a network connection (e.g., via the Internet). In at least one embodiment, direct connections can be provided by a vehicle-to-vehicle (V2V) communication link, while indirect connections can be provided by an infrastructure-to-vehicle (I2V) communication link. In general, V2V communication provides information about the vehicles immediately ahead (e.g., vehicles that are directly in front of the vehicle 1100 and in the same lane), while I2V communication provides information about the traffic further ahead.In at least one embodiment, a CACC system can include one or both I2V and V2V information sources. In at least one embodiment, a CACC system can be more reliable with regard to information about vehicles ahead of the vehicle 1100 and has the potential to improve the smoothness of traffic flow and reduce congestion on the road.

[0165] In at least one embodiment, an FCW system is designed to alert a driver to a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or one or more radar sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to provide driver feedback, such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an FCW system can provide a warning, such as an audible tone, a visual warning, a vibration, and / or a rapid braking pulse.

[0166] In at least one embodiment, an AEB system detects an impending head-on collision with another vehicle or object and can automatically apply the brakes if a driver does not take corrective action within a specific time or distance parameter. In at least one embodiment, an AEB system can use one or more forward-facing cameras and / or one or more radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, it can typically first alert a driver so that they can take corrective action to avoid a collision. If a driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the impact of a predicted collision.In at least one embodiment, an AEB system may include techniques such as dynamic brake support and / or braking for an impending accident.

[0167] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to warn a driver if the vehicle crosses the lane markings. In at least one embodiment, an LDW system is not activated if a driver indicates an intentional lane departure, such as by activating a turn signal. In at least one embodiment, an LDW system may use forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to provide feedback to the driver, such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system.In at least one embodiment, an LKA system provides steering inputs or brakes to correct vehicle 1100 when vehicle 1100 begins to leave the lane.

[0168] In at least one embodiment, a BSW system detects vehicles in the car's blind spot and warns the driver. In at least one embodiment, a BSW system can provide a visual, audible, and / or tactile warning signal to indicate that merging into or changing lanes is unsafe. In at least one embodiment, a BSW system can issue an additional warning when a driver uses a turn signal. In at least one embodiment, a BSW system can use one or more rear-facing cameras and / or one or more radar sensors 1160, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to provide feedback to the driver, such as a display, a speaker, and / or a vibrating component.

[0169] In at least one embodiment, an RCTW system can provide a visual, audible, and / or tactile notification when an object is detected outside the field of view of the reversing camera while the vehicle is reversing. In at least one embodiment, an RCTW system includes an AEB system to ensure that vehicle brakes are applied to avoid a collision. In at least one embodiment, an RCTW system can use one or more rear-facing radar sensors coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to provide feedback to the driver, such as a display, a speaker, and / or a vibrating component.

[0170] In at least one embodiment, conventional ADAS systems can produce false positive results, which, while annoying and distracting for a driver, are generally not catastrophic because conventional ADAS systems warn the driver and give them the opportunity to decide whether a safety problem actually exists and to act accordingly. In at least one embodiment, in the event of conflicting results, vehicle 1100 decides for itself whether to follow the result of a primary computer or a secondary computer (e.g., a first controller or a second controller of controllers 1136). In at least one embodiment, the ADAS system 1138 can, for example, be a backup and / or secondary computer that provides information about perception to a rationality module of the backup computer.In at least one embodiment, a backup computer rationality monitor can run redundant diverse software on hardware components to detect errors in perception and dynamic driving tasks. In at least one embodiment, outputs from the ADAS system 1138 can be provided to a monitoring MCU. In at least one embodiment, if outputs from a primary computer and outputs from a secondary computer conflict, a monitoring MCU determines how to reconcile the conflicts to ensure safe operation.

[0171] In at least one embodiment, a primary computer can be configured to provide a monitoring MCU with a confidence score indicating the primary computer's confidence in a chosen result. In at least one embodiment, if the confidence score exceeds a threshold, this monitoring MCU can follow the primary computer's instruction, regardless of whether the secondary computer provides a contradictory or inconsistent result. In at least one embodiment, if a confidence score does not reach the threshold and the primary and secondary computers display different results (e.g., a conflict), a monitoring MCU can mediate between the computers to determine a suitable result.

[0172] In at least one embodiment, a monitoring MCU can be configured to operate one or more neural networks trained and configured to determine, based at least in part on outputs from a primary computer and outputs from a secondary computer, the conditions under which the latter will trigger false alarms. In at least one embodiment, one or more neural networks in a monitoring MCU can learn when the output of a secondary computer can be trusted and when it cannot. For example, in at least one embodiment, if this secondary computer is a radar-based FCW system, one or more neural networks in this monitoring MCU can learn when the FCW system identifies metallic objects that do not actually pose a hazard, such as a drain grate or manhole cover, which triggers an alarm.In at least one embodiment, a neural network in a monitoring MCU can learn, when a secondary computer is a camera-based LDW system, to disregard an LDW when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In at least one embodiment, a monitoring MCU can include at least one DLA or GPU suitable for running one or more neural networks with allocated memory. In at least one embodiment, a monitoring MCU can include and / or be contained as a component of one or more SoCs 1104.

[0173] In at least one embodiment, the ADAS system 1138 can include a secondary computer that executes the ADAS functionality using classical computer vision rules. In at least one embodiment, this secondary computer can use classical computer vision rules (if-then), and the presence of one or more neural networks in a monitoring MCU can improve reliability, safety, and performance. In at least one embodiment, for example, diverse implementation and intentional non-identity make the overall system more fault-tolerant, particularly with respect to errors caused by a function of the software (or software-hardware interfaces).In at least one embodiment, for example, if a software bug or error occurs in the software on a primary computer and non-identical software code on a secondary computer provides a consistent overall result, a monitoring MCU can have greater confidence that an overall result is correct and that a bug in the software or hardware on a primary computer does not cause a material error.

[0174] In at least one embodiment, an output from the ADAS system 1138 can be fed into a perception block of a primary computer and / or into a dynamic driving task block of a primary computer. For example, in at least one embodiment, if the ADAS system 1138 displays a frontal collision warning due to an object directly in front of the vehicle, a perception block can use this information in object identification. In at least one embodiment, a secondary computer can have its own trained neural network, thus reducing the risk of false positive results, as described herein.

[0175] In at least one embodiment, vehicle 1100 may further include infotainment SoC 1130 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, in at least one embodiment the infotainment system SoC 1130 may not be an SoC and may, without limitation, include two or more discrete components. In at least one embodiment, infotainment SoC 1130 may, without limitation, include a combination of hardware and software that can be used to provide vehicle 1100 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g.,Navigation systems, rear parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.) are provided. For example, the Infotainment SoC 1130 could include radios, turntables, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, a hands-free system, a head-up display (“HUD”), an HMI display 1134, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. In at least one embodiment, the Infotainment SoC 1130 can also be used to provide information (e.g.,to provide information (visually and / or acoustically) such as information from the ADAS system 1138, information on autonomous driving, such as planned vehicle maneuvers, road layouts, environmental information (e.g. intersection information, vehicle information, road information, etc.), and / or other information.

[0176] In at least one embodiment, the infotainment SoC 1130 can include any quantity and type of GPU functionality. In at least one embodiment, the infotainment SoC 1130 can communicate with other devices, systems, and / or components of the vehicle 1100 via the bus 1102. In at least one embodiment, the infotainment SoC 1130 can be coupled with a monitoring MCU so that a GPU of an infotainment system can perform some self-driving functions if one or more primary controllers 1136 (e.g., primary and / or backup computers of the vehicle 1100) fail. In at least one embodiment, the infotainment SoC 1130 can place the vehicle 1100 into a chauffeur-to-safe-stop mode, as described herein.

[0177] In at least one embodiment, vehicle 1100 may further include an instrument cluster 1132 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, instrument cluster 1132 may, without limitation, include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 1132 may, without limitation, include any number and any combination of a series of instruments, such as speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, gearshift position indicator, seatbelt warning light(s), parking brake warning light(s), engine malfunction light(s), supplemental restraint system information (e.g., airbag), lighting controls, safety system controls, navigation information, etc.In some examples, information from the infotainment SoC 1130 and the instrument cluster 1132 can be displayed and / or shared. In at least one embodiment, the instrument cluster 1132 can be included as part of the infotainment SoC 1130, or vice versa.

[0178] In at least one embodiment, an embodiment of at least one of Fig. 11C include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0179] Fig. 11D is a diagram of a system for communication between one or more cloud-based servers and the autonomous vehicle 1100. Fig. 11A according to at least one embodiment. In at least one embodiment, the system can include, without limitation, one or more servers 1178, one or more networks 1190, and any number and type of vehicles, including vehicle 1100. In at least one embodiment, one or more servers 1178 can include, without limitation, a plurality of GPUs 1184(A)-1184(H) (here collectively referred to as GPUs 1184), PCIe switches 1182(A)-1182(D) (here collectively referred to as PCIe switches 1182), and / or CPUs 1180(A)-1180(B) (here collectively referred to as CPUs 1180). In at least one embodiment, GPUs 1184, CPUs 1180 and PCIe switches 1182 can be interconnected using high-speed intermediate links, such as, but not limited to, the NVLink interfaces 1188 and / or PCIe links 1186 developed by NVIDIA.In at least one embodiment, GPUs 1184 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 1184 and the PCIe switches 1182 are connected via PCIe intermediate links. Although eight GPUs 1184, two CPUs 1180, and four PCIe switches 1182 are illustrated, this is not to be understood as a limitation. In at least one embodiment, each of the servers 1178 can contain any number of GPUs 1184, CPUs 1180, and / or PCIe switches 1182 in any combination without limitation. In at least one embodiment, for example, one or more servers 1178 could each contain eight, sixteen, thirty-two, and / or more GPUs 1184.

[0180] In at least one embodiment, one or more servers 1178 can receive image data from vehicles via one or more networks 1190. This data can be representative of images showing unexpected or altered road conditions, such as recently commenced roadworks. In at least one embodiment, one or more servers 1178 can transmit updated or otherwise updated map information 1194 to vehicles, neural networks 1192, including, without limitation, information about traffic and road conditions. In at least one embodiment, updates to the map information 1194 can include, without limitation, updates to the HD map 1122, such as information about construction sites, potholes, detours, floods, and / or other obstacles.In at least one embodiment, neural networks 1192 and / or map information 1194 can result from new training and / or experience represented in the data received from any number of vehicles in an environment, and / or be based at least partially on training performed in a data center (e.g. using one or more servers 1178 and / or other servers).

[0181] In at least one embodiment, one or more Server 1178 can be used to train machine learning models (e.g., neural networks) based at least partially on training data. In at least one embodiment, training data can be generated by the vehicles and / or generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged (e.g., if the associated neural network benefits from supervised learning) and / or subjected to other preprocessing. In at least one embodiment, any amount of training data is not tagged and / or preprocessed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once machine learning models are trained, machine learning models of vehicles can be used (e.g.,to vehicles via one or more networks 1190) and / or machine learning models can be used by one or more servers 1178 for remote monitoring of vehicles.

[0182] In at least one embodiment, one or more Server 1178 can receive data from vehicles and apply that data to current neural networks in real time for real-time intelligent inference. In at least one embodiment, one or more Server 1178 can include deep learning supercomputers and / or dedicated AI computers powered by one or more GPUs 1184, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more Server 1178 can include a deep learning infrastructure that utilizes CPU-driven data centers.

[0183] In at least one embodiment, the deep learning infrastructure of one or more servers 1178 can perform fast, real-time inference and can use this capability to evaluate and verify the state of processors, software, and / or associated hardware in the vehicle 1100. For example, in at least one embodiment, the deep learning infrastructure can receive periodic updates from the vehicle 1100, such as a sequence of images and / or objects that the vehicle 1100 has located within that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques).In at least one embodiment, the deep learning infrastructure can execute its own neural network to identify objects and compare them with objects identified by vehicle 1100, and if results do not match and the deep learning infrastructure concludes that the AI ​​in vehicle 1100 is not working correctly, one or more servers 1178 can send a signal to vehicle 1100 instructing a fail-safe computer of vehicle 1100 to take control, notify the passengers and perform a safe parking maneuver.

[0184] In at least one embodiment, one or more servers 1178 can include one or more GPUs 1184 and one or more programmable inference accelerators (e.g., NVIDIA TensorRT-3 devices). In at least one embodiment, a combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In at least one embodiment, for example, when performance is less critical, servers driven by CPUs, FPGAs, and other processors can be used for inference. In at least one embodiment, one or more hardware structures 815 are used to perform one or more embodiments. Details relating to one or more hardware structures 815 are set forth herein in conjunction with Fig. 8A and / or 8B provided. COMPUTER SYSTEMS

[0185] Fig. Figure 12 is a block diagram illustrating an exemplary computer system, which may be a system with interconnected devices and components, a system-on-a-chip (SoC), or any other combination thereof, formed with a processor, which may include execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, a computer system 1200 may, without limitation, include a component, such as a processor 1202, to employ execution units, including logic for executing algorithms on process data, according to the present disclosure, as in the embodiment described herein.In at least one embodiment, the Computer System 1200 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors, available from Intel Corporation, Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes and the like) may also be used. In at least one embodiment, the Computer System 1200 may run a version of a WINDOWS operating system, available from Microsoft Corporation, Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software and / or graphical user interfaces may also be used.

[0186] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip, network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0187] In at least one embodiment, the computer system 1200 can, without limitation, include the processor 1202, which can, without limitation, include one or more execution units 1208 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 1200 is a single-processor desktop or server system, but in another embodiment, the computer system 1200 can be a multiprocessor system. In at least one embodiment, the processor 1202 can, without limitation, include a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing a combination of instruction sets, or another processing device, such as a digital signal processor.In at least one embodiment, the processor 1202 can be coupled to a processor bus 1210, which can transmit data signals between the processor 1202 and other components in the computer system 1200.

[0188] In at least one embodiment, the processor 1202 can include, without limitation, an internal Level 1 ("L1") cache (“cache”) 1204. In at least one embodiment, the processor 1202 can have a single internal cache or multiple levels of an internal cache. In at least one embodiment, the cache can be located external to the processor 1202. Other embodiments can also include, depending on the specific implementation and requirements, a combination of both internal and external caches. In at least one embodiment, a register bank 1206 can store different types of data in different registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.

[0189] In at least one embodiment, the execution unit 1208, which includes without limitation logic for performing integer and floating-point operations, is also located in processor 1202. In at least one embodiment, processor 1202 may also include a microcode ("ucode") read-only memory ("ROM") that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 1208 may include logic for handling a packed instruction set 1209. In one embodiment, by including a packed instruction set 1209 in an instruction set of a general-purpose processor, together with associated switching technology for executing instructions, operations used by many multimedia applications can be performed using packed data in processor 1202.In at least one embodiment, many multimedia applications can be accelerated and run more efficiently by using a full width of a processor data bus to perform operations on packed data, thereby eliminating the need to transfer smaller units of data over the processor data bus to perform one or more operations on a single data element each.

[0190] In at least one embodiment, the execution unit 1208 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 1200 can include a working memory 1220 without restriction. In at least one embodiment, the working memory 1220 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or another type of working memory device. In at least one embodiment, the working memory 1220 can store one or more instructions 1219 and / or data 1221, represented by data signals that can be executed by the processor 1202.

[0191] In at least one embodiment, a system logic chip can be coupled to the processor bus 1210 and the main memory 1220. In at least one embodiment, a system logic chip can, without restriction, include a memory controller hub (“MCH”) 1216, and the processor 1202 can communicate with the MCH 1216 via the processor bus 1210. In at least one embodiment, the MCH 1216 can provide a high-bandwidth main memory path 1218 to the main memory 1220 for instruction and data storage and for storing graphics instructions, data, and textures. In at least one embodiment, the MCH 1216 can route data signals between the processor 1202, the main memory 1220, and other components in the computer system 1200, and for bridging data signals between the processor bus 1210, the main memory 1220, and a system I / O interface 1222.In at least one embodiment, a system logic chip can provide a graphics port for coupling with a graphics controller. In at least one embodiment, the MCH 1216 can be coupled to the main memory 1220 via the high-bandwidth main memory path 1218, and the graphics / video card 1212 can be coupled to the MCH 1216 via an Accelerated Graphics Port (“AGP”) intermediate connection 1214.

[0192] In at least one embodiment, the computer system 1200 can use the system I / O interface 1222 as a proprietary hub interface bus to couple the MCH 1216 to an I / O controller hub (“ICH”) 1230. In at least one embodiment, the ICH 1230 can provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus can, without limitation, include a high-speed I / O bus for connecting peripheral devices to the main memory 1220, a chipset, and the processor 1202. Examples can include, without limitation, an audio controller 1229, a firmware hub (“Flash BIOS”) 1228, a wireless transceiver 1226, a data storage device 1224, an old I / O controller 1223 containing user input and keyboard interfaces 1225, a serial expansion port 1227 such as Universal Serial Bus (“USB”), and a network controller 1234.In at least one embodiment, the data storage device 1224 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device or other mass storage device.

[0193] Illustrated in at least one embodiment Fig. 12 a system comprising interconnected hardware devices or “chips”, whereas in other embodiments Fig. 12 possibly illustrates an exemplary SoC. In at least one embodiment, in Fig. Figure 12 illustrates devices that can be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of the 1200 computer system are interconnected using Compute Express Link (CXL) interconnects.

[0194] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the computer system 1200 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0195] In at least one embodiment, an embodiment of at least one of the Fig. 11C and / or 12 contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0196] Fig. Figure 13 is a block diagram illustrating an electronic device 1300 for using a processor 1310, according to at least one embodiment. In at least one embodiment, the electronic device 1300 can be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device.

[0197] In at least one embodiment, the electronic device 1300 can, without limitation, include the processor 1310, which is communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1310 is coupled using a bus or interface, such as an I2C bus, a system management bus (“SMBus”), a low-pin-count (LPC) bus, a serial peripheral interface (“SPI”), a high-definition audio (“HDA”) bus, a serial advance technology attachment (“SATA”) bus, a universal serial bus (“USB”) (versions 1, 2, 3, etc.), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, the following is illustrated: Fig. 13 a system that includes interconnected hardware devices or “chips”, while Fig. 13 may illustrate an exemplary SoC in other embodiments. In at least one embodiment, in Fig. 13 illustrated devices are connected to each other using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 13 interconnected using Compute Express Link (CXL) intermediaries.

[0198] In at least one embodiment, Fig. 13 a display 1324, a touchscreen 1325, a touchpad 1330, a near field communication (NFC) unit 1345, a sensor hub 1340, a thermal sensor 1346, an Express chipset (EC) 1335, a Trusted Platform Module (TPM) 1338, a BIOS / firmware / flash memory (BIOS, FW Flash) 1322, a DSP 1360, a drive 1320, such as a solid-state drive (SSD) or a hard disk drive (HDD), a wireless local area network (WLAN) unit 1350, a Bluetooth unit 1352, a wireless wide area network (WWAN) unit 1356, a global positioning system (GPS) unit 1355, a camera (USB 3.0 camera) 1354, such as a The system may include a USB 3.0 camera and / or a low-power double data rate (LPDDR) memory unit (LPDDR3), for example, implemented in an LPDDR3 standard. These components can each be implemented in any suitable manner.

[0199] In at least one embodiment, other components can be communicatively coupled to the processor 1310 via the components described herein. In at least one embodiment, an accelerometer 1341, an ambient light sensor (ALS) 1342, a compass 1343, and a gyroscope 1344 can be communicatively coupled to the sensor hub 1340. In at least one embodiment, a thermal sensor 1339, a fan 1337, a keyboard 1336, and a touchpad 1330 can be communicatively coupled to the EC 1335. In at least one embodiment, loudspeakers 1363, headphones 1364, and a microphone (“Mic”) 1365 can be communicatively coupled to an audio unit (“audio codec and Class-D amplifier”) 1362, which in turn can be communicatively coupled to the DSP 1360. In at least one embodiment, the audio unit 1362 can, for example and without limitation, include an audio encoder / decoder (“codec”) and a Class-D amplifier.In at least one embodiment, a SIM card (“SIM”) 1357 can be communicatively coupled with the WWAN unit 1356. In at least one embodiment, components such as the WLAN unit 1350 and Bluetooth unit 1352, as well as the WWAN unit 1356, can be implemented in a next-generation form factor (“NGFF”).

[0200] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the electronic device 1300 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0201] In at least one embodiment, an embodiment of at least one of Fig. 13 include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0202] Fig. Figure 14 illustrates a computer system 1400 according to at least one embodiment. In at least one embodiment, the computer system 1400 is configured to implement various processes and procedures described in this disclosure.

[0203] In at least one embodiment, the computer system 1400 comprises, without limitation, at least one central processing unit (“CPU”) 1402 connected to a communication bus 1410 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), Peripheral Component Interconnect Express (“PCI Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 1400 also includes, without limitation, a main memory 1404 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in the main memory 1404, which may be in the form of random-access memory (“RAM”).In at least one embodiment, a network interface subsystem (“network interface”) 1422 provides an interface to other computing devices and networks to receive data from other systems with the computer system 1400 and to transmit data to them.

[0204] In at least one embodiment, the computer system 1400 includes, without limitation, input devices 1408, a parallel processing system 1412, and display devices 1406, which may be implemented using a conventional cathode ray tube (CRT), a liquid crystal display (LCD), a light-emitting diode (LED) display, a plasma display, or other suitable display technologies. In at least one embodiment, user input is received from input devices 1408, such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein may be arranged on a single semiconductor platform to form a processing system.

[0205] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are given herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the computer system 1400 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or neural network architectures or neural network use cases as described herein.

[0206] In at least one embodiment, an embodiment of at least one of Fig. 14 include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0207] Fig. Figure 15 illustrates a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 includes, without limitation, a computer 1510 and a USB stick 1520. In at least one embodiment, the computer 1510 can include, without limitation, any number and any type of processors (not shown) and main memory (not shown). In at least one embodiment, the computer 1510 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.

[0208] In at least one embodiment, the USB stick 1520 includes, without limitation, a processing unit 1530, a USB interface 1540, and USB interface logic 1550. In at least one embodiment, the processing unit 1530 can be any system, device, or instruction-executing apparatus capable of executing instructions. In at least one embodiment, the processing unit 1530 can include, without limitation, any number and any type of processing cores (not shown). In at least one embodiment, the processing unit 1530 comprises an application-specific integrated circuit (“ASIC”) optimized for performing any set and any type of machine learning-related operations.For example, in at least one embodiment, the processing unit 1530 is a tensor processing unit (“TPC”) optimized for performing machine learning inference operations. For example, in at least one embodiment, the processing unit 1530 is a vision processing unit (“VPU”) optimized for performing machine vision and machine learning inference operations.

[0209] In at least one embodiment, the USB interface 1540 can be any type of USB connector or USB socket. In at least one embodiment, the USB interface 1540 is, for example, a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1540 is a USB 3.0 connector. In at least one embodiment, the USB interface logic 1550 can include any number and any type of logic that enables the processing unit 1530 to establish an interface with devices (e.g., a computer 1510) via the USB connector 1540.

[0210] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the computer system 1500 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0211] In at least one embodiment, an embodiment of at least one of Fig. 15 include or cause one or more processors, circuits, or systems to conduct inference or training data of neural networks based at least partially on identifying different types of inference or training data within the inference or training data of neural networks that are to be denoised separately, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0212] Fig. Figure 16A illustrates an exemplary architecture in which a plurality of GPUs 1610(1)-1610(N) are communicatively coupled to a plurality of multi-core processors 1605(1)-1605(M) via high-speed links 1640(1)-1640(N) (e.g., buses, point-to-point intermediate links, etc.). In at least one embodiment, the high-speed links 1640(1)-1640(N) support a communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. In at least one embodiment, various intermediate link protocols can be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, “N” and “M” represent positive integers whose values ​​may differ from figure to figure. In at least one embodiment, one or more GPUs in a plurality of GPUs 1610(1)-1610(N) include one or more graphics cores (also simply referred to as "cores") 1900, as in Fig. 19A and Fig. 19B disclosed. In at least one embodiment, one or more graphics cores 1900 can be designated as streaming multiprocessors (“SMs”), stream processors (“SPs”), stream processing units (“SPUs”), compute units (“CUs”), execution units (“EUs”) and / or slices, wherein a slice in this context can refer to a portion of processing resources in a processing unit (e.g. 16 cores, a ray tracing unit, a thread director or scheduler).

[0213] Additionally, and in at least one embodiment, two or more GPUs 1610 are interconnected via the high-speed links 1629(1)-1629(2), which can be implemented using similar or different protocols / links than those used for the high-speed links 1640(1)-1640(N). Likewise, two or more multi-core processors 1605 can be interconnected via a high-speed link 1628, which can be symmetric multi-processor buses (SMP buses) operating at 20 GB / s, 30 GB / s, 120 GB / s or higher. Alternatively, all communication between different in Fig. System components shown in Figure 16A are connected using similar protocols / links (e.g., via a common intermediate link fabric).

[0214] In at least one embodiment, each multi-core processor 1605 is communicatively coupled to a processor memory 1601(1)-1601(M) via memory interfaces 1626(1)-1626(M), and each GPU 1610(1)-1610(N) is communicatively coupled to the GPU memory 1620(1)-1620(N) via GPU memory interfaces 1650(1)-1650(N). In at least one embodiment, the memory interfaces 1626 and 1650 can use similar or different technologies for accessing the memory. For example, and without limitation, the processor memory 1601(1)-1601(M) and the GPU memory 1620 can be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g.The memory may be GDDR5, GDDR6) or high-bandwidth memory (HBM), and / or non-volatile memory, such as 3D XPoint or NanoRAM. In at least one embodiment, part of the processor memory 1601 may be volatile memory and another part non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0215] As described herein, different multi-core processors 1605 and GPUs 1610 can each be physically coupled to a specific memory 1601, 1620, and / or a unified memory architecture can be implemented in which a virtual system address space (also called the "effective address space") is distributed across different physical memory units. For example, the processor memory units 1601(1)-1601(M) can each comprise 64 GB of the system memory address space, and the GPU memory units 1620(1)-1620(N) can each comprise 32 GB of the system memory address space, resulting in a total of 256 GB of addressable memory if M=2 and N=4. Other values ​​for N and M are possible.

[0216] Fig. Figure 16B illustrates additional details for an intermediate connection between a multi-core processor 1607 and a graphics acceleration module 1646 according to an exemplary embodiment. In at least one embodiment, the graphics acceleration module 1646 can include one or more GPU chips integrated on a line card coupled to the processor 1607 via the high-speed link 1640 (e.g., a PCIe bus, NVLink, etc.). Alternatively, in at least one embodiment, the graphics acceleration module 1646 can be integrated on a package or chip with the processor 1607.

[0217] In at least one embodiment, the processor 1607 includes a plurality of cores 1660A-1660D (which may be referred to as "execution units"), each with a translation lookaside buffer (TLB) 1661A-1661D and one or more caches 1662A-1662D. In at least one embodiment, the cores 1660A-1660D may include various other components for executing instructions and processing data, which are not illustrated. In at least one embodiment, the caches 1662A-1662D may include Level 1 (L1) and Level 2 (L2) caches. Additionally, one or more shared caches 1656 may be included in the caches 1662A-1662D and shared by sets of cores 1660A-1660D. For example, one embodiment of the 1607 processor includes 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches.In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. In at least one embodiment, the processor 1607 and the graphics acceleration module 1646 are connected to the main memory 1614, which is the processor main memory 1601(1)-1601(M). Fig. May include 16A.

[0218] In at least one embodiment, coherence for data and instructions stored in various caches 1662A-1662D, 1656, and in main memory 1614 is maintained via communication between the cores over a coherence bus 1664. In at least one embodiment, for example, each cache can have its own associated cache coherence logic / circuit to communicate over the coherence bus 1664 in response to detected read or write operations in specific cache rows. In at least one embodiment, a cache snooping protocol is implemented over the coherence bus 1664 to monitor access to the cache.

[0219] In at least one embodiment, a proxy circuit 1625 communicatively couples the graphics acceleration module 1646 to the coherence bus 1664, enabling the graphics acceleration module 1646 to participate in a cache coherence protocol as an equal partner of the cores 1660A-1660D. In particular, in at least one embodiment, an interface 1635 provides connectivity to the proxy circuit 1625 via a high-speed link 1640, and an interface 1637 connects the graphics acceleration module 1646 to the high-speed link 1640.

[0220] In at least one embodiment, an accelerator integration circuit 1636 provides cache management, memory access, context management, and interrupt management services for a plurality of graphics processing units 1631(1)-1631(N) of the graphics acceleration module 1646. In at least one embodiment, each graphics processing unit 1631(1)-1631(N) may comprise a separate graphics processing unit (GPU). In at least one embodiment, a plurality of graphics processing units 1631(1)-1631(N) of the graphics acceleration module 1646 comprises one or more graphics cores 1900, as described in conjunction with Fig. 19A and Fig. 19B explained. In at least one embodiment, the graphics processing machines 1631(1)-1631(N) can alternatively comprise different types of graphics processing machines within a GPU, such as graphics execution units, media processing machines (e.g., video encoders / decoders), samplers, and blith machines. In at least one embodiment, the graphics acceleration module 1646 can be a GPU with a plurality of graphics processing machines 1631(1)-1631(N), or the graphics processing machines 1631(1)-1631(N) can be individual GPUs integrated in a common package, line card, or chip.

[0221] In at least one embodiment, the accelerator integration circuit 1636 includes a memory management unit (MMU) 1639 for performing various memory management functions, such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing the memory 1614. In at least one embodiment, the MMU 1639 may also include a translation buffer (TLB) (not shown) for caching translations from virtual / effective to physical / real addresses. In at least one embodiment, a cache 1638 may store instructions and data for efficient access by graphics processing machines 1631(1)-1631(N).In at least one embodiment, the data stored in the cache 1638 and the graphics memory 1633(1)-1633(M) are kept coherent with the core caches 1662A-1662D, 1656, and the system memory 1614, possibly using a pickup unit 1644. As mentioned earlier, this can be done via the proxy circuit 1625 on behalf of the cache 1638 and the memory 1633(1)-1633(M) (e.g., sending updates to the cache 1638 regarding modifications / accesses to cache lines in the processor caches 1662A-1662D and 1656, and receiving updates from the cache 1638).

[0222] In at least one embodiment, a set of registers 1645 stores context data for threads executed by the graphics processing machines 1631(1)-1631(N), and a context management circuit 1648 manages thread contexts. For example, the context management circuit 1648 can perform backup and restore operations to save and restore the contexts of different threads during context switches (e.g., when a first thread is saved and a second thread is stored so that a second thread can be executed by a graphics processing machine). For example, during a context switch, the context management circuit 1648 can store current register values ​​in a specific region of memory (e.g., identified by a context pointer). It can then restore the register values ​​when returning to a context.In at least one embodiment, an interruption management circuit 1647 receives and processes interruptions received from system devices.

[0223] In at least one embodiment, virtual / effective addresses from a graphics processing machine 1631 are translated into real / physical addresses in the system memory 1614 by the MMU 1639. In at least one embodiment, the accelerator integration circuit 1636 supports multiple (e.g., 4, 8, 16) graphics acceleration modules 1646 and / or other accelerator devices. In at least one embodiment, the graphics acceleration module 1646 can be dedicated to a single application running on the processor 1607 or shared by multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources from graphics processing machines 1631(1)-1631(N) are shared with multiple applications or virtual machines (VMs).In at least one embodiment, resources can be divided into “slices” that are assigned to different VMs and / or applications based on processing requirements and priorities assigned to those VMs and / or applications.

[0224] In at least one embodiment, the accelerator integration circuit 1636 acts as a bridge to a system for the graphics acceleration module 1646 and provides address translation and system memory cache services. Additionally, in at least one embodiment, the accelerator integration circuit 1636 can provide virtualization facilities for a host processor to manage the virtualization of graphics processing machines 1631(1)-1631(N), aborts, and memory management.

[0225] In at least one embodiment, each host processor can directly address these resources using an effective address value, since the hardware resources of the graphics processing machines 1631(1)-1631(N) are explicitly allocated to a real address space seen by the host processor 1607. In at least one embodiment, a function of the accelerator integration circuit 1636 is to physically separate the graphics processing machines 1631(1)-1631(N) so that they appear to a system as independent units.

[0226] In at least one embodiment, one or more graphics memory units 1633(1)-1633(M) are coupled to each of the graphics processing machines 1631(1)-1631(N), and N=M. In at least one embodiment, the graphics memory units 1633(1)-1633(M) store instructions and data that are processed by each graphics processing machine 1631(1)-1631(N). In at least one embodiment, the graphics memory units 1633(1)-1633(M) can be volatile memory, such as DRAMs (including stacked DRAMs), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memory, such as 3D XPoint or NanoRAM.

[0227] In at least one embodiment, biasing techniques can be used to reduce data traffic over the high-speed link 1640 and to ensure that the data stored in the graphics memory 1633(1)-1633(M) is data most frequently used by graphics processing units 1631(1)-1631(N) and preferably not (or at least not frequently) used by cores 1660A-1660D. Likewise, in at least one embodiment, a biasing mechanism attempts to keep data required by cores (and preferably not by graphics processing units 1631(1)-1631(N)) in caches 1662A-1662D, 1656, and the main memory 1614.

[0228] Fig. Figure 16C illustrates another exemplary embodiment in which the accelerator integration circuit 1636 is integrated within the processor 1607. In this embodiment, the graphics processing machines 1631(1)-1631(N) communicate directly with the accelerator integration circuit 1636 via the high-speed link 1640, using interface 1637 and interface 1635 (which can be any form of bus or interface protocol). In at least one embodiment, the accelerator integration circuit 1636 can perform operations similar to those described in Figure 16C. Fig. 16B are described, but possibly with a higher throughput, since it is located in close proximity to the coherence bus 1664 and the caches 1662A-1662D, 1656. In at least one embodiment, an accelerator integration circuit supports different programming models, including a dedicated process programming model (no virtualization of the graphics acceleration module) and shared programming models (with virtualization), which may include programming models controlled by the accelerator integration circuit 1636 and programming models controlled by the graphics acceleration module 1646.

[0229] In at least one embodiment, the graphics processing machines 1631(1)-1631(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can forward other application requests to graphics processing machines 1631(1)-1631(N), thereby providing virtualization within a VM / partition.

[0230] In at least one embodiment, graphics processing machines 1631(1)-1631(N) can be shared by multiple VM / application partitions. In at least one embodiment, shared models can use a system hypervisor to virtualize graphics processing machines 1631(1)-1631(N) to allow access by any operating system. In at least one embodiment, in single-partition systems without a hypervisor, the graphics processing machines 1631(1)-1631(N) belong to an operating system. In at least one embodiment, an operating system can virtualize graphics processing machines 1631(1)-1631(N) to provide access to any process or application.

[0231] In at least one embodiment, the graphics acceleration module 1646 or an individual graphics processing machine 1631(1)-1631(N) selects a process element using a process handle. In at least one embodiment, process elements are stored in memory 1614 and are addressable using a technique described herein for translating effective addresses into real addresses. In at least one embodiment, a process handle can be an implementation-specific value provided to a host process when its context is registered with the graphics processing machine 1631(1)-1631(N) (i.e., by calling system software to add a process element to a process element link list). In at least one embodiment, the lower 16 bits of a process handle can be an offset of a process element within a process element link list.

[0232] Fig. Figure 16D illustrates an exemplary accelerator integration slice 1690. In at least one embodiment, a "slice" comprises a specified portion of the processing resources of the accelerator integration circuit 1636. In at least one embodiment, an application is an effective address space 1682 within the main memory 1614 in which process elements 1683 are stored. In at least one embodiment, process elements 1683 are stored in response to GPU calls 1681 from applications 1680 running on the processor 1607. In at least one embodiment, a process element 1683 contains a process state for the corresponding application 1680. In at least one embodiment, a work descriptor (WD) 1684 contained in a process element 1683 can be a single job requested by an application or a pointer to a queue of jobs.In at least one embodiment, the WD 1684 is a pointer to a queue for job requests in an effective address space 1682 of an application.

[0233] In at least one embodiment, the graphics acceleration module 1646 and / or individual graphics processing machines 1631(1)-1631(N) can be shared by all or a subset of processes in a system. In at least one embodiment, an infrastructure for setting process states and sending a WD 1684 to a graphics acceleration module 1646 to start a job in a virtualized environment can be included.

[0234] In at least one embodiment, a dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single graphics acceleration module 1646 or a single graphics processing machine 1631 belongs to a single process. In at least one embodiment, a hypervisor initializes the accelerator integration circuit 1636 for an owning partition, and an operating system initializes the accelerator integration circuit 1636 for an owning process when the graphics acceleration module 1646 is assigned to a single process.

[0235] In at least one embodiment, a WD picker 1691 retrieves the next WD 1684 from the accelerator integration slice 1690 during operation. This WD 1684 contains an indication of the work to be performed by one or more graphics processing units of the graphics acceleration module 1646. In at least one embodiment, data from the WD 1684 can be stored in the registers 1645 and used by the MMU 1639, the interrupt management circuit 1647, and / or the context management circuit 1648, as illustrated. For example, one embodiment of the MMU 1639 includes segment / page scrolling circuits for accessing segment / page tables 1686 within a virtual address space 1685 of the OS. In at least one embodiment, the interrupt management circuit 1647 can process interrupt events 1692 received from the graphics acceleration module 1646.In at least one embodiment, when performing graphics operations, an effective address 1693, generated by a graphics processing machine 1631(1)-1631(N), is translated by the MMU 1639 into a real address.

[0236] In at least one embodiment, registers 1645 are duplicated for each graphics processing engine 1631(1)-1631(N) and / or each graphics acceleration module 1646 and can be initialized by a hypervisor or an operating system. In at least one embodiment, each of these duplicated registers can be included in an accelerator integration slice 1690. Exemplary registers that can be initialized by a hypervisor are shown in Table 1. Table 1 - Registers initialized by the hypervisor Register-Nr. Beschreibung 1 Slice-Steuerungsregister 2 Reale Adresse (RA) des Bereichszeigers geplanter Prozesse 3 Autoritätsmasken-Überschreibungsregister 4 Unterbrechungsvektor-Tabelleneintrags-Offset 5 Unterbrechungsvektor-Tabelleneintragslimit 6 Statusregister 7 ID der logischen Partition 8 Reale Adresse (RA) des Hypervisor-Beschleuniger-Nutzungsaufzeichnungs-Zeigers 9 Speicherungsbeschreibungsregister

[0237] Examples of registers that can be initialized by an operating system are shown in Table 2. Table 2 - Registers initialized by the operating system Register-Nr. Description 1 Process and thread identification 2 Effective Address (EA) of the Context Backup / Restore Pointer 3 Virtual Address (VA) of the Accelerator Usage Record Pointer 4 Virtual address (VA) of the memory segment table pointer 5 Authority mask 6 Work descriptor

[0238] In at least one embodiment, each WD 1684 is specific to a particular graphics acceleration module 1646 and / or graphics processing machines 1631(1)-1631(N). In at least one embodiment, it contains all the information that a graphics processing machine 1631(1)-1631(N) needs to perform its work, or it can be a pointer to a memory location where an application has set up an instruction queue containing the work to be performed.

[0239] Fig.Figure 16E illustrates additional details for an exemplary embodiment of a split model. This embodiment includes a real hypervisor address space 1698 in which a process element list 1699 is stored. In at least one embodiment, the real hypervisor address space 1698 is accessible via a hypervisor 1696 that virtualizes graphics acceleration module machines for the operating system 1695.

[0240] In at least one embodiment, shared programming models allow all or a subset of processes from all or a subset of partitions in a system to use a graphics acceleration module 1646. In at least one embodiment, there are two programming models in which the graphics acceleration module 1646 is shared by multiple processes and partitions, namely time-sliced ​​and graphically directed shared use.

[0241] In at least one embodiment, the system hypervisor 1696 in this model includes the graphics acceleration module 1646 and makes its functionality available to all operating systems 1695. In at least one embodiment, a graphics acceleration module 1646 can meet certain requirements so that the graphics acceleration module 1646 supports virtualization by the system hypervisor 1696, such as (1) an application job request must be autonomous (i.e.,(1) the state does not need to be maintained between jobs), or the Graphics Acceleration Module 1646 must provide a mechanism for saving and restoring the context, (2) an application job request is guaranteed by the Graphics Acceleration Module 1646 to be complete within a specified time period, including any translation errors, or the Graphics Acceleration Module 1646 provides the ability to displace the processing of a job, and (3) the Graphics Acceleration Module 1646 must be guaranteed fairness between processes when operating in a directed shared programming model.

[0242] In at least one embodiment, application 1680 is required to perform a system call of operating system 1695 with a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes a targeted acceleration function for a system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value.In at least one embodiment, WD is formatted specifically for the 1646 graphics acceleration module and can be in the form of a 1646 graphics acceleration module instruction, an effective address pointer to a user-defined structure, an effective address pointer to a queue of instructions, or any other data structure to describe the work to be performed by the 1646 graphics acceleration module.

[0243] In at least one embodiment, an AMR value is an AMR state to be used for a current process. In at least one embodiment, a value passed to an operating system resembles an application that sets an AMR. In at least one embodiment, if the accelerator integration circuit 1636 (not shown) and the implementations of the graphics acceleration module 1646 do not support a user authority mask override register (UAMOR), an operating system can apply a current UAMOR value to an AMR value before passing an AMR in a hypervisor call. In at least one embodiment, the hypervisor 1696 can optionally apply a current authority mask override register (AMOR) value before placing an AMR into the process element 1683.In at least one embodiment, the CSRP is one of the registers 1645, which is an effective address of a range in the effective address space 1682. an application contains a pointer so that the 1646 graphics acceleration module can save and restore the context state. In at least one embodiment, this pointer is optional if no state needs to be saved between jobs or if a job is overridden. In at least one embodiment, the context save / restore area can be memory with locked pages.

[0244] Upon receiving a system call, the operating system 1695 can check whether the application 1680 is registered and has been granted permission to use the graphics acceleration module 1646. In at least one embodiment, the operating system 1695 then calls the hypervisor 1696 with the information shown in Table 3. Table 3 - Operating system call parameters to the hypervisor Parameter No. Description 1 A working descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked) 3 An effective address (EA) of the context backup / restore area pointer (CSRP) 4 A process ID (PID) and an optional thread ID (TID) 5 A virtual address (VA) of the accelerator usage record pointer (AURP) 6 Virtual address of the memory segment table pointer (SSTP) 7 A Logical Interruption Service Number (LISN)

[0245] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 1696 checks whether the operating system 1695 is registered and has been granted permission to use the graphics acceleration module 1646. In at least one embodiment, the hypervisor 1696 then places the process element 1683 in a process element link list for a corresponding type of graphics acceleration module 1646. In at least one embodiment, a process element can contain information shown in Table 4. Table 4 - Process element information Element No. Description 1 A working descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked). 3 An effective address (EA) of the context backup / restore area pointer (CSRP) 4 A process ID (PID) and an optional thread ID (TID) 5 A virtual address (VA) of the accelerator usage record pointer (AURP) 6 Virtual address of the memory segment table pointer (SSTP) 7 A Logical Interruption Service Number (LISN) 8 Interruption vector table derived from the hypervisor's call parameters 9 A status register (SR) value 10 A logical partition ID (LPID) 11 Real Address (RA) of the Hypervisor Accelerator Usage Record Pointer 12 Storage Descriptor Register (SDR)

[0246] In at least one embodiment, the hypervisor initializes a plurality of registers 1645 of the accelerator integration slice 1690.

[0247] As in Fig.As illustrated in Figure 16F, in at least one embodiment a unified memory is used, which is addressable via a common virtual memory address space used for accessing physical processor memory 1601(1)-1601(N) and GPU memory 1620(1)-1620(N). In this implementation, operations performed on the GPUs 1610(1)-1610(N) use the same virtual / effective memory address space to access the processor memory 1601(1)-1601(M) and vice versa, thus simplifying programmability. In at least one embodiment, a first part of a virtual / effective address space is allocated to the processor memory 1601(1), a second part to the second processor memory 1601(N), a third part to the GPU memory 1620(1), and so on.In at least one embodiment, this distributes an entire virtual / effective memory space (sometimes also referred to as effective address space) across each of the processor memory 1601 and GPU memory 1620, so that each processor or GPU can access each physical memory with a virtual address assigned to that memory.

[0248] In at least one embodiment, the bias / coherence management circuits 1694A-1694E in one or more MMUs 1639A-1639E ensure cache coherence between caches of one or more host processors (e.g., 1605) and GPUs 1610 and implement biasing techniques that specify physical memory locations where certain types of data should be stored. In at least one embodiment, although multiple instances of the bias / coherence management circuits 1694A-1694E are present in Fig.As illustrated in Figure 16F, bias / coherence circuits can be implemented within an MMU of one or more host processors 1605 and / or within the accelerator integration circuit 1636.

[0249] One embodiment allows GPU memory 1620 to be allocated as part of system memory and accessed using shared virtual memory (SVM) technology, without incurring the performance degradation associated with full system cache coherence. In at least one embodiment, the ability to access GPU memory 1620 as system memory without the cumbersome cache coherence overhead provides an advantageous operating environment for GPU offloading. In at least one embodiment, this arrangement allows the host processor 1605 software to set up operands and access computation results without the overhead of traditional I / O DMA data copies.In at least one embodiment, such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses. In at least one embodiment, the ability to access GPU memory 1620 without cache coherence overhead can be critical to the execution time of an offloaded computation. In at least one embodiment, for example, in cases with significant streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by a GPU 1610. In at least one embodiment, the efficiency of operand setup, the efficiency of accessing the results, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offloading.

[0250] In at least one embodiment, the selection of GPU bias and host processor bias is controlled by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granular structure (e.g., controlled with a memory page granularity) containing 1 or 2 bits per GPU-connected memory page. In at least one embodiment, a bias table can be implemented in a stolen memory area of ​​one or more GPU memory 1620s, with or without a bias cache in a GPU 1610 (e.g., for caching frequently / recently used entries of a bias table). Alternatively, in at least one embodiment, an entire bias table can be maintained in a GPU.

[0251] In at least one embodiment, prior to the actual access to a GPU memory, a bias table entry is accessed, which is associated with each access to a GPU-connected memory 1620, thereby triggering subsequent operations. In at least one embodiment, local requests from a GPU 1610 that find their page in the GPU bias are forwarded directly to a corresponding GPU memory 1620. In at least one embodiment, local requests from a GPU that find their page in the host bias are forwarded to a processor 1605 (e.g., via a high-speed link, as described herein). In at least one embodiment, requests from the processor 1605 that find a requested page in the host processor bias complete a request like a normal memory read.Alternatively, requests directed to a GPU-biased page can be forwarded to a GPU 1610. In at least one embodiment, a GPU can then convert a page into a host processor bias if it is not currently using that page. In at least one embodiment, a bias state of a page can be changed either by a software-based mechanism, a hardware-based software-based mechanism, or, in a limited number of cases, by a purely hardware-based mechanism.

[0252] In at least one embodiment, a bias-state-changing mechanism uses an API call (e.g., OpenCL) which in turn calls a GPU device driver. This driver then sends a message (or a queued instruction) to the GPU, instructing it to change a bias state and, on some transitions, to perform a cache-flushing operation on a host. In at least one embodiment, a cache-flushing operation is used for a transition from the host processor 1605 bias to the GPU bias, but not for the reverse transition.

[0253] In at least one embodiment, cache coherence is maintained by temporarily rendering GPU-biased pages uncacheable by the host processor 1605. In at least one embodiment, the processor 1605 can request access to these pages from the GPU 1610, which may or may not grant access immediately. Therefore, in at least one embodiment, to reduce communication between the processor 1605 and the GPU 1610, it is advantageous to ensure that GPU-biased pages are those requested by a GPU, but not by the host processor 1605, and vice versa.

[0254] One or more hardware structures 815 are used to carry out one or more embodiments. Details regarding one or more hardware structures 815 can be found herein in conjunction with Fig. 8A and / or 8B will be provided.

[0255] Fig.Figure 17 illustrates exemplary integrated circuits and associated graphics processing units (GPUs) that can be manufactured according to various embodiments described herein using one or more IP cores. In addition to what is illustrated, at least one embodiment may include other logic and circuitry, including additional GPUs / cores, peripheral interface controllers, or general-purpose processor cores.

[0256] Fig.Figure 17 is a block diagram illustrating an exemplary integrated system-on-a-chip circuit 1700, which can be manufactured according to at least one embodiment using one or more IP cores. In at least one embodiment, the integrated circuit 1700 includes one or more application processors 1705 (e.g., CPUs), at least one graphics processor 1710, and may additionally include an image processor 1715 and / or a video processor 1720, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 1700 includes peripheral device or bus logic comprising a USB controller 1725, a UART controller 1730, an SPI / SDIO controller 1735, and an I22S / I22C controller 1740.In at least one embodiment, the integrated circuit 1700 can include a display device 1745 coupled to one or more High-Definition Multimedia Interface (HDMI) controllers 1750 and / or Mobile Industry Processor Interface (MIPI) display interfaces 1755. In at least one embodiment, storage can be provided by a flash memory subsystem 1760, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface can be provided via a memory controller 1765 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security machine 1770.

[0257] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the integrated circuit 1700 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0258] In at least one embodiment, an embodiment of at least one of the Fig. 16A, Fig. 16B, Fig. 16C, Fig. 16D, Fig. 16E, Fig.16F and / or 17 contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to the Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0259] Fig.Figures 18A-18B illustrate exemplary integrated circuits and associated graphics processing units (GPUs) that can be manufactured according to various embodiments described herein using one or more IP cores. In addition to what is illustrated, at least one embodiment may include other logic and circuitry, including additional GPUs / cores, peripheral interface controllers, or general-purpose processor cores.

[0260] Fig. Figures 18A-18B are block diagrams illustrating exemplary graphics processors for use within a SoC, according to the embodiments described herein. Fig. Figure 18A illustrates an exemplary graphics processor 1810 of an integrated system-on-a-chip circuit which can be manufactured according to at least one embodiment using one or more IP cores. Fig.Figure 18B illustrates an additional exemplary graphics processor 1840 of an integrated system-on-a-chip circuit, which can be manufactured according to at least one embodiment using one or more IP cores. In at least one embodiment, the graphics processor 1810 is of Fig. 18A is a low-performance graphics processor core. In at least one embodiment, the graphics processor 1840 is of Fig. 18B is a higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 1810, 1840 can be a variant of the graphics processor 1710. Fig. Be 17.

[0261] In at least one embodiment, the graphics processor 1810 includes a vertex processor 1805 and one or more fragment processors 1815A-1815N (e.g., 1815A, 1815B, 1815C, 1815D to 1815N-1 and 1815N). In at least one embodiment, the graphics processor 1810 can execute different shader programs via separate logic, such that the vertex processor 1805 is optimized for executing operations for vertex shader programs, while one or more fragment processors 1815A-1815N execute fragment shading operations (e.g., pixel shading operations) for fragment or pixel shader programs. In at least one embodiment, the Vertex Processor 1805 performs a vertex processing stage of a 3D graphics pipeline and generates primitive and vertex data.In at least one embodiment, one or more fragment processors 1815A-1815N use the primitive and vertex data generated by the vertex processor 1805 to create a framebuffer that is displayed on a display device. In at least one embodiment, one or more processors 1815A-1815N are optimized for executing fragment shader programs, such as those provided in an OpenGL API, which can be used to perform operations similar to those of a pixel shader program, such as those provided in a Direct3D API.

[0262] In at least one embodiment, the graphics processor 1810 additionally includes one or more memory management units (MMUs) 1820A-1820B, cache(s) 1825A-1825B, and interconnects 1830A-1830B. In at least one embodiment, one or more MMUs 1820A-1820B provide the allocation of virtual to physical addresses for the graphics processor 1810, including for the vertex processor 1805 and / or one or more fragment processors 1815A-1815N, which may refer to vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1825A-1825B. In at least one embodiment, one or more MMU(s) 1820A-1820B can be synchronized with other MMUs within a system comprising one or more MMUs that are connected to one or more application processors 1705, image processors 1715 and / or video processors 1720. Fig. 17 are assigned such that each processor 1705-1720 can participate in a shared or unified virtual memory. In at least one embodiment, one or more interconnects 1830A-1830B enable the graphics processor 1810 to interface with other IP cores within the SoC, either via an internal bus of the SoC or via a direct connection.

[0263] In at least one embodiment, the 1840 graphics processor includes one or more shader cores 1855A-1855N (e.g., 1855A, 1855B, 1855C, 1855D, 1855E, 1855F to 1855N-1 and 1855N), as shown in Fig.Figure 18B shows a unified shader core architecture in which a single core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or computational shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1840 includes a task manager 1845 between the cores, which acts as a thread distributor to distribute execution threads to one or more shader cores 1855A-1855N, and a tiling unit 1858 to accelerate tiling operations for tile-based rendering, where rendering operations for a scene are subdivided in image space, for example, to take advantage of local spatial coherence within a scene or to optimize the use of internal caches.

[0264] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the graphics processor 1810 and / or 1840 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0265] In at least one embodiment, an embodiment of at least one of the Fig.18A and / or 18B contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to the Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0266] Fig. Figures 19A-19B illustrate additional exemplary graphics processor logic according to the embodiments described herein. In at least one embodiment, the logic is described in and in conjunction with Fig.19A-19B illustrated and described components integrated into a single system, such as a graphics processing unit (GPU), a SoC, or another type of processor. Fig. Figure 19A illustrates a graphics core 1900, which in at least one embodiment is located within the graphics processor 1710. Fig. 17 may be included and in at least one embodiment a unified shader core 1855A-1855N as in Fig. It could be 18B. Fig.Figure 19B illustrates a highly parallel general-purpose graphics processing unit (“GPGPU,” which may also be referred to as a “graphics processing unit”) 1930, which in at least one embodiment is suitable for use on a multi-chip module. In at least one embodiment, the graphics processing unit 1930 is a GPGPU comprising a graphics processor. In at least one embodiment, the integrated circuit 1700 comprises a graphics core 1900, for example, to form an integrated circuit and / or a system-on-a-chip (SoC), wherein such an integrated circuit and / or such an SoC performs the operations described herein.

[0267] In at least one embodiment, the graphics core 1900 includes a shared instruction cache 1902, a texture unit 1918, and a cache / shared memory 1920 (e.g., including L1, L2, L3, last-level cache, or other caches) that are common execution resources within the graphics core 1900. In at least one embodiment, the graphics core 1900 can include multiple slices 1901A-1901N or a partition for each core, and a graphics processor can include multiple instances of the graphics core 1900. In at least one embodiment, each slice 1901A-1901N relates to the graphics core 1900. In at least one embodiment, the slices 1901A-1901N have sub-slices that are parts of a slice 1901A-1901N. In at least one embodiment, the slices 1901A-1901N are independent of other slices or dependent on other slices.In at least one embodiment, the slices 1901A-1901N can include support logic comprising a local instruction cache 1904A-1904N, a thread scheduler (sequencer) 1906A-1906N, a thread distributor 1908A-1908N and a set of registers 1910A-1910N. In at least one embodiment, the slices 1901A-1901N can include a set of additional function units (AFUs 1912A-1912N), floating-point units (FPUs 1914A-1914N), arithmetic logic units (ALUs 1916A-1916N) for integers, address computational units (ACUs 1913A-1913N), double-precision floating-point units (DPFPUs 1915A-1915N), and matrix processing units (MPUs 1917A-1917N). In at least one embodiment, MPUs 1917A-1917N are referred to as matrix machines.

[0268] In at least one embodiment, each slice 1901A-1901N includes one or more machines for floating-point and integer vector operations and one or more machines for accelerating convolution and matrix operations in AI, machine learning, or large dataset workloads. In at least one embodiment, one or more slices 1901A-1901N include one or more vector machines for computing a vector (e.g., computing mathematical operations on vectors). In at least one embodiment, a vector machine can compute a vector operation in 16-bit floating-point (also known as "FP16"), 32-bit floating-point (also known as "FP32"), or 64-bit floating-point (also known as "FP64").In at least one embodiment, one or more slices 1901A-1901N comprise 16 vector machines paired with 16 matrix math units to compute matrix / tensor operations, wherein the vector machines and math units are exposed via matrix extensions. In at least one embodiment, a slice is a specified portion of the processing resources of a processing unit, e.g., 16 cores and a ray tracing unit, or 8 cores, a thread scheduler, a thread distributor, and additional functional units for a processor. In at least one embodiment, the graphics core 1900 comprises one or more matrix machines for computation of matrix operations, e.g., in the computation of tensor operations.

[0269] In at least one embodiment, one or more slices 1901A-1901N include one or more ray tracing units to compute ray tracing operations (e.g., 16 ray tracing units per slice of slices 1901A-1901N). In at least one embodiment, a ray tracing unit computes ray tracing, triangle crossing, bounding box crossing, or other ray tracing operations.

[0270] In at least one embodiment, one or more slices 1901A-1901N include a media slice that encodes, decodes and / or transcodes data; scales data and / or converts the format of the data; and / or performs video quality operations on video data.

[0271] In at least one embodiment, one or more slices 1901A-1901N are connected to L2 cache and memory fabric, link connectors, stacks of high-bandwidth memory (HBM) (e.g., HBM2e, HDM3), and a media machine. In at least one embodiment, one or more slices 1901A-1901N include multiple cores (e.g., 16 cores) and multiple ray tracing units (e.g., 16) paired with each core. In at least one embodiment, one or more slices 1901A-1901N have one or more L1 caches. In at least one embodiment, one or more slices 1901A-1901N include one or more vector machines; one or more instruction caches for storing instructions; and one or more L1 caches for caching data. one or more shared local memories (SLMs) for storing data, e.g.according to the instructions; one or more samplers to draw data as a sample; one or more ray tracing units to perform ray tracing operations; one or more geometries to perform operations in geometry pipelines and / or to apply geometric transformations at vertices or polygons; one or more rasterizers to describe an image in vector graphics format (e.g., shape) and convert it into a raster image (e.g., a series of pixels, points, or lines that, when displayed together, produce an image represented by shapes); one or more hierarchical depth buffers (Hiz) for buffering data; and / or one or more pixel backends. In at least one embodiment, a Slice 1901A-1901N includes a memory fabric, e.g., an L2 cache.

[0272] In at least one embodiment, the FPUs 1914A-1914N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPUs 1915A-1915N perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALUs 1916A-1916N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPUs 1917A-1917N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations.In at least one embodiment, MPUs 1917-1917N can perform a variety of matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general matrix-to-matrix multiplication (GEMM). In at least one embodiment, AFUs 1912A-1912N can perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0273] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig.8A and / or 8B provided. In at least one embodiment, the logic 815 in the graphics kernel 1900 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0274] In at least one embodiment, the graphics core 1900 includes an intermediate link and a link fabric sublayer connected to a switch and a GPU-to-GPU bridge, enabling multiple graphics processors 1900 (e.g., 8) to be interconnected without adhesive bonding, with load / store units (LSUs), data transfer units, and synchronization semantics across multiple graphics processors 1900. In at least one embodiment, the intermediate links include standardized intermediate links (e.g., PCIe) or a combination thereof.

[0275] In at least one embodiment, the 1900 graphics core includes multiple tiles. In at least one embodiment, a tile is a single chip or one or more chips, wherein individual chips can be connected by an interconnect (e.g., an embedded multi-die interconnect bridge, EMIB). In at least one embodiment, the 1900 graphics core includes a compute tile, a memory tile (e.g., which a memory tile can access exclusively through different tiles or different chipsets, such as a Rambo tile), a substrate tile, a base tile, an HMB tile, a link tile, and an EMIB tile, wherein all tiles are packaged together in the 1900 graphics core as part of a GPU. In at least one embodiment, the 1900 graphics core can include multiple tiles in a single package (also referred to as a "multi-tile package").In at least one embodiment, a compute tile can comprise eight 1900 graphics cores and an L1 cache; and a base tile can comprise a host interface with PCIe 5.0, HBM2e, MDFI, and EMIB, a link tile with eight links and eight ports with an embedded switch. In at least one embodiment, the tiles are connected by a face-to-face (F2F) chip-on-chip connection using microbumps (e.g., copper pillars) with a fine pitch of 36 micrometers. In at least one embodiment, the 1900 graphics core includes a memory fabric that contains a memory module and is a tile accessible to multiple tiles.In at least one embodiment, the 1900 graphics core stores, accesses, or loads its own hardware contexts in main memory, wherein a hardware context is a set of data that is loaded from registers before a process is resumed, and wherein a hardware context can specify a state of the hardware (e.g., the state of a GPU).

[0276] In at least one embodiment, the graphics core includes 1900 serializer / deserializer circuits (SERDES) that convert a serial data stream into a parallel data stream or convert a parallel data stream into a serial data stream.

[0277] In at least one embodiment, the 1900 graphics core includes a coherent, unified high-speed fabric (GPU to GPU), load / store units, bulk data transmission and synchronization semantics, and connected GPUs via an embedded switch, wherein a GPU-GPU bridge is controlled by a controller.

[0278] In at least one embodiment, the 1900 graphics kernel performs an API, wherein said API abstracts the hardware of the 1900 graphics kernel and accesses libraries containing instructions for performing mathematical operations (e.g., Math Kernel library), operations in deep neural networks (e.g., Deep Neural Network library), vector operations, collective communication, thread building blocks, video processing, data analysis library, and / or ray tracing operations.

[0279] In at least one embodiment, an embodiment of at least one of Fig.19A include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0280] Fig.Figure 19B illustrates GPGPU 1930, which can be configured to perform highly parallel computing operations on an array of graphics processing units in at least one embodiment. In at least one embodiment, the GPGPU 1930 can be directly connected to other instances of the GPGPU 1930 to create a multi-GPU cluster, which improves the training speed for deep neural networks. In at least one embodiment, the GPGPU 1930 includes a host interface 1932 to enable communication with a host processor. In at least one embodiment, the host interface 1932 is a PCI Express interface. In at least one embodiment, the host interface 1932 can be a vendor-specific communication interface or communication fabric.In at least one embodiment, the GPGPU 1930 receives instructions from a host processor and uses a global scheduler 1934 (which may be referred to as a thread sequencer and / or asynchronous computation machine) to distribute the execution threads associated with these instructions across a set of compute clusters 1936A-1936H. In at least one embodiment, the compute clusters 1936A-1936H share a cache memory 1938. In at least one embodiment, the cache memory 1938 can serve as a higher-level cache for cache memory within the compute clusters 1936A-1936H. In at least one embodiment, the compute clusters 1936A-1936H comprise a slice or are referred to as "slices". In at least one embodiment, the GPGPU 1930 is part of a SoC, for example, part of the integrated circuit 1700. Fig. 17).

[0281] In at least one embodiment, the GPGPU 1930 includes the main memory 1944A-1944B, which is coupled to compute clusters 1936A-1936H via a set of main memory controllers 1942A-1942B (e.g., one or more controllers for HBM2e). In at least one embodiment, the main memory 1944A-1944B can include various types of main memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate memory (GDDR).

[0282] In at least one embodiment, the computing clusters 1936A-1936H each include a set of graphics cores, such as the 1900 graphics core from Fig.19A, which can include several types of integer and floating-point logic units capable of performing arithmetic operations in a range of precisions, including those suitable for machine learning calculations. For example, in at least one embodiment, at least a subset of floating-point units in each of the computing clusters 1936A-1936H can be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of floating-point units can be configured to perform 64-bit floating-point operations.

[0283] In at least one embodiment, multiple instances of GPGPU 1930 can be configured to operate as a compute cluster. In at least one embodiment, the communication used by the compute clusters 1936A-1936H for synchronization and data exchange varies between embodiments. In at least one embodiment, multiple instances of GPGPU 1930 communicate via the host interface 1932. In at least one embodiment, GPGPU 1930 includes an I / O hub 1939 that couples GPGPU 1930 to a GPU link 1940, enabling a direct connection to other instances of GPGPU 1930. In at least one embodiment, GPU link 1940 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1930.In at least one embodiment, the GPU link 1940 is coupled with a high-speed intermediate connection to send and receive data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 1930 are located in separate data processing systems and communicate with a network device accessible via the host interface 1932. In at least one embodiment, the GPU link 1940 can be configured to provide a connection to a host processor in addition to, or as an alternative to, the host interface 1932.

[0284] In at least one embodiment, the GPGPU 1930 can be configured to train neural networks. In at least one embodiment, the GPGPU 1930 can be used within an inference platform. In at least one embodiment where the GPGPU 1930 is used for inference, the GPGPU 1930 can include fewer compute clusters 1936A-1936H with respect to its use for training a neural network. In at least one embodiment, the memory technology allocated to the main memory 1944A-1944B can differ between inference and training configurations, with higher-bandwidth memory technologies being provided for training configurations. In at least one embodiment, an inference configuration can support specific inference instructions from the GPGPU 1930.For example, in at least one embodiment, an inference configuration can support one or more 8-bit integer dot product instructions that can be used during inference operations for deployed neural networks.

[0285] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the GPGPU 1930 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0286] In at least one embodiment, an embodiment of at least one of the Fig. 19A and / or 19B contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised, at least in part, based on the identification of different types of inference or training data within the inference or training data of neural networks that are to be denoised separately, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to the Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0287] Fig.Figure 20 is a block diagram illustrating a computing system 2000 according to at least one embodiment. In at least one embodiment, the computing system 2000 includes a processing subsystem 2001 comprising one or more processors 2002 and a system memory 2004, which communicate via an intermediate link path that may include a memory hub 2005. In at least one embodiment, the memory hub 2005 may be a separate component within a chipset component or be integrated into one or more processors 2002. In at least one embodiment, the memory hub 2005 is coupled to an I / O subsystem 2011 via a communication link 2006. In at least one embodiment, the I / O subsystem 2011 includes an I / O hub 2007 that enables the computing system 2000 to receive input from one or more input devices 2008.In at least one embodiment, the I / O hub 2007 enables a display controller, which may be included in one or more processors 2002, to provide outputs to one or more display devices 2010A. In at least one embodiment, one or more display devices 2010A coupled to the I / O hub 2007 may include a local, internal, or embedded display device.

[0288] In at least one embodiment, the processing subsystem 2001 includes one or more parallel processors 2012 coupled to the memory hub 2005 via a bus or other communication link 2013. In at least one embodiment, the communication link 2013 can use any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or it can be a vendor-specific communication interface or communication fabric. In at least one embodiment, one or more parallel processors 2012 form a computationally focused parallel or vector processing system that can include a large number of processing cores and / or processing clusters, such as a many-integrated-core (MIC) processor.In at least one embodiment, some or all of the parallel processors 2012 form a graphics processing subsystem that can output pixels to one or more display devices 2010A coupled via an I / O hub 2007. In at least one embodiment, one or more parallel processors 2012 may also include a display controller and a display interface (not shown) to enable a direct connection to one or more display devices 2010B. In at least one embodiment, one or more parallel processors 2012 include one or more cores, such as the graphics cores 1900 described herein.

[0289] In at least one embodiment, a system memory unit 2014 can be connected to an I / O hub 2007 to provide a storage mechanism for the computing system 2000. In at least one embodiment, an I / O switch 2016 can be used to provide an interface mechanism that enables connections between the I / O hub 2007 and other components, such as a network adapter 2018 and / or a wireless network adapter 2019, which may be integrated into the platform, and various other devices that can be added via one or more add-in devices 2020. In at least one embodiment, the network adapter 2018 can be an Ethernet adapter or another wired network adapter.In at least one embodiment, the wireless network adapter 2019 may include one or more Wi-Fi, Bluetooth, near field communication (NFC) or other network devices, which include one or more wireless radio devices.

[0290] In at least one embodiment, the Computing System 2000 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video recording devices, and similar devices, which may also be connected to the I / O Hub 2007. In at least one embodiment, communication paths connecting various components in Fig.20 interconnect, implemented using any suitable protocols, such as PCI-based (Peripheral Component Interconnect) protocols (e.g. PCI-Express) or other bus or point-to-point communication interfaces and / or protocol(s), such as NV-Link high-speed interlink or interlink protocols.

[0291] In at least one embodiment, Parallel Processors 2012 include circuitry optimized for graphics and video processing, including, for example, video output circuitry, and constitute a graphics processing unit (GPU), e.g., Parallel Processors 2012 include a Graphics Core 1900. In at least one embodiment, Parallel Processors 2012 include circuitry optimized for general-purpose processing. In at least one embodiment, components of the Computing System 2000 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more Parallel Processors 2012, the Memory Hub 2005, one or more Processors 2002, and the I / O Hub 2007 can be integrated into an integrated system-on-a-chip (SoC) circuit.In at least one embodiment, components of the Computing System 2000 can be integrated into a single package to form a System-in-Package (SIP) configuration. In at least one embodiment, at least some of the components of the Computing System 2000 can be integrated into a Multi-Chip Module (MCM) that can be connected to other Multi-Chip Modules to form a modular Computing System.

[0292] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig.8A and / or 8B provided. In at least one embodiment, the logic 815 in the Computing System 2000 can be used to infer or predict operations at least partly based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0293] In at least one embodiment, an embodiment of at least one of Fig.20 include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented. PROCESSORS

[0294] Fig.Figure 21A illustrates a parallel processor 2100 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2100 can be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 2100 is a variant of one or more parallel processors 212, which are implemented in Fig. Figure 20 shows an exemplary embodiment. In at least one embodiment, a parallel processor 2100 includes one or more graphics cores 1900.

[0295] In at least one embodiment, the parallel processor 2100 includes a parallel processing unit 2102. In at least one embodiment, the parallel processing unit 2102 includes an I / O unit 2104, which enables communication with other devices, including other instances of the parallel processing unit 2102. In at least one embodiment, the I / O unit 2104 can be directly connected to other devices. In at least one embodiment, the I / O unit 2104 connects to other devices via a hub or switch interface, such as a memory hub 2105. In at least one embodiment, connections between the memory hub 2105 and the I / O unit 2104 form a communication link 2113.In at least one embodiment, the I / O unit 2104 is connected to a host interface 2106 and a memory crossbar 2116, wherein the host interface 2106 receives commands aimed at performing processing operations, and the memory crossbar 2116 receives commands aimed at performing memory operations.

[0296] In at least one embodiment, the host interface 2106, upon receiving an instruction buffer via the I / O unit 2104, can direct work operations to a frontend 2108 to execute these instructions. In at least one embodiment, the frontend 2108 is coupled to a scheduler 2110 (which can be referred to as a sequencer) configured to distribute instructions or other work items to a processing cluster arrangement 2112. In at least one embodiment, the scheduler 2110 ensures that the processing cluster arrangement 2112 is properly configured and in a valid state before tasks are distributed to a cluster of the processing cluster arrangement 2112. In at least one embodiment, the scheduler 2110 is implemented via firmware logic running on a microcontroller.In at least one embodiment, the scheduler 2110 implemented in a microcontroller is configurable to perform complex scheduling and workload distribution operations with coarse and fine granularity, enabling fast pre-suppression and context switching of threads running on the processing cluster 2112. In at least one embodiment, the host software can detect workloads for scheduling on the processing cluster 2112 via one of several graphics processing paths. In at least one embodiment, workloads can then be automatically distributed across the processing cluster 2112 by the logic of the scheduler 2110 within a microcontroller containing the scheduler 2110.

[0297] In at least one embodiment, the processing cluster arrangement 2112 can contain up to "N" processing clusters (e.g., cluster 2114A, cluster 2114B to cluster 2114N), where "N" is a positive integer (which may be a different integer "N" than the one used in other figures). In at least one embodiment, each cluster 2114A-2114N of the processing cluster arrangement 2112 can execute a large number of concurrent threads. In at least one embodiment, the scheduler 2110 can allocate work to the clusters 2114A-2114N of the processing cluster arrangement 2112 by using various scheduling and / or workload distribution algorithms that can vary depending on the workload generated for each type of program or computation.In at least one embodiment, scheduling can be performed dynamically by the scheduler 2110 or partially supported by the compiler logic during the compilation of the program logic configured for execution by the processing cluster arrangement 2112. In at least one embodiment, different clusters 2114A-2114N of the processing cluster arrangement 2112 can be assigned for processing different types of programs or for performing different types of calculations.

[0298] In at least one embodiment, the processing cluster arrangement 2112 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster arrangement 2112 is configured to perform parallel general-purpose computing operations. For example, in at least one embodiment, the processing cluster arrangement 2112 can include logic for performing processing tasks, including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.

[0299] In at least one embodiment, the processing cluster arrangement 2112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster arrangement 2112 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as tiling logic and other vertex processing logic. In at least one embodiment, the processing cluster arrangement 2112 may be configured to execute shader programs with respect to graphics processing, such as, but not limited to, vertex shaders, tiling shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2102 may transfer data from system memory via the I / O unit 2104 for processing.In at least one embodiment, transmitted data can be stored in main memory on a chip (e.g., in the parallel processor main memory 2122) during processing and then written back to main memory.

[0300] In at least one embodiment, when the parallel processing unit 2102 is used to perform graphics processing, the scheduler 2110 can be configured to divide a processing workload into approximately equal tasks to enable better distribution of graphics processing operations across multiple clusters 2114A-2114N of the processing cluster arrangement 2112. In at least one embodiment, parts of the processing cluster arrangement 2112 can be configured to perform different types of processing.For example, in at least one embodiment, a first part can be configured to perform vertex shading and topology generation, a second part can be configured to perform tiling and geometry shading, and a third part can be configured to perform pixel shading or other screen-space operations to produce a rendered image for display. In at least one embodiment, intermediate data produced by one or more clusters 2114A-2114N can be stored in buffers to allow the transfer of intermediate data between clusters 2114A-2114N for further processing.

[0301] In at least one embodiment, the processing cluster arrangement 2112 can receive processing tasks to be executed via the scheduler 2110, which receives commands defining processing tasks from the frontend 2108. In at least one embodiment, processing tasks can include indices of data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., which program is to be executed). In at least one embodiment, the scheduler 2110 can be configured to retrieve indices according to the tasks or to receive indices from the frontend 2108. In at least one embodiment, the frontend 2108 can be configured to ensure that the processing cluster arrangement 2112 is configured in a valid state before an incoming command buffer (e.g.,Batch buffer, push buffer, etc.) is initiated with a specified workload.

[0302] In at least one embodiment, each of the one or more instances of the parallel processing unit 2102 can be coupled to a parallel processor memory 2122. In at least one embodiment, the parallel processor memory 2122 can be accessed via a memory crossbar 2116, which can receive memory requests from both the processing cluster arrangement 2112 and the I / O unit 2104. In at least one embodiment, the memory crossbar 2116 can access the parallel processor memory 2122 via a memory interface 2118. In at least one embodiment, the memory interface 2118 can include several partition units (e.g., partition unit 2120A, partition unit 2120B to partition unit 2120N), each of which can be coupled to a portion (e.g., a memory unit) of the parallel processor memory 2122.In at least one embodiment, the number of partition units 2120A-2120N is configured to correspond to the number of memory units, such that a first partition unit 2120A has a corresponding first memory unit 2124A, a second partition unit 2120B has a corresponding memory unit 2124B, and an Nth partition unit 2120N has a corresponding Nth memory unit 2124N. In at least one embodiment, the number of partition units 2120A-2120N can be different from the number of memory units.

[0303] In at least one embodiment, the memory units 2124A-2124N can include various types of memory devices, including dynamic random-access memory (DRAM) or graphics random-access memory, such as synchronous graphics random-access memory (SGRAM), including graphics double data rate memory (GDDR). In at least one embodiment, the memory units 2124A-2124N can also include 3D stacked memory, including, but not limited to, high-bandwidth memory (HBM), HBM2e, or HDM3. In at least one embodiment, render targets, such as framebuffers or texture maps, can be stored across the memory units 2124A-2124N, so that partition units 2120A-2120N can write portions of each render target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 2122.In at least one embodiment, a local instance of the parallel processor memory 2122 can be excluded in favor of a unified memory design that uses the system memory in conjunction with a local cache memory.

[0304] In at least one embodiment, each of the clusters 2114A-2114N of the processing cluster arrangement 2112 can process data written to any of the memory units 2124A-2124N within the parallel processor memory 2122. In at least one embodiment, the memory crossbar 2116 can be configured to transfer an output from each cluster 2114A-2114N to a partition unit 2120A-2120N or to another cluster 2114A-2114N, which can perform additional processing operations on the output. In at least one embodiment, each cluster 2114A-2114N can communicate with the memory interface 2118 via the memory crossbar 2116 to read from or write to various external memory devices.In at least one embodiment, the memory crossbar 2116 has a connection to the memory interface 2118 for communication with the I / O unit 2104, as well as a connection to a local instance of the parallel processor memory 2122, enabling processing units in different processing clusters 2114A-2114N to communicate with the system memory or other memory that is not local to the parallel processing unit 2102. In at least one embodiment, the memory crossbar 2116 can use virtual channels to separate traffic flows between the clusters 2114A-2114N and the partition units 2120A-2120N.

[0305] In at least one embodiment, multiple instances of the Parallel Processing Unit 2102 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of the Parallel Processing Unit 2102 can be configured to work together, even if different instances have different numbers of processor cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the Parallel Processing Unit 2102 can include higher-precision floating-point units relative to other instances.In at least one embodiment, systems containing one or more instances of the Parallel Processing Unit 2102 or the Parallel Processor 2100 can be implemented in a variety of configurations and form factors, including, but not limited to, desktop, laptop or handheld PCs, servers, workstations, game consoles and / or embedded systems.

[0306] Fig. 21B is a block representation of a partition unit 2120 according to at least one embodiment. In at least one embodiment, the partition unit 2120 is an instance of one of the partition units 2120A-2120N of Fig.21A. In at least one embodiment, the partition unit 2120 includes an L2 cache 2121, a framebuffer interface 2125, and a ROP 2126 (raster operation unit). In at least one embodiment, the L2 cache 2121 is a read / write cache configured to perform load and store operations received from the memory crossbar 2116 and the ROP 2126. In at least one embodiment, read errors and urgent write-back requests are issued from the L2 cache 2121 to the framebuffer interface 2125 for processing. In at least one embodiment, updates can also be sent to a framebuffer for processing via the framebuffer interface 2125. In at least one embodiment, the framebuffer interface 2125 is connected to one of the memory units in the parallel processor's main memory, such as the memory units 2124A-2124N of Fig.21A (e.g., within the parallel processor memory 2122).

[0307] In at least one embodiment, the ROP 2126 is a processing unit that performs raster operations such as stenciling, Z-testing, blending, etc. In at least one embodiment, the ROP 2126 then outputs processed graphics data, which is stored in the graphics memory. In at least one embodiment, the ROP 2126 includes compression logic to compress depth or color data written to memory and to decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that uses one or more of several compression algorithms. In at least one embodiment, the type of compression performed by the ROP 2126 can vary based on statistical properties of the data to be compressed.For example, in at least one embodiment, delta color compression is performed on depth and color data on a per-tile basis.

[0308] In at least one embodiment, the ROP 2126 is in each processing cluster (e.g., clusters 2114A-2114N of Fig. 21A) instead of in the partition unit 2120. In at least one embodiment, read and write requests for pixel data are transmitted via the memory crossbar 2116 instead of pixel fragment data. In at least one embodiment, processed graphics data can be displayed on a display device, such as one or more display devices 2110 from Fig. 20, routed for further processing by one or more processors 2002 or for further processing by one of the processing entities within the parallel processor 2100 Fig. Routed to 21A.

[0309] Fig.21C is a block representation of a processing cluster 2114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of the processing clusters 2114A-2114N of Fig.21A. In at least one embodiment, the processing cluster 2114 can be configured to execute many threads in parallel, where "thread" refers to an instance of a particular program running on a particular set of input data. In at least one embodiment, single-instruction, multiple-data (SIMD) instruction-issuing techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single-instruction, multiple-thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads, using a common instruction unit configured to issue instructions to a set of processing machines in each processing cluster.

[0310] In at least one embodiment, the operation of the processing cluster 2114 can be controlled via a pipeline manager 2132, which distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2132 receives instructions from the scheduler 2110. Fig.21A and manages the execution of these instructions via a graphics multiprocessor 2134 and / or a texture unit 2136. In at least one embodiment, the graphics multiprocessor 2134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, different types of SIMT parallel processors with different architectures can be included within the processing cluster 2114. In at least one embodiment, one or more instances of the graphics multiprocessor 2134 can be included within a processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 can process data, and a data crossbar 2140 can be used to distribute processed data to one of several possible destinations, including other shader units.In at least one embodiment, the pipeline manager 2132 can facilitate the distribution of processed data by prescribing destinations for processed data to be distributed via the data crossbar 2140.

[0311] In at least one embodiment, each graphics multiprocessor 2134 within the processing cluster 2114 can include an identical set of functional execution logic (e.g., arithmetic logic units, load / store units, etc.). In at least one embodiment, the functional execution logic can be configured in a pipelined manner, allowing new instructions to be issued before previous instructions have completed. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same functional unit hardware can be used to perform different operations, and any combination of functional units can be present.

[0312] In at least one embodiment, the instructions transmitted to the processing cluster 2114 form a thread. In at least one embodiment, a set of threads executed across a set of parallel processing machines is a thread group. In at least one embodiment, a thread group executes a common program with different input data. In at least one embodiment, each thread within a thread group can be assigned to a different processing machine within a graphics multiprocessor 2134. In at least one embodiment, a thread group can contain fewer threads than the number of processing machines within a graphics multiprocessor 2134.In at least one embodiment, if a thread group contains fewer threads than the number of processing machines, one or more of the processing machines may be idle during the cycles in which that thread group is processed. In at least one embodiment, a thread group may also contain more threads than the number of processing machines within a graphics multiprocessor 2134. In at least one embodiment, processing may be performed over successive clock cycles if a thread group contains more threads than the number of processing machines within the graphics multiprocessor 2134. In at least one embodiment, processing may be performed over successive clock cycles if a thread group contains more threads than the number of processing machines in the graphics multiprocessor 2134.

[0313] In at least one embodiment, the graphics multiprocessor 2134 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2134 can forgo an internal cache and use a cache memory (e.g., L1 cache 2148) within the processing cluster 2114. In at least one embodiment, each graphics multiprocessor 2134 also has access to L2 caches in partition units (e.g., partition units 2120A-2120N of Fig.21A), which are shared by all processing clusters 2114 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2134 can also access a global memory outside the chip, which may include one or more local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 2102 can be used as global memory. In at least one embodiment, the processing cluster 2114 includes multiple instances of the graphics multiprocessor 2134 and can share common instructions and data, which can be stored in the L1 cache 2148.

[0314] In at least one embodiment, each processing cluster 2114 can include an MMU 2145 (memory management unit) configured to assign virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2145 can be accessed in the memory interface 2118 of Fig.21A. In at least one embodiment, the MMU 2145 includes a set of page table entries (PTEs) used to map a virtual address to a physical address of a tile and optionally a cache row index. In at least one embodiment, the MMU 2145 may include address translation buffers (TLBs) or caches that may reside in the graphics multiprocessor 2134, the L1 cache 2148, or the processing cluster 2114. In at least one embodiment, a physical address is processed to distribute access to surface data locally, enabling efficient request nesting between partition units. In at least one embodiment, a cache row index may be used to determine whether a request for a cache row is a hit or a failure.

[0315] In at least one embodiment, a processing cluster 2114 can be configured such that each graphics multiprocessor 2134 is coupled to a texture unit 2136 to perform texture allocation operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 2134 and, if necessary, retrieved from an L2 cache, local parallel processor memory, or system memory.In at least one embodiment, each graphics multiprocessor 2134 outputs processed tasks to the data crossbar 2140 to make processed tasks available to another processing cluster 2114 for further processing or to store processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 2116. In at least one embodiment, a preROP 2142 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 2134 and forward data to ROP units that may be located in the partition units described herein (e.g., partition units 2120A-2120N in ). Fig. 21A). In at least one embodiment, the preROP 2142 unit can perform color blending optimizations, organize pixel color data, and perform address translations.

[0316] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the graphics processing cluster 2114 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0317] In at least one embodiment, an embodiment of at least one of the Fig. 20, Fig. 21A, Fig.21B and / or 21C contain or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to the Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0318] Fig.Figure 21D shows a graphics multiprocessor 2134 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 2134 is coupled to the pipeline manager 2132 of the processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 has an execution pipeline which, without limitation, includes an instruction cache 2152, an instruction unit 2154, an address allocation unit 2156, a register bank 2158, one or more cores of a general-purpose graphics processing unit (GPGPU) 2162, and one or more load / store units 2166, wherein one or more load / store units 2166 can perform load / store operations to load / store instructions corresponding to the execution of an operation.In at least one embodiment, GPGPU cores 2162 and load / store units 2166 are coupled to a cache memory 2172 and a shared memory 2170 via a memory and cache intermediate 2168. In at least one embodiment, the GPGPU cores 2162 are part of a SoC, for example, part of the integrated circuit 1700 from [company name missing]. Fig. 17.

[0319] In at least one embodiment, the instruction cache 2152 receives a stream of instructions for execution from the pipeline manager 2132. In at least one embodiment, instructions are cached in the instruction cache 2152 and distributed for execution by an instruction unit 2154. In at least one embodiment, the instruction unit 2154 can distribute instructions as thread groups (e.g., warps, wavefronts, waves), with each thread of the thread group being assigned to a different execution unit within the GPGPU cores 2162. In at least one embodiment, an instruction can access any address space—local, shared, or global—by specifying an address within a unified address space.In at least one embodiment, the address allocation unit 2156 can be used to translate addresses in a unified address space into a unique working memory address that can be accessed by load / store units 2166.

[0320] In at least one embodiment, the register bank 2158 provides a set of registers for functional units of the graphics multiprocessor 2134. In at least one embodiment, the register bank 2158 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 2162, load / store units 2166) of the graphics multiprocessor 2134. In at least one embodiment, the register bank 2158 is partitioned among the individual functional units such that each functional unit is assigned a dedicated portion of the register bank 2158. In at least one embodiment, the register bank 2158 is partitioned between different warps (which may be referred to as wavefronts and / or waves) executed by the graphics multiprocessor 2134.

[0321] In at least one embodiment, GPGPU cores 2162 can each include floating-point units (FPUs) and / or arithmetic logic units (ALUs) for integers, which are used to execute instructions of the graphics multiprocessor 2134. In at least one embodiment, the GPGPU cores 2162 can have a similar architecture or differ in their architecture. In at least one embodiment, a first set of GPGPU cores 2162 includes a single-precision FPU and an integer ALU, while a second set of GPGPU cores includes a double-precision FPU. In at least one embodiment, the FPUs can implement the floating-point arithmetic of the IEEE 754-2008 standard or enable variable-precision floating-point arithmetic.In at least one embodiment, the graphics multiprocessor 2134 may additionally include one or more fixed-function or special-purpose function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores 2162 may also include fixed-function or special-purpose logic.

[0322] In at least one embodiment, GPGPU cores 2162 include SIMD logic capable of executing a single instruction on multiple data records. In at least one embodiment, GPGPU cores 2162 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated at compile time by a shader compiler or automatically when programs written and compiled for SPMD (Single Program Multiple Data) or SIMT architectures are executed. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction.For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel over a single SIMD8 logic unit.

[0323] In at least one embodiment, the memory and cache intermediate 2168 is an intermediate network that connects each functional unit of the graphics multiprocessor 2134 to the register bank 2158 and the shared memory 2170. In at least one embodiment, the memory and cache intermediate 2168 is a crossbar intermediate that allows the load / store unit 2166 to implement load and store operations between the shared memory 2170 and the register bank 2158. In at least one embodiment, the register bank 2158 can operate at the same frequency as the GPGPU cores 2162, enabling very low latency data transmission between the GPGPU cores 2162 and the register bank 2158.In at least one embodiment, the shared memory 2170 can be used to enable communication between threads running on functional units within the graphics multiprocessor 2134. In at least one embodiment, the cache memory 2172 can be used as a data cache, for example, to cache texture data communicated between functional units and the texture unit 2136. In at least one embodiment, the shared memory 2170 can also be used as a programmatically managed cache. In at least one embodiment, threads running on GPGPU cores 2162 can programmatically store data in the shared memory in addition to automatically cached data stored in the cache memory 2172.

[0324] In at least one embodiment, a parallel processor or GPGPU, as described herein, is communicatively coupled to host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, a GPU can be communicatively coupled to a host processor / cores via a bus or other intermediate connection (e.g., a high-speed intermediate such as PCIe or NVLink). In at least one embodiment, a SoC includes a parallel processor or GPGPU, as described herein, wherein said parallel processor or GPGPU is implemented on the named SoC. In at least one embodiment, a GPU can be integrated as cores on a package or chip and communicatively coupled to cores via an internal processor bus / intermediate connection within a package or chip.In at least one embodiment, processor cores, regardless of how a GPU is connected, can assign work to that GPU in the form of instruction sequences contained in a work descriptor. In at least one embodiment, this GPU then uses dedicated circuitry / logic to efficiently process these instructions.

[0325] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig.8A and / or 8B provided. In at least one embodiment, the logic 815 in the graphics multiprocessor 2134 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0326] In at least one embodiment, an embodiment of at least one of Fig.21D cause or include one or more processors, circuits, or systems causing inference or training data of neural networks to be denoised, at least in part, based on identifying different types of inference or training data within the inference or training data of neural networks that are to be denoised separately, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0327] Fig.Figure 22 illustrates a multi-GPU computing system 2200 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2200 can include a processor 2202 coupled to multiple general-purpose graphics processing units (GPGPUs) 2206A-D via a host interface switch 2204. In at least one embodiment, the host interface switch 2204 is a PCI Express switch device that couples the processor 2202 to a PCI Express bus through which the processor 2202 can communicate with the GPGPUs 2206A-D. In at least one embodiment, the GPGPUs 2206A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2216. In at least one embodiment, GPU-to-GPU links 2216 connect to the respective GPGPUs 2206A-D via a dedicated GPU link.In at least one embodiment, P2P GPU links 2216 enable direct communication between the respective GPGPUs 2206A-D without requiring communication via the host interface bus 2204, to which the processor 2202 is connected. In at least one embodiment, where the GPU-to-GPU traffic is routed to P2P GPU links 2216, the host interface bus 2204 remains available for accessing system memory or for communication with other instances of the multi-GPU computing system 2200, for example, via one or more network devices. Although in at least one embodiment the GPGPUs 2206A-D are connected to the processor 2202 via the host interface switch 2204, in at least one embodiment the processor 2202 includes direct support for P2P GPU links 2216 and can be directly connected to the GPGPUs 2206A-D.In at least one embodiment, the GPGPUs 2206A-D are part of a SoC, such as part of the integrated circuit 1700 in . Fig. 17, wherein the GPGPUs 2206A-D perform the operations described herein.

[0328] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the multi-GPU computing system 2200 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0329] In at least one embodiment, the multi-GPU computing system 2200 contains one or more graphics cores 1900.

[0330] In at least one embodiment, an embodiment of at least one of Fig. 22 include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0331] Fig.Figure 23 is a block diagram of a graphics processor 2300 according to at least one embodiment. In at least one embodiment, the graphics processor 2300 includes a ring interconnect 2302, a pipeline front end 2304, a media machine 2337, and graphics cores 2380A-2380N. In at least one embodiment, the ring interconnect 2302 couples the graphics processor 2300 to other processing units containing other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2300 is one of many processors integrated in a multi-core processing system. In at least one embodiment, the graphics processor 2300 includes the graphics core 1900.

[0332] In at least one embodiment, the graphics processor 2300 receives instruction stacks via a ring interconnect 2302. In at least one embodiment, incoming instructions are interpreted by an instruction streamer 2303 in a pipeline front-end unit 2304. In at least one embodiment, the graphics processor 2300 includes scalable execution logic to perform 3D geometry processing and media processing via one or more graphics cores 2380A-2380N. In at least one embodiment, the instruction streamer 2303 provides instructions for 3D geometry processing to the geometry pipeline 2336. In at least one embodiment, the instruction streamer 2303 provides instructions for at least some media processing to a video front-end 2334 coupled to the media machine 2337.In at least one embodiment, the media machine 2337 includes a video quality engine (VQE) 2330 for post-processing videos and images and a multi-format encoder / decoder (MFX) 2333 to provide hardware-accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2336 and the media machine 2337 each generate execution threads for thread execution resources provided by at least one graphics core 2380.

[0333] In at least one embodiment, the graphics processor 2300 includes scalable thread execution resources with graphics cores 2380A-2380N (which may be modular and are sometimes referred to as core slices), each comprising multiple subcores 2350A-2350N, 2360A-2360N (sometimes referred to as core subslices). In at least one embodiment, the graphics processor 2300 may have any number of graphics cores 2380A. In at least one embodiment, the graphics processor 2300 includes a graphics core 2380A with at least one first subcore 2350A and one second subcore 2360A. In at least one embodiment, the graphics processor 2300 is a low-performance processor with a single subcore (e.g., 2350A). In at least one embodiment, the graphics processor 2300 contains multiple graphics cores 2380A-2380N, each containing a set of first subcores 2350A-2350N and a set of second subcores 2360A-2360N.In at least one embodiment, each subcore in the first subcores 2350A-2350N contains at least one first set of execution units 2352A-2352N and media / texture samplers 2354A-2354N. In at least one embodiment, each subcore in the second subcores 2360A-2360N contains at least one second set of execution units 2362A-2362N and samplers 2364A-2364N. In at least one embodiment, each subcore 2350A-2350N, 2360A-2360N shares a set of shared resources 2370A-2370N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic. In at least one embodiment, the graphics processor 2300 includes load / store units in the pipeline front end 2304.

[0334] Logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding Logic 815 are provided herein in conjunction with Fig. 8A and / or 8B provided. In at least one embodiment, the logic 815 in the graphics processor 2300 can be used to infer or predict operations at least partially based on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.

[0335] In at least one embodiment, an embodiment of at least one of Fig.23 include or cause one or more processors, circuits, or systems to cause inference or training data of neural networks to be denoised separately, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks, using a corresponding number of neural diffusion networks according to various embodiments described above in relation to Fig. The noise reduction measures discussed in points 1-7 will be implemented.

[0336] Fig.Figure 24 is a block diagram illustrating the microarchitecture for a 2400 processor, which, according to at least one embodiment, can include logic circuits for executing instructions. In at least one embodiment, the 2400 processor can execute instructions, including x86 instructions, ARM instructions, specialized instructions for application-specific integrated circuits (ASICs), etc. In at least one embodiment, the 2400 processor can include registers for storing packed data, such as 64-bit wide MMX™ registers in microprocessors enabled by MMX technology from Intel Corporation, Santa Clara, California. In at least one embodiment, MMX registers, available in both integer and floating-point form, can operate with packed data elements that accompany single-instruction, multiple-data ("SIMD"), and streaming SIMD extensions ("SSE") instructions.In at least one embodiment, 128-bit wide XMM registers can contain such packed data operands with respect to SSE2, SSE3, SSE4, AVX, or beyond (generically referred to as "SSEx") technology. In at least one embodiment, the processor can execute 2400 instructions for accelerating machine learning or deep learning algorithms, training, or inference.

[0337] In at least one embodiment, the processor 2400 includes an in-order frontend (“frontend”) 2401 for retrieving instructions to be executed and preparing instructions for later use in a processor pipeline. In at least one embodiment, the frontend 2401 can include multiple units. In at least one embodiment, an instruction pre-retriever 2426 retrieves instructions from memory and feeds instructions into an instruction decoder 2428, which in turn decodes or interprets instructions. For example, in at least one embodiment, the instruction decoder 2428 decodes a received instruction into one or more operations, referred to as “micro-instructions” or “micro-operations” (also called “micro-ops” or “u-ops” or “µ-ops”), that a machine can execute.In at least one embodiment, the instruction decoder 2428 analyzes an instruction into an operation code and corresponding data and control fields that can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, a trace cache 2430 can assemble decoded Uops into program-ordered sequences or traces in a Uop queue 2434 for execution. In at least one embodiment, when the trace cache 2430 encounters a complex instruction, a microcode ROM 2432 provides the Uops required to complete an operation.

[0338] In at least one embodiment, some instructions can be converted into a single micro-op, while others require multiple micro-ops to complete the full operation. In at least one embodiment, the instruction decoder 2428 can access the microcode ROM 2432 to execute an instruction if more than four micro-ops are required to complete it. In at least one embodiment, an instruction can be decoded into a small number of micro-ops for processing in the instruction decoder 2428. In at least one embodiment, an instruction can be stored in the microcode ROM 2432 should a number of micro-ops be required to perform such an operation.In at least one embodiment, the trace cache 2430 refers to a programmable logic array (PLA) as an entry point to determine a correct micro-instruction pointer for reading microcode sequences to complete one or more instructions from the microcode ROM 2432 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2432 has completed sequencing micro-ops for an instruction, the front end 2401 of a machine can resume fetching micro-ops from the trace cache 2430.

[0339] In at least one embodiment, the out-of-order execution machine (“out-of-order machine”) can prepare 2403 instructions for execution. In at least one embodiment, the out-of-order execution logic includes a number of buffers to smooth and reorder the flow of instructions to optimize performance as they traverse a pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution machine 2403 includes, without limitation, an assigner / register renamer 2440, a memory up-op queue 2442, an integer / floating-point up-op queue 2444, a memory scheduler 2446, a fast scheduler 2402, a slow / general floating-point scheduler (“slow / general FP scheduler”) 2404, and a simple floating-point scheduler (“simple FP scheduler”) 2406.In at least one embodiment, the fast scheduler 2402, the slow / general floating-point scheduler 2404, and the simple floating-point scheduler 2406 are collectively referred to herein as "Uop schedulers 2402, 2404, 2406". In at least one embodiment, the assignor / register renamer 2440 allocates machine buffers and resources that each Uop requires for execution. In at least one embodiment, the assignor / register renamer 2440 renames logical registers to entries in a register bank. In at least one embodiment, the assigner / register renamer 2440 also assigns an entry for each Uop in one of two Uop queues: the main memory Uop queue 2442 for main memory operations and the integer / floating point Uop queue 2444 for operations that do not involve main memory, before the main memory scheduler 2446 and the Uop schedulers 2402, 2404, 2406.In at least one embodiment, the Uop schedulers 2402, 2404, and 2406 determine when a Uop is ready for execution based on the readiness of its dependent input register operand sources and the availability of execution resources that Uops need to complete their operation. In at least one embodiment, the fast scheduler 2402 can schedule on each half of a main clock cycle, while the slow / general floating-point scheduler 2404 and the simple floating-point scheduler 2406 can schedule once per main processor clock cycle. In at least one embodiment, the Uop schedulers 2402, 2404, and 2406 decide on the allocation of ports to schedule Uops for execution.

[0340] In at least one embodiment, the execution block 2411 includes, without limitation, an integer register / bypass network 2408, a floating-point register / bypass network (“FP register / bypass network”) 2410, address generation units (“AGUs”) 2412 and 2414, fast arithmetic logic units (ALUs) (“fast ALUs”) 2416 and 2418, a slow arithmetic logic unit (“slow ALU”) 2420, a floating-point ALU (“FP”) 2422, and a floating-point move unit (“FP move”) 2424. In at least one embodiment, the integer register / bypass network 2408 and the floating-point register / bypass network 2410 are also referred to herein as “registers 2408, 2410”.In at least one embodiment, the AGUs 2412 and 2414, fast ALUs 2416 and 2418, slow ALU 2420, floating-point ALU 2422, and floating-point motion unit 2424 are also referred to herein as "execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424". In at least one embodiment, the execution block 2411 can, without restriction, include any number (including zero) and any type of register banks, bypass networks, address generation units, and execution units in any combination.

[0341] In at least one embodiment, the register networks 2408 and 2410 can be arranged between the Uop schedulers 2402, 2404, and 2406 and the execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424. In at least one embodiment, the integer register bank / bypass network 2408 performs operations on integers. In at least one embodiment, the floating-point register bank / bypass network 2410 performs operations on floating-point numbers. In at least one embodiment, each of the register networks 2408 and 2410 can, without restriction, include a bypass network that can redirect or forward just-completed results that have not yet been written to a register bank to new dependent Uops. In at least one embodiment, the register networks 2408 and 2410 can communicate data with each other.In at least one embodiment, the integer register bank / bypass network 2408 can, without restriction, include two separate register banks: one register bank for thirty-two bits of low-order data and a second register bank for thirty-two bits of high-order data. In at least one embodiment, the floating-point register bank / bypass network 2410 can, without restriction, include 128-bit wide entries, since floating-point instructions typically have operands 64 to 128 bits wide.

[0342] In at least one embodiment, the execution units 2412, 2414, 2416, 2418, 2420, 2422, 2424 can execute instructions. In at least one embodiment, the register networks 2408, 2410 store operand values ​​for integer and floating-point data that require micro-instructions for execution. In at least one embodiment, the processor 2400 can include any number and combination of executi...

Claims

[1] Processor, encompassing: one or more circuits for causing inference or training data of neural networks to be denoised, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks that are to be denoised separately using a corresponding number of neural diffusion networks. [2] Processor according to claim 1, wherein, in order to cause the inference or training data of neural networks to be denoised, the neural networks serve one or more circuits to: Obtaining the training data of neural networks; Identifying the different types of training data within the training data of neural networks; and Training the appropriate number of neural diffusion networks, wherein, to train the appropriate number of neural diffusion networks, the one or more circuits update the weights of the appropriate number of neural diffusion networks using a distribution matching with respect to a multi-stage neural diffusion network and the different types of training data. [3] Processor according to claim 2, wherein, in order to train the corresponding number of neural diffusion networks, one or more circuits further initialize the weights of the corresponding number of neural diffusion networks as if the multi-level neural diffusion networks use a score-matching technique with respect to the multi-level neural diffusion network and the different types of training data. [4] Processor according to one of claims 2 or 3, wherein, to train the corresponding number of neural diffusion networks, the one or more circuits further update the weights of the corresponding number of neural diffusion networks using an adversarial distribution matching with respect to the multi-stage neural diffusion network and the different types of training data. [5] Processor according to any of the preceding claims, wherein, in order to cause the inference or training data of neural networks to be denoised, the neural networks serve one or more circuits to: Receiving an inference request; and Selecting one of the appropriate number of neural diffusion networks based on one of the different types of inference data to be used in order to denoise the inference data of neural networks. [6] Processor according to any of the preceding claims, wherein the inference or training data of neural networks to be denoised are video data. [7] Processor according to any of the preceding claims, wherein the identification of the different types of inference or training data within the inference or training data of neural networks is based on a user configuration. [8] Procedures, comprehensive: To cause inference or training data of neural networks to be denoised, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks that are to be denoised separately using an appropriate number of neural diffusion networks. [9] The method of claim 8, wherein causing the inference or training data of neural networks to be denoised comprises: Obtaining the training data of neural networks; Identifying the different types of training data within the training data of neural networks; and Training the appropriate number of neural diffusion networks, wherein training the appropriate number of neural diffusion networks includes updating the weights of the appropriate number of neural diffusion networks using a distribution matching with respect to a multi-stage neural diffusion network and the different types of training data. [10] Method according to claim 9, wherein the training of the corresponding number of neural diffusion networks further comprises initializing the weights of the corresponding number of neural diffusion networks as if multi-level neural diffusion networks use score matching with respect to the multi-level neural diffusion network and the different types of training data. [11] Method according to one of claims 9 or 10, wherein training the corresponding number of neural diffusion networks further comprises updating the weights of the corresponding number of neural diffusion networks using an opponent distribution matching with respect to the multi-level neural diffusion network and the different types of training data. [12] Method according to any one of claims 8 to 11 wherein causing the inference or training data of neural networks to be denoised comprises: Receiving an inference request; and Selecting one of the appropriate number of neural diffusion networks based on one of the different types of inference data to be used in order to denoise the inference data of neural networks. [13] Method according to any one of claims 8 to 12, wherein the inference or training data of neural networks to be denoised are video data. [14] Method according to claims 8 to 13, wherein the identification of the different types of inference or training data within the inference or training data of neural networks is based on a user configuration. [15] System, encompassing: one or more processors that cause the inference or training data of neural networks to be denoised, based at least in part on identifying different types of inference or training data within the inference or training data of neural networks that are to be denoised separately using a corresponding number of neural diffusion networks; and one or more working memories to store the weights of the neural diffusion networks. [16] System according to claim 15, wherein, in order to cause the inference or training data of neural networks to be denoised, one or more processors shall perform the following: Obtaining the training data of neural networks; Identifying the different types of training data within the training data of neural networks; Training the appropriate number of neural diffusion networks, wherein, to train the appropriate number of neural diffusion networks, one or more processors update the weights of the appropriate number of neural diffusion networks using distribution matching with respect to a multi-level neural diffusion network and the different types of training data. [17] System according to claim 16, wherein, in order to train the corresponding number of neural diffusion networks, the one or more processors further initialize the weights of the corresponding number of neural diffusion networks as if multi-level neural diffusion networks use a score-matching technique with respect to the multi-level neural diffusion network and the different types of training data. [18] System according to one of claims 16 or 17, wherein, to train the corresponding number of neural diffusion networks, the one or more processors further update the weights of the corresponding number of neural diffusion networks using an adversarial distribution matching with respect to the multi-level neural diffusion network and the different types of training data. [19] System according to any one of claims 15 to 18, wherein, in order to cause the inference or training data of neural networks to be denoised, the neural networks serving one or more circuits to: Receiving an inference request; and Selecting one of the appropriate number of neural diffusion networks based on one of the different types of inference data to be used in order to denoise the inference data of neural networks. [20] System according to any one of claims 15 to 19, wherein the inference or training data of neural networks to be denoised are video data.