Graphics processing unit for detecting cheating using neural networks

By using a trained neural network on a graphics processing unit (GPU) to detect illicit information in images, the problem of slow detection and privacy issues in existing technologies is solved, achieving fast and reliable detection of illicit information and reducing false positive reports.

CN114588636BActive Publication Date: 2025-12-09NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111486135.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-07
Filing Date
2021-12-07
Publication Date
2025-12-09
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and reliably detect the presence of illegal information in images used on computing devices, especially in online computer games where privacy issues may arise and detection methods are slow.

Method used

By employing a graphics processing unit (GPU) combined with a trained neural network, the system analyzes rendered images in real time, detects and reports illegal information, reduces false positives using confidence scores, and enhances the detection capabilities of the neural network through training datasets.

Benefits of technology

It enables rapid and reliable detection of illegal information in images, reduces false positive reports, protects user privacy, and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114588636B_ABST
    Figure CN114588636B_ABST
Patent Text Reader

Abstract

Graphics processing units for detecting cheating using neural networks are disclosed, specifically apparatuses, systems, and techniques for detecting cheating in computer games are disclosed. In at least one embodiment, one or more circuits detect cheating by one or more users of a computer game based at least in part on one or more images generated by the computer game using one or more neural networks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] At least one embodiment pertains to a graphics processing unit configured to prevent illicit information from being used in images using a neural network. For example, at least one embodiment pertains to operations encountered in training and using a neural network executing on a graphics processing unit to prevent the use of cheating or other illicit information within images displayed to users of computing devices. BACKGROUND

[0002] In various computing applications, multiple participants can be involved in collective activities in which some information is shared with other participants while other information remains private and available only to some participants. For example, in online computer games, it can not be allowed to have information about the location of an opponent player, their health, armor, ammunition status, etc. However, a gamer can acquire different cheating software to gain access to such illicit (cheating) information. According to a survey, more than 10% of gamers engage in cheating, and more than three quarters of gamers consider cheating to be a serious problem, and are likely to stop playing the game if they suspect that other players are cheating. If based on reports from the gamers, detecting cheating can be a slow and unreliable process. In contrast, a faster approach that uses logs of memory on the computers of the cheaters and / or of the activities of the participants can involve privacy issues. BRIEF DESCRIPTION OF DRAWINGS

[0003] Figure 1A illustrates inference and / or training logic, according to at least one embodiment;

[0004] Figure 1B illustrates inference and / or training logic, according to at least one embodiment;

[0005] Figure 2 illustrates training and deployment of a neural network, according to at least one embodiment;

[0006] Figure 3 is a block diagram of an example computer system in which detection of illicit information in images rendered for display to one or more users can be performed, according to at least one embodiment;

[0007] Figure 4 is a diagram of an image including illicit information rendered by a graphics processing unit on a display of a user device, according to at least one embodiment;

[0008] Figure 5A is a diagram of an enhanced illicit training image having a reduced amount of illicit information for effective training of a neural network to detect illicit information in images rendered for display to one or more users, according to at least one embodiment;

[0009] Figure 5B is a diagram of an augmented real training image including a pasted artifact for effective training of a neural network to detect illicit information in images rendered for display to one or more users, according to at least one embodiment;

[0010] Figure 6 is a diagram of a system including one or more neural network models for detecting illicit information in images rendered by a graphics processing unit on a display of a user device, according to at least one embodiment;

[0011] Figure 7 is a flow diagram of an example method of detecting illicit information in images rendered by a graphics processing unit on a display of a user device using one or more neural networks, according to at least one embodiment;

[0012] Figure 8 is a flow diagram of an example method of training one or more neural networks to detect illicit information in images rendered by a graphics processing unit on a display of a user device, according to at least one embodiment;

[0013] Figure 9 shows an example data center system, according to at least one embodiment;

[0014] Figure 10A shows an example of an autonomous vehicle, according to at least one embodiment;

[0015] Figure 10B shows an example of a camera position and field of view of an autonomous vehicle, according to at least one embodiment; Figure 10A

[0016] Figure 10C is a block diagram illustrating an example system architecture of an autonomous vehicle, according to at least one embodiment; Figure 10A

[0017] Figure 10D is a diagram illustrating a system for communication between one or more cloud-based servers and an autonomous vehicle, according to at least one embodiment; Figure 10A

[0018] Figure 11 is a block diagram illustrating a computer system, according to at least one embodiment;

[0019] Figure 12 is a block diagram illustrating a computer system, according to at least one embodiment;

[0020] Figure 13 shows a computer system, according to at least one embodiment; ​​​

[0021] Figure 14 A computer system, in accordance with at least one embodiment, is shown;

[0022] Figure 15A A computer system, in accordance with at least one embodiment, is shown;

[0023] Figure 15B A computer system, in accordance with at least one embodiment, is shown;

[0024] Figure 15C A computer system, in accordance with at least one embodiment, is shown;

[0025] Figure 15D A computer system, in accordance with at least one embodiment, is shown;

[0026] Figure 15E and Figure 15F A shared programming model, in accordance with at least one embodiment, is shown;

[0027] Figure 16 An exemplary integrated circuit and associated graphics processor, in accordance with at least one embodiment, is shown.

[0028] Figures 17A-17B An exemplary integrated circuit and associated graphics processor, in accordance with at least one embodiment, is shown.

[0029] Figures 18A-18B Additional exemplary graphics processor logic, in accordance with at least one embodiment, is shown;

[0030] Figure 19 A computer system, in accordance with at least one embodiment, is shown;

[0031] Figure 20A A parallel processor, in accordance with at least one embodiment, is shown;

[0032] Figure 20B A partition unit, in accordance with at least one embodiment, is shown;

[0033] Figure 20C A processing cluster, in accordance with at least one embodiment, is shown;

[0034] Figure 20D A graphics multiprocessor, in accordance with at least one embodiment, is shown;

[0035] Figure 21 A multi-GPU system, in accordance with at least one embodiment, is shown;

[0036] Figure 22 A graphics processor, in accordance with at least one embodiment, is shown;

[0037] Figure 23 is a block diagram illustrating a processor microarchitecture for a processor, in accordance with at least one embodiment;

[0038] Figure 24 a deep learning application processor is shown, in accordance with at least one embodiment;

[0039] Figure 25 is a block diagram illustrating an example neuromorphic processor, in accordance with at least one embodiment;

[0040] Figure 26 at least a portion of a graphics processor is shown, in accordance with one or more embodiments;

[0041] Figure 27 at least a portion of a graphics processor is shown, in accordance with one or more embodiments;

[0042] Figure 28 at least a portion of a graphics processor is shown, in accordance with one or more embodiments;

[0043] Figure 29 is a block diagram illustrating a graphics processing engine of a graphics processor, in accordance with at least one embodiment;

[0044] Figure 30 is a block diagram illustrating at least a portion of a graphics processor core, in accordance with at least one embodiment;

[0045] Figures 31A-31B thread execution logic is shown, in accordance with at least one embodiment, including an array of processing elements of a graphics processor core.

[0046] Figure 32 a parallel processing unit (“PPU”) is shown, in accordance with at least one embodiment;

[0047] Figure 33 a general processing cluster (“GPC”) is shown, in accordance with at least one embodiment;

[0048] Figure 34 a memory partition unit of a parallel processing unit (“PPU”) is shown, in accordance with at least one embodiment;

[0049] Figure 35 a streaming multiprocessor is shown, in accordance with at least one embodiment.

[0050] Figure 36 is an example dataflow graph of an advanced computing pipeline, in accordance with at least one embodiment;

[0051] Figure 37 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, in accordance with at least one embodiment. DETAILED DESCRIPTION

[0052] Inference and training logic

[0053] Figure 1A Inference and / or training logic 115 is shown performing inference and / or training operations associated with one or more embodiments.

[0054] In at least one embodiment, inference and / or training logic 115 can include, without limitation, code and / or data storage 101 for storing forward and / or output weight and / or input / output data, and / or other parameters of neurons or layers of a neural network being trained and / or configured for inference in aspects of one or more embodiments. In at least one embodiment, training logic 115 can include or be coupled to code and / or data storage 101 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic unit(s) (ALUs) or simply ALUs). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 101 stores weight parameters and / or input / output data of each layer of a neural network being trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 101 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0055] In at least one embodiment, any portion of code and / or data storage 101 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 101 can be cache memory, dynamic random- addressable memory (“DRAM”), static random- addressable memory (“SRAM”), non-volatile memory (such as flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 101 is internal or external to a processor, e.g., or comprised of DRAM, SRAM, flash or some other storage type, can depend on available storage space to store on-chip or off-chip, latency requirements of performing training and / or inference functions, batch size of data used in inference and / or training of a neural network, or some combination of these factors.

[0056] In at least one embodiment, inference and / or training logic 115 can include, without limitation, code and / or data storage 105 to store backward and / or output weights and / or input / output data for neurons or layers of a neural network trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 105 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, training logic 115 can include or be coupled to code and / or data storage 105 to store graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic units (ALUs)).

[0057] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 105 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 105 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 105 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash) or other storage. In at least one embodiment, whether code and / or data storage 105 is internal or external to a processor, e.g., whether made up of DRAM, SRAM, Flash, or some other storage type, is a choice dictated by design parameters including whether available storage is on-chip or off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data being used in inferencing and / or training of a neural network, or some combination of these factors.

[0058] In at least one embodiment, code and / or data storage 101 and code and / or data storage 105 can be separate storage structures. In at least one embodiment, code and / or data storage 101 and code and / or data storage 105 can be the same storage structure. In at least one embodiment, code and / or data storage 101 and code and / or data storage 105 can be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 101 and code and / or data storage 105 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory.

[0059] In at least one embodiment, inference and / or training logic 115 can include, without limitation, one or more arithmetic logic units (“ALUs”) 110 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom. Results of such operations can result in activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 120, which are functions of input / output and / or weight parameter data stored in code and / or data storage 101 and / or code and / or data storage 105. In at least one embodiment, activations stored in activation storage 120 are generated by ALUs 110 in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALUs 110, where weight values stored in code and / or data storage 105 and / or code and / or data storage 101 are used as operands having other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 105 or code and / or data storage 101 or other on-chip or off-chip storage.

[0060] In at least one embodiment, one or more ALUs 110 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment one or more ALUs 110 can be external to a processor or other hardware logic device or circuit using them (e.g., a co-processor). In at least one embodiment, one or more ALUs 110 can be included within execution units of a processor, or otherwise included in a group of ALUs accessible by execution units of a processor, which can be within a same processor or distributed between different types of processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 101, code and / or data storage 105, and activation storage 120 can share a processor or other hardware logic device or circuit, while in another embodiment they can be in different processors or other hardware logic devices or circuits, or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 120 can be included with other on-chip or off-chip data storage, including L1, L2, or L3 cache(s) of a processor or system memory. Moreover, inference and / or training code can be stored with other code accessible to a processor or other hardware logic or circuit, and can be fetched and / or processed using fetch, decode, schedule, execution, exit, and / or other logic circuits of a processor.

[0061] In at least one embodiment, activation storage 120 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., Flash) or other storage. In at least one embodiment, activation storage 120 can be entirely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, whether activation storage 120 is internal or external to a processor, or contains DRAM, SRAM, Flash, or other storage types, can be chosen depending on storage available on-chip or off-chip, latency requirements of training and / or inference functions, batch size of data used in inference and / or training neural networks, or some combination of these factors.

[0062] In at least one embodiment, Figure 1A Inference and / or training logic 115 shown in FIG. 1A can be used in conjunction with a special-purpose integrated circuit (‘ASIC’) such as a Tensor Processing Unit (‘TPU’) from Google, an inference processing unit (‘IPU’) from Graphcore TM AI accelerator, or a ‘Lake Crest’ processor from Intel Corp. In at least one embodiment, inference and / or training logic 115 can be used in a system-on-a-chip (‘SoC’) system that integrates hardware processing circuits (e.g., one or more of: application-specific integrated circuit (‘ASIC’), central processing unit (‘CPU’), graphics processing unit (‘GPU’), tensor processing unit (‘TPU’), etc.) with circuitry and / or storage for acceleration of software processes. Figure 1AThe inference and / or training logic 115 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware or other hardware (such as field programmable gate array (“FPGA”)).

[0063] Figure 1B An inference and / or training logic 115 according to at least one embodiment is illustrated. In at least one embodiment, the inference and / or training logic 115 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 1B The inference and / or training logic 115 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM Inference processing units (IPUs) or from Intel Corp. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 1B The inference and / or training logic 115 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 115 includes, but is not limited to, code and / or data storage 101 and code and / or data storage 105, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 1B In at least one embodiment shown, each of code and / or data storage 101 and code and / or data storage 105 is associated with dedicated computing resources (e.g., computing hardware 102 and computing hardware 106), respectively. In at least one embodiment, each of computing hardware 102 and computing hardware 106 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data storage 101 and code and / or data storage 105, respectively, and the results of the function execution are stored in active storage 120.

[0064] In at least one embodiment, each of code and / or data stores 101 and 105 and corresponding compute hardware 102 and 106, respectively, correspond to different layers of a neural network, such that activations resulting from one storage / compute pair 101 / 102 of code and / or data store 101 and compute hardware 102 are provided as input to next “storage / compute pair 105 / 106” of code and / or data store 105 and compute hardware 106 to reflect a conceptual organization of a neural network. In at least one embodiment, each storage / compute pair 101 / 102 and 105 / 106 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in inference and / or training logic 115 after or in parallel with storage compute pairs 101 / 102 and 105 / 106.

[0065] Neural network training and deployment

[0066] Figure 2 Training and deployment of a deep neural network is shown, in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 206 is trained using a training dataset 202. In at least one embodiment, training framework 204 is a PyTorch framework, while in other embodiments, training framework 204 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, training framework 204 trains untrained neural network 206 and enables it to train using processing resources described herein to generate a trained neural network 208. In at least one embodiment, weights can be chosen randomly or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.

[0067] In at least one embodiment, an untrained neural network 206 is trained using supervised learning, where a training dataset 202 includes inputs paired with desired outputs for inputs, or where a training dataset 202 includes inputs with known outputs and the output of the neural network 206 is manually graded. In at least one embodiment, an untrained neural network 206 is trained in a supervised manner and processes an input from a training dataset 202 and compares a resulting output to a set of expected or desired outputs. In at least one embodiment, an error is then propagated back through the untrained neural network 206. In at least one embodiment, a training framework 204 adjusts weights that control the untrained neural network 206. In at least one embodiment, a training framework 204 includes tools for monitoring how well an untrained neural network 206 is converging towards a model (e.g., a trained neural network 208) that is suitable for generating correct answers (e.g., results 214) based on input data (e.g., new datasets 212). In at least one embodiment, a training framework 204 trains an untrained neural network 206 repeatedly while adjusting weights to improve outputs of the untrained neural network 206 using a loss function and an adjustment algorithm (e.g., stochastic gradient descent). In at least one embodiment, a training framework 204 trains an untrained neural network 206 until the untrained neural network 206 reaches a desired accuracy. In at least one embodiment, a trained neural network 208 can then be deployed to implement any number of machine learning operations.

[0068] In at least one embodiment, an untrained neural network 206 is trained using unsupervised learning, where an untrained neural network 206 attempts to train itself using unlabeled data. In at least one embodiment, an unsupervised learning training dataset 202 will include input data without any associated output data or “ground truth” data. In at least one embodiment, an untrained neural network 206 can learn groupings within a training dataset 202 and can determine how individual inputs relate to the untrained dataset 202. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 208 that is capable of performing operations useful for reducing a dimensionality of new datasets 212. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in new datasets 212 that deviate from a normal pattern of new datasets 212.

[0069] In at least one embodiment, semi-supervised learning can be used, which is a technique in which a mix of labeled and unlabeled data is included in training dataset 202. In at least one embodiment, training framework 204 can be used to perform incremental learning, for example, through a transferred learning technique. In at least one embodiment, incremental learning enables trained neural network 208 to adapt to new dataset 212 without forgetting knowledge that was imprinted into trained neural network 208 during initial training.

[0070] Detecting cheating with neural networks

[0071] In at least one embodiment, an application publisher (e.g., a game publisher or developer) can maintain ownership (e.g., copyright and control) of images (e.g., game images) generated by an application (e.g., game software) installed on a user’s computer. In at least one embodiment, an image can be subject to analysis by detection software and / or other software before, in parallel with, and / or after the image is displayed to a user (e.g., a gamer). In at least one embodiment, one or more trained neural networks at local (pixel-level) and / or global (image-level) scales can be used to detect the presence of a cheating image, a cheating plug-in, or any other falsified information included in a genuine image displayed to a user. In at least one embodiment, an image can be generated by a graphics processing unit (GPU), which can be a specialized processor used to efficiently render graphical output (e.g., three-dimensional images of a game environment and scenery) on a user’s display. In at least one embodiment, a GPU can operate in conjunction with a central processing unit (CPU) of a user’s computer; the CPU can perform various tasks that require fast serial processing. For example, a CPU can determine a gamer’s path, a line of fire, a hit or miss on a target, track a detailed inventory, communicate with other gamers’ computing devices (e.g., over the Internet or other network), etc. In at least one embodiment, a GPU can perform tasks that are amenable to parallel processing, such as performing matrix operations (e.g., pixel-level rotations and translations) that determine how a game environment should be displayed to a user, and generating a final (combined) image to be provided to a user. In at least one embodiment, based on instructions received from a CPU, a GPU can also add various ancillary data to a final image, such as information about available weapons, ammunition, current health status, and so on. In at least one embodiment, a CPU can be executing instructions generated by game software. In at least one embodiment, a user can also instantiate cheating software that is executed along with game software, can cause a CPU to output modified instructions to a GPU, and cause a GPU to render (or directly write into GPU memory) an image that includes illicit or falsified information in the form of a window, a plug-in, a semi-transparent overlay, and so on.

[0072] In at least one embodiment, a GPU or any other processing unit that renders images to a user’s display can provide the rendered images to one or more trained neural networks. For example, an image can be retrieved from a buffer that temporarily stores the image prior to rendering the image on a display. In at least one embodiment, a trained neural network can be used to determine whether an image contains illicit (e.g., cheating or fraudulent) information. In at least one embodiment, if a trained neural network determines that an image includes cheating information, a GPU (or other processing logic) can generate a report and provide (e.g., over a network) the generated report to a server (e.g., a game server or a publishing server of a game provider) to notify the server of the detected instance of cheating. A trained neural network can also provide a confidence score that indicates a level of confidence in a determination as to whether an image includes or does not include cheating information. In at least one embodiment, a provided confidence score can be used to reduce a number of false positive instances of cheating detection. Specifically, if a level of confidence is low (e.g., below a first threshold), an instance that is suspected of cheating can not be reported. In at least one embodiment, if a level of confidence (e.g., across a particular set of images) falls below a certain value (e.g., a second threshold), a server can take this as an indication that new cheating software has become available or that a new or modified application has been provided (e.g., a new game scenario or has been added to a game) and that the neural network is no longer able to be used to provide accurate determinations and should be retrained using recent images (e.g., images generated by the new cheating software).

[0073] While examples used herein relate to detecting cheating information in game images, similar devices and techniques can be used to detect illicit (or fraudulent) information in images in any other context in which a user can be able to exceed a scope of authorized access. In at least one embodiment, such contexts can include, for example, accessing financial information, medical information, insurance information, classified information, or any other secure and / or private information. In at least one embodiment, such contexts can also include a paid subscription, digital information protected by access prevention measures (passwords, signatures, multi-factor authentication measures, etc.), online exams in which examinees are not allowed to possess certain information or access certain resources (e.g., content of a user’s computer). In at least one embodiment, illicit information can include any type of information that a user is not allowed or not expected to have legal access to. In at least one embodiment, illicit information can be any information access that can be restricted by law, contract, employment regulations, or any other type of rule or agreement.

[0074] Figure 3is a block diagram of an exemplary computer system 100 in accordance with at least one embodiment in which detection of illicit information in images rendered for display to one or more users can be performed. In at least one embodiment, as shown in Figure 3 exemplary computer system 100 can include a publishing server 302 (e.g., a game publishing server), a game server 304, one or more user machines 310, 312, 313,..., a training server 340, and a training image repository 350, some or all of which can be connected to a network 308. In at least one embodiment, network 308 can be a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), or a combination thereof. In at least one embodiment, a user machine (e.g., user machine 310) can connect to network 308 via a network adapter 312. In at least one embodiment, user machine 310 can include a memory 314, a processing device (e.g., CPU 316), and an input / output (I / O) module 318. CPU 316 can be any device capable of executing instructions for coded arithmetic, logic, or I / O operations. In at least one embodiment, memory device 314 can be a volatile or non-volatile memory such as a random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), or any other device capable of storing data. In at least one embodiment, CPU 316 can follow a von Neumann architecture model and can include one or more arithmetic logic units (ALUs), a control unit, and a plurality of registers. In at least one embodiment, CPU 316 can be a single core processor capable of executing one instruction at a time, or can be a multi-core processor capable of executing multiple instructions simultaneously. In at least one embodiment, CPU 316 can be implemented as a single integrated circuit, two or more integrated circuits, or can be a component of a multi-chip module. In at least one embodiment, user machine 310 can also include one or more I / O devices 319 supported by I / O module 318. In at least one embodiment, I / O devices 319 can include a keyboard, a mouse device, a pointer, a touchscreen, a touchpad, a camera, a microphone, or any other type of sensor capable of detecting user input.

[0075] In at least one embodiment, the user machine 310 may include a GPU 320. The GPU 320 may have multiple cores 322, each capable of executing multiple threads simultaneously. The GPU cores 322 may access GPU memory 324, which may include private (thread-specific) registers, shared registers, caches (e.g., L1, L2, L3, etc.), and / or other memory devices. In at least one embodiment, the GPU memory 324 may include non-volatile memory storing a neural network (NN) model 325 for detecting illegal images (e.g., images containing at least some illegal information that the user is not authorized to possess). The NN model 325 may be executed by one or more cores 322. In at least one embodiment, the NN model 325 may be executed by... Figure 3 The processing device, not shown, executes independently or in conjunction with core 322. In at least one embodiment, GPU 320 may include buffer 326 for storing game (or any other) images prepared by GPU 320 for display on display 329. In at least one embodiment, buffer 229 may be a first-in-first-out (FIFO) buffer. In at least one embodiment, the image may be any raster image, including any pixel-based digital format image such as, but not limited to, BMP, TIFF, CIB, JPEG, DIMAP, GIF, NITF, PNG, etc. In at least one embodiment, the image may be a vectorized image, in which objects are depicted via mathematical relationships, such as by defining lines, curves, predefined shapes, or any combination thereof. In at least one embodiment, instead of NN model 325 or other than NN model 325, GPU memory 324 may also include one or more machine learning models different from the trained neural network, such as tree classifiers, Bayesian classifiers, regression classifiers, stochastic gradient descent classifiers, K-nearest neighbor classifiers, or other algorithms designed to detect illegal information.

[0076] In at least one embodiment, NN model 325 can be trained by training server 340. In at least one embodiment, training server 340 can be (and / or include) a rackmount server, a router computer, a personal computer, a laptop computer, a tablet computer, a desktop computer, a media center, or any combination thereof. In at least one embodiment, training server 340 can include a training engine 342. In at least one embodiment, training engine 342 can build machine learning models, such as NN model 325. In at least one embodiment, NN model 325 can be trained by training engine 342 using training data that includes training inputs 347 and corresponding target outputs 348. In at least one embodiment, target outputs 348 can include correct associations (mapping data 349) of training images with presence or absence of illicit information in training images. In at least one embodiment, training engine 342 can derive patterns in training inputs 347 that map training inputs 347 to target outputs 348 (which can include correct associations to be predicted during training), and train NN model 325 to capture such patterns. In at least one embodiment, patterns can be subsequently used by trained NN model 325 for future inferences performed on new images. For example, after accessing one or more new images in buffer 326, trained NN model 325 can be able to determine presence or absence of illicit information in new images. In at least one embodiment, training server 340 can be able to retrain NN model 325 when new information (e.g., regarding availability of new cheating software, changes / updates to existing games, or emergence of new games or new game scenarios) becomes available. In at least one embodiment, after initial training or subsequent retraining, training server 340 can be able to broadcast modified parameters of NN model 325 to different GPUs (such as GPU 320) of user machines (such as user machine 310) to update NN models installed thereon.

[0077] In at least one embodiment, training images 352 can be stored in a training image repository 350, which training server 340 can access directly or via network 308. In at least one embodiment, training image repository 350 can be a persistent storage device capable of storing training images as well as metadata for stored training images. In at least one embodiment, training image repository 350 can be hosted by one or more storage devices, such as main memory, disk, tape or hard drives based on magnetic or optical storage, NAS, SAN, etc. While depicted as separate from training server 340, in at least one embodiment, image repository 350 can be part of training server 340. In at least one embodiment, image repository 350 can be a network-attached file server, while in other embodiments, image repository 350 can be some other type of persistent storage device, such as an object-oriented database, a relational database, etc., which can be hosted by one or more different machines in communication with training server 340 via network 308.

[0078] In at least one embodiment, training images 352 can be generated by replaying games using game logs 306 stored on game servers 304. Specifically, game logs 306 can be used to replay game contexts, re-render images using game software, and add illicit information to produce a set of training images. In at least one embodiment, illicit information can be added to truly (not illicit) game images (or images from other sources). In at least one embodiment, different overlays (e.g., obtained from publicly available video and graphics resources) can be added to training images to train NN model 325, thereby disregarding overlays that do not include illicit information.

[0079] In at least one embodiment, training images 352 can be generated using cheat software 346, such as any cheat software that can be used by a cheating user (e.g., a gamer) or any other unauthorized user of illicit information. In at least one embodiment, as described in more detail below, some of training images 352 can be genuine (non-cheat) images generated by legitimate software (e.g., game software 345). In at least one embodiment, a genuine image can be an image that does not have any illicit information. In at least one embodiment, a genuine image can be generated by any legitimate software that is different from game software, such as financial software, insurance software, accounting software, healthcare software, etc. In at least one embodiment, training images 352 can include images that include illicit information (training cheat images) generated by cheat software 346 and added to genuine images by training image generator 344. In at least one embodiment, training image generator 344 can vary an amount of cheat information in a training cheat image by combining (superimposing, appending, overlaying, replacing, etc.) portions of illicit information generated by cheat software 346 with portions of non-cheat images generated by game software 345. In at least one embodiment, this can be done by training NN model 325 to detect cheat images that even have small amounts of illicit information, in order to prevent NN model 325 from focusing on the most obvious cheat information, which can be turned off by sophisticated cheaters who are seeking only a small competitive advantage and who configure their cheat software to provide only a small fraction of available cheat information. In at least one embodiment, to train NN model 325 to ignore paste artifacts (caused by pasting) within training cheat images, training non-cheat images can be supplemented with similar artifacts. More specifically, non-illicit portions (e.g., portions without illicit information) of cheat images can be pasted into similar portions of training non-cheat images. As a result, training cheat images and training non-cheat images can differ only in the presence of illicit information (which NN model 325 is trained to detect) while having the same or similar paste artifacts (which NN model 325 is therefore trained to ignore).

[0080] In at least one embodiment, GPU 320 may include a certification module 328 for certifying that one or more users of user machine 310 are using game hardware (e.g., GPU 320) capable of detecting cheating. In at least one embodiment, certification module 328 may transmit a message to publishing server 302 and / or game server 304 indicating that GPU 320 is a certified anti-cheat GPU. In at least one embodiment, the transmitted message may identify GPU 320 by its unique identifier (ID), which may be the media access control (MAC) address of GPU 320, or some other form of identifier, such as a unique ID assigned to a specific GPU 320 installed on user machine 310. In at least one embodiment, the message may be transmitted at the start of a user's game session. In at least one embodiment, additional messages may be transmitted while a user's game session is in progress (e.g., periodically) and / or when a user's game session ends. Such additional messages may include reports indicating whether any instances of cheating have been detected.

[0081] In at least one embodiment, combined Figures 3 to 8 The described system and method can be implemented in a device other than the GPU 320, such as a display driver board (DDB) that renders images on a display 329. In at least one embodiment, the NN model 325 can reside on the DDB and can access the images rendered by the DDB on the display 329. In at least one embodiment, combined with Figures 3 to 8 The described system and method can be used to detect cheating in images rendered by a device different from the detection device (e.g., the detection device may be separate from GPU 320). In at least one embodiment, the detection device (e.g., GPU 320) may have access to an image being rendered on display 329 by another device and / or application (e.g., DDB). In at least one embodiment, cheating software may display illegal information on display 329 separately from the displayed image, e.g., in a separate window, using an overlay generated by a display controller different from GPU 320, etc. In at least one embodiment, GPU 320 (or another detection device, e.g., a camera pointed at display 329) may retrieve the image rendered on display 329 (or stored for subsequent rendering on display 329) for example via the display driver board, and perform the detection of cheating (or other illegal information) in a manner similar to how illegal information is detected when GPU 320 renders the final image (as described below).

[0082] Figure 4 This is a schematic diagram of an image according to at least one embodiment, the image including illegal information rendered by a graphics processing unit on the display of a user's device. Figure 4 depicted is an illegal image 400 of an example game environment, which can include depictions of opponent players, structures, buildings, obstacles, paths, natural items, and any other objects that can exist in a particular game. Although Figure 4 An environment of a computer game is shown, but in at least one embodiment, an image of any other type of information— financial, educational, medical, etc.— can be displayed. In at least one embodiment, the displayed environment (e.g., game environment) can include illegal information that provides an unfair advantage to the user. For example, Figure 4 The depicted game environment can include global illegal information 402, such as information related to the game as a whole, e.g., the number and characteristics of opponent players, a map of the game environment that the user should not have access to, pointers to ammunition and medical supplies, etc. In at least one embodiment, the game environment can also include various local illegal information, such as information about a particular player or particular object on the screen, e.g., information 404-410. For example, local illegal information 404 can include information associated with a representation (e.g., depiction) of an opponent player 420 of health / ammunition / status, local illegal information 406 can indicate to the user the presence of an opponent hidden behind a wall 430, local illegal information 408 can indicate to the user the presence of a communication located under a structure 440, local illegal information 410 can indicate to the user the presence of an opponent within a structure 450, etc. In at least one embodiment, the illegal information is added by CPU 316 (and possibly rendered by GPU 320 according to instructions from CPU 316), which executes software installed on user machine 310 and stored in memory 314. The illegal information can be pasted, added, overlaid, etc. on the true image produced by game software 345. In at least one embodiment, the illegal information can be presented in any format, e.g., text, rasterized pictures, vectorized objects, video, animation, etc., or any combination thereof. In at least one embodiment, GPU 320 can be the last device to render the image on display 329. Thus, GPU 320 can have possession of the graphical information that is actually displayed to the user, as no additional components process the displayed image, e.g., no components add illegal information to the true image or intercept illegal information (possibly removing illegal information after it is displayed to the user).

[0083] In at least one embodiment, cheating software 346 installed on training server 340 can generate images similar to the images viewed by the cheating player (or any other beneficiary or user of illegal information). For example, cheating software 346 can be the same or similar to the cheating software used by the cheating user. In at least one embodiment, the unmodified illegal training images produced by cheating software 346 (similar to the images viewed by the cheating user) can be stored in memory 334 and / or transmitted to user machine 310. In at least one embodiment, the illegal information can be added to the true images by CPU 316 (and possibly rendered by GPU 320 according to instructions from CPU 316), which executes software installed on user machine 310 and stored in memory 314. Figure 4The images shown (e.g., the images shown in FIG. 5) can be used to train the NN model 325. In at least one embodiment, at least some of the training images can be based on images produced by the cheating software 346 but additionally modified by the training image generator 344. More specifically, the amount of illicit information present in the training images can vary, e.g., be increased or decreased, as compared to the amount of illicit information in the original images produced by the cheating software 346. In at least one embodiment, the increased amount can be used to facilitate initial training of the NN model 325 (e.g., to speed up initial training), while the decreased amount can be used to train the NN model 325 to even detect small amounts of presence of illicit information. In at least one embodiment, the NN model 325 trained using varying amounts of illicit information can effectively identify even discreet players who, in an effort to remain undetected, seek a relatively modest advantage over competitors by limiting the amount of illicit information displayed. In at least one embodiment, the NN model 325 can be trained using images having progressively decreasing amounts of illicit information.

[0084] In at least one embodiment, the varying amounts of illicit information in the training images can be implemented using the enhanced techniques and processes described below. Figure 5A is a schematic diagram of an enhanced illicit training image 500 having a reduced amount of illicit information, in accordance with at least one embodiment, for effectively training a neural network to detect illicit information in images rendered for display to one or more users. In Figure 5A depicted in FIG. 5 is an image that can be generated from a similar image as Figure 4The image 400 depicted is an image of a game environment obtained from an illicit image. In at least one embodiment, some illicit information in the image 400 is preserved. For example, shown as preserved are local illicit information 406 (indicating the presence of an opponent), local illicit information 408 (indicating underhanded communication), and local illicit information 410 (also indicating the presence of an opponent). In at least one embodiment, the amount of illicit information in the training image can be limited, for example, by modifying the settings of the cheating software. In at least one embodiment, some illicit information can be removed. For example, that which can be removed are global illicit information 402 as well as local illicit information 404 (showing the health / ammo / status / etc. of an opponent). In at least one embodiment, the removal of a portion of the illicit information can be performed if a genuine / non-illicit (training) image pair is available. In at least one embodiment, the genuine training image can be generated by the game software and the illicit training image can be generated by the cheating software. In at least one embodiment, the removal of the illicit information can be performed by replacing the areas of the illicit training image with corresponding areas from a genuine image (e.g., a genuine image of the same or similar environment produced by the game software 345). In at least one embodiment, a replacement area 502 can cover (replace, paste, overlay, superimpose, etc.) the global illicit information 402 and a replacement area 504 can cover the local illicit information 404.

[0085] In at least one embodiment, the obtained enhanced illicit training image 500 can have a pasting / overlaying / superimposing / etc. artifact present at the location where the pasting or replacement has occurred. In at least one embodiment, the genuine images produced by the game software 345 do not have exact pixel-to-pixel synchronization with the illicit images produced by the cheating software (alone or in combination with the game software 345). Similarly, in at least one embodiment, the illicit images and / or the genuine images can undergo compression, which can also distort the pixel-to-pixel synchronization. This can result in a pasting artifact being present in the enhanced illicit image. The pasting artifact in the illicit training image can cause the NN model 325 to identify the illicit image by focusing on the artifact (rather than focusing on the illicit information) when analyzed pixel-by-pixel by the NN model 325, even when the human eye cannot see it. In at least one embodiment, to prevent such misperception by the NN model 325, a similar artifact can be pasted into the genuine training image.

[0086] Figure 5B is a schematic diagram of an enhanced genuine training image 550 including a pasting artifact, in accordance with at least one embodiment, for effectively training a neural network to detect illicit information in images rendered for display to one or more users. In Figure 5BImage 550 depicts a game environment that can be obtained from a genuine image after one or more pastes / replacements have been made. For example, shown as added are replacement regions 554, 556, and 558. Adding replacement regions introduces pasting artifacts into the genuine image, and - with the presence of pasting artifacts in both genuine and illegitimate training images - causes the NN model 325 to ignore the pasting artifacts and distinguish between genuine and illegitimate training images based on the presence of illegitimate information. In at least one embodiment, replacement regions 554, 556, and 558 are taken from illegitimate images, e.g., illegitimate images corresponding to the same computer game, setting, player, scenery, etc. In at least one embodiment, replacement regions 554, 556, and 558 are taken from such regions of illegitimate images that are free of illegitimate information, in order to prevent the augmented genuine training image 550 from being contaminated with pixels corresponding to illegitimate information.

[0087] In at least one embodiment, to produce augmented genuine training images and augmented illegitimate training images, the training engine 342 can launch (automatically or in response to a command by a developer) game software 345 to produce genuine images associated with the game. Additionally (separately or in parallel), in at least one embodiment, the training images (and / or a developer) can launch cheat software 346 to produce (separately or in cooperation with game software 345) one or more underlying illegitimate images, which can be unmodified illegitimate images produced by game software. In at least one embodiment, game software 345 and cheat software 346 can generate images corresponding to the same or similar game contexts. In at least one embodiment, one instance of game software 345 (a first context) can start from a particular initial position of a player, weapon condition, health, ammo, etc. Meanwhile, in at least one embodiment, cheat software 346 can launch another instance of game software 345 (a second context, which can be a copy of the first context), and can start executing the same (or similar) game context at the same (or similar) time with the same (or similar) player, weapon, health, ammo, etc. In at least one embodiment, the training images generator 344 can be used to sample both contexts at about the same (or at substantially the same) time to extract images (e.g., snapshots) provided by both instances of the game. In at least one embodiment, the training server 340 can have separate GPUs that generate images for separate game contexts.

[0088] In at least one embodiment, training image generator 344 can be used to remove (e.g., at direction of a developer or automatically) some illegal information from a base illegal image (generated by a second context) to varying degrees. In at least one embodiment, a base illegal image can be an unmodified illegal image produced by game software. In at least one embodiment, one base illegal image can be used to produce multiple illegal training images, e.g., where 10%, 20%, 50%, 70%, etc. of illegal information is removed. In at least one embodiment, additional illegal information can be added to a base image to generate illegal training images with 110%, 120%, etc. of illegal information. In at least one embodiment, some of the illegal information from the same base image can be pasted into additional locations within the base image. In at least one embodiment, illegal information can be taken from other similar images, e.g., earlier or later images within the same game context or images obtained from other game contexts.

[0089] In at least one embodiment, training image generator 344 can also be used to add (e.g., paste, overlay, superimpose, etc.) illegal portions of illegal images to authentic images (e.g., generated by a first context). In at least one embodiment, the amount, size, location of insertion into authentic images can be comparable to the amount of removal (or insertion) into illegal images, such that the augmented illegal images have about the same (or comparable) amount and type of artifacts caused by insertion as the augmented authentic images.

[0090] In at least one embodiment, as a result of image manipulations described in previous paragraphs, training server 340 can generate one or more training images 352, including one or more augmented illegal training images (such as image 500) and one or more augmented authentic training images (such as image 550), which can be used by training engine 342 to train one or more NN models 325. In at least one embodiment, augmentation of authentic and illegal training images can be done automatically, without requiring detailed input from a developer. For example, a developer can label regions in a base illegal image that include illegal information, and training image generator 344 can randomly select portions of illegal information to remove from the base illegal image, or randomly select portions of a base authentic image to replace with non-illegal regions of an illegal image.

[0091] Figure 6 is a schematic diagram of a system 600 including one or more neural network models for detecting illegal information in images rendered by a graphics processing unit on a display of a user’s device, in accordance with at least one embodiment. In at least one embodiment, the operations performed at Figure 6depiction of the training of one or more NN models (or, hereinafter, neural networks). In at least one embodiment, Figure 6 The depicted neural network can be Figure 3 NN model 325. In at least one embodiment, after training, the trained neural network can be incorporated into GPU 210, which can be included in user machine 310. In at least one embodiment, GPU 210 can perform a dual function: rendering images on display 329 while (before, after) applying the neural network to the rendered images to detect the presence of illicit (e.g., cheating) information therein. In at least one embodiment, image 602 can be input into system 600. In at least one embodiment, image 602 can be one of a training image 352 (used during a training phase) or a new image (during an inference phase). In at least one embodiment, image 602 can be input into first network 610. In at least one embodiment, first network 610 can be trained to classify image 602 at a local level. The local level can refer to individual pixels of image 602. Alternatively, the local level can refer to super-pixels (aggregations of pixels), which can include multiple pixels, such as a 2x2, 4x4, 16x16, 8x16, or any other super-pixel size block. Hereinafter, in embodiments using a super-pixel representation of images, references to one or more pixels are also to be understood to apply to one or more super-pixels.

[0092] In at least one embodiment, an output of first network 610 can be a feature map 620 of local scores 622 indicative of a probability that a particular pixel (or super-pixel) contains (or belongs to) illicit information. In at least one embodiment, a local score 622 can be a number in a range from 0 to 1 (e.g., a probability) or in any other pre-defined range (e.g., 0 to 10, related to a probability, although not necessarily equal to it), where a value close to one end of the range (e.g., 0) indicates that a pixel is very unlikely to belong to a group of pixels representing illicit information, and a value close to the other end of the range (e.g., 1) indicates that a pixel is very likely to belong to a group of pixels representing illicit information. In at least one embodiment, input image 602 can first be converted into a digital representation of a pixel map of image 602. In at least one embodiment, a number of pixels can depend on a resolution of image 602, e.g., an image can have 2048 x 1024 pixels, 1680 x 1050 pixels, or any other resolution. In at least one embodiment, each pixel can be characterized by one or more intensity values. In at least one embodiment, black and white pixels can be characterized as one intensity value representing darkness of a pixel, e.g., a value of 0 (or 1) can correspond to a white pixel and a value of 1 (or 0) can correspond to a black pixel. In at least one embodiment, intensity values can assume continuous (or quasi-continuous) values between 0 and 1 (or between any other chosen limits). Likewise, color pixels can be represented by multiple intensity values. For example, in an RGB scheme, there can be three separate intensity values for red, green, and blue, respectively. In a CMYK scheme, there can be four intensity values.

[0093] In at least one embodiment, a number of input nodes 612 of first network 610 can equal a number of pixels (or a total number of parameters describing all pixels). In at least one embodiment, a number of output nodes 614 can also equal a number of pixels, with each pixel assigned a score 622. In at least one embodiment, a number of output nodes 614 can be less than a number of pixels, e.g., if local scores 622 are assigned to superpixels. In at least one embodiment, first network 610 can be a convolutional network (CNN). In at least one embodiment, parameters of first network 610 can include a depth of convolution, a filter size, filter values (weights of convolution), a stride length, a bias, etc. In at least one embodiment, first network 610 can use multiple layers of neurons, such as an input layer, an output layer, and one or more hidden layers. In at least one embodiment, first network 610 can have 3, 4, 5, 6, 7, etc. layers of neurons or any other (lower or higher) number of layers. In at least one embodiment, some layers can be max-pooling layers and / or average-pooling layers. In at least one embodiment, inputs of different layers can be provided to subsequent layers using activation functions, such as a softplus function, a sigmoid function (e.g., a logistic sigmoid function), a hyperbolic tangent function, or different rectifier linear units (ReLUs), such as a standard ReLU, a leaky ReLU, an exponential ReLU, etc. In at least one embodiment, a leaky ReLU can be a parametric ReLU with one or more parameters to be determined during training.

[0094] In at least one embodiment, a matching illegal / legitimate pair of training images can be used, such that ground truth can be extracted that identifies different pixels as belonging to a non-illegal region (in both illegal and legitimate images) and / or an illegal region (in illegal images). In at least one embodiment, a matching pair of training images can involve a same (or similar) game scene, episode, scene, etc. In at least one embodiment, a legitimate image in a matching pair of training images can be generated by game software, while an illegal image in the matching pair can be generated by cheating software (or by cheating software executed in conjunction with game software). In at least one embodiment, during training of first network 610, output local scores 622 can be compared to target local scores (not shown) using one or more loss functions. In at least one embodiment, a loss function can be a binary cross-entropy loss function. More specifically, if c j is a likelihood (e.g., a probability) that jth pixel belongs to a region (of image 602) containing illegal information, computed by first network 610, and c j is a likelihood that the pixel belongs to a region not containing illegal information, respectively, and t j is a target likelihood of the same result (ground truth), then a loss function can be a sum over multiple pixels of image 602.

[0095]

[0096] In at least one embodiment, the plurality of pixels can include all pixels of image 602. In at least one embodiment, a loss function other than cross-entropy loss function can be used. For example, first network 602 (and second network 630) can employ a mean squared error loss function, a weighted mean squared error loss function, a mean absolute error loss function, a Huber loss function, a hinge loss function, a multi-class cross-entropy loss function, a Kullback-Liebler loss function, etc.

[0097] In at least one embodiment, instead of defining pixel classification at a local level, ground truth can identify all pixels of non-compliant training images having a first global value (e.g., value 1) and all pixels of non-non-compliant training images having a second global value (e.g., value 0). In at least one embodiment, during training, first network 610 can learn to identify pixels as belonging probabilistically to a non-compliant or a true image. For example, pixels similar to those encountered by first network 610 in both non-compliant images and true images with equal (e.g., 50%) probability (during training) can be classified as pixels belonging to a non-compliant image. In at least one embodiment, pixels similar to those encountered by first network 610 in non-compliant images with a probability w > 50% can be classified as pixels belonging to a non-compliant region. In at least one embodiment, to avoid false positives, the probability w can have to exceed a certain predetermined threshold w T (which can be 60%, 70%, 75%, etc.) or heuristically determined to provide an optimal balance between having too many false positives and losing too many pixels of non-compliant regions. In at least one embodiment, the described approach can be used without extracting local classifications of pixels (using matched non-compliant / true pairs of training images) in order to speed up training / re-training of first neural network 610. In at least one embodiment, the described approach can be used even when training is performed using non-matched true and non-compliant images. Specifically, true training images and non-compliant training images can be obtained from different game scenes, episodes, game versions, etc., or even from different games.

[0098] In at least one embodiment, during training of first network 610, training engine 342 can use backpropagation 625 to adjust different parameters (biases, weights, strides, filter biases, receptive depth, etc.) of first network until an observed difference (loss function) between computed output and target output (ground truth) is minimized. For example, first network 610 can utilize one or more matrix filters, parameters (matrix elements, depth, stride, etc.) of which can be adjusted during training.

[0099] In at least one embodiment, local scores 622 of feature map 620 can be input into second network 630, which can use local feature map 620 to make a final determination about image 602. In at least one embodiment, second network 620 can be trained to classify image 602 at a global level. A global level can refer to an image-level determination of whether image 602 includes any illicit information. Thus, in at least one embodiment, an output of second network 630 can be a binary output, e.g., 1 or yes if image 602 includes illicit information or 0 or no if image 602 does not have illicit information. In at least some embodiments, other outputs can be used. For example, global score 640 can be a sliding scale value, e.g., an output value from 0 to 0.5 can indicate that image 602 does not have illicit information, while a value from 0.5 to 1.0 can indicate that illicit information is present.

[0100] In at least one embodiment, a number of input nodes 632 of second network 630 can be equal to a number of pixels (or a total number of parameters describing all pixels). In at least one embodiment, second network 630 can use multiple layers of neurons, such as an input layer, a final layer, and one or more hidden layers (not shown). In at least one embodiment, an output node of second network 630 can include output node 636 that outputs global score 640. In at least one embodiment, second network 630 can include one or more fully connected layers of neurons. For example, an input fully connected layer (nodes 632) can uniformize local scores 622 into a single vector, which can then be processed by a next stage (e.g., a hidden fully connected layer or a final fully connected layer), where applying weights and biases to the output of the input layer can be used to predict global score 640. In at least one embodiment, second network 630 can use additional hidden layers to improve perception of feature map 620. In at least one embodiment, first neural network 610 and / or second neural network 630 can have skip connections that connect non-consecutive layers of neurons. For brevity and simplicity, a single skip connection of a node in a first hidden layer of first neural network 610 and a node in a second layer of second neural network 630 is shown, but it should be understood that any number of skip connections can be present in system 600, including any number of skip connections between non-consecutive layers of first network 610, second network 630, or skip connections connecting various layers of first network 610 and second network 630.

[0101] In at least one embodiment, during training of second network 630, various loss functions can be used, e.g., as described with respect to first network 610. In at least one embodiment, a loss function can be a binary cross-entropy loss function. In at least one embodiment, during training of second network 630, training engine 342 can use backpropagation 645 to adjust individual parameters (biases, weights, etc.) of second network until observed differences (loss function) between computed output and target output (ground truth) are minimized.

[0102] In at least one embodiment, training of first network 610 and second network 630 can be performed separately. For example, first network 610 can be trained using mapping data 349 that includes known locations (e.g., bounding boxes) of regions of pixels that contain illicit information. In at least one embodiment, known locations can be regions having complex shapes, such as polygons (including irregular polygons, concave polygons, etc.), non-polygon shapes drawn with a single closed line, shapes drawn with multiple closed lines, and / or any shape drawn with one or more open lines. In at least one embodiment, known locations can be identified (or otherwise mapped, using any mapping scheme or resource, such as a mapping table) as containing illicit information, e.g., using any appropriate identification at pixel level (such as a value of 1 for pixels that display illicit information and a value of 0 for pixels that do not display any illicit information). In at least one embodiment, a “region” can refer to an entire image, where a value of 1 identifies an image as having at least some illicit information, and a value of 0 identifies a clean image (without illicit information). In at least one embodiment, second network 630 can be trained based on binary mapping data 349 that includes identifying training images as either clean or illicit. In at least one embodiment, training of first network 610 or second network 630 (or both networks) can be performed using training images having different amounts of illicit information. For example, illicit training images can be grouped into multiple batches defined by an amount of illicit information remaining in the training images. In at least one embodiment, batches can have a decreasing amount of illicit content. In at least one embodiment, after training first network 610 and / or second network 630 using illicit training images that have 100% (or more) of illicit information present, a next batch of training images can be used (e.g., images that have 80% of illicit information remaining), and so on. In at least one embodiment, a transition from one batch to a next batch can occur after a certain target success rate is reached. In at least one embodiment, such a process can continue until a last batch (with X% of illicit information remaining) is applied. In at least one embodiment, a minimum percentage X of illicit information can be set based on a number of considerations. In at least one embodiment, X can be set based on what a most diligent cheater can use in order to avoid detection. In another embodiment, X can be set based on sufficiency of training. For example, if it is determined that reliable identification of illicit images becomes problematic when percentage drops below X, such a value X can be used as a minimum percentage for a last training batch. In at least one embodiment, some illicit information can be partially obscured rather than completely removed. For example, a region of pixels that includes illicit information can be given a reduced contrast so that such pixels are less recognizable from a background.In at least one embodiment, the color representation of pixels of illicit information can be changed, e.g., pixels of white, yellow, orange, red, etc. can be modified to include more colors of the spectrum of the blue, green, etc. portions.

[0103] In at least one embodiment, system 600 can also predict a confidence level 645 indicating how accurately output global score 640 predicts the likelihood of presence or absence of illicit information in image 602. In at least one embodiment, confidence level 645 can characterize the probability that output global score 640 is correct for an ensemble of images similar to image 602. In at least one embodiment, confidence level 645 can be output by an additional output node 638 (depicted by the hatched circle). In at least one embodiment, during training, some of the training images can be used to generate an ensemble for determining the confidence level. In at least one embodiment, various methods of identifying the confidence level can be used.

[0104] In at least one embodiment, random noise can be added to seed training image I0 to produce training image II. Similarly, another training image I2, etc. can be generated from image I0 using a different instance of added noise, until an ensemble of training images {I k} is obtained. In at least one embodiment, each image in ensemble {I k} can be analyzed by first network 610 and second network 630, and for each image a global score 640 can be output. Subsequently, in at least one embodiment, statistics of the ensemble of global scores {GS k} of ensemble {I k} can be determined and compared to the output confidence level. In at least one embodiment, the confidence level to compare can be the confidence level of seed image I0. Alternatively, in at least one embodiment, the confidence level to compare can be the average confidence level of the entire ensemble {I k}. In at least one embodiment, an additional loss function can then be used to evaluate the difference between the confidence level and the statistics of the output global scores {GS k}. For example, if a set of 10 global scores {GS k} of ensemble {I k} of 10 images correctly identifies the images as illicit 7 times out of 10, then the target confidence level can be 0.7. In at least one embodiment, if the confidence level output by node 638 (either for image I0 or for the entire ensemble {I kIf the average is 0.4, then the difference 0.7 - 0.4 = 0.3 can be backpropagated through the second network 630 and / or the first network 610 until the difference is minimized. In at least one embodiment, any suitable loss function can be used to evaluate the difference (e.g., cross-entropy, squared error, etc.). In at least one embodiment, as the parameters of the first network 610 and / or the second network 630 are modified, the output global score {GS k} can change along with the output confidence level. In at least one embodiment, the parameters of the networks can continue to be modified until both the difference in the output global score and the confidence level are minimized. Subsequently, in at least one embodiment, one or more additional seed images can be selected, respective sets of training images can be generated, and the training of the neural networks can be repeated. In at least one embodiment, as a result of the overall training, the system 600 acquires the ability to output a confidence level 645 that predicts the likelihood that the global score 640 accurately determines the type of the image 602 (fake vs. real).

[0105] In at least one embodiment, the confidence level can be determined based on introducing randomness into the neural network architecture. More specifically, the training images can be processed multiple times by a set of neural networks in which at least some (or all) of the nodes are subject (with some probability) to removal (dropout). Removal of a node means that all of the incoming and outgoing network connections of the node are eliminated. In at least one embodiment, the probability of removal can depend on the layer in which the removed node is located, and can be greater for hidden layers than for input / output layers. In at least one embodiment, the confidence level can be determined based on how successfully the individual neural networks of the generated set identify the images (or the pixels of the images) as fake or real. In at least one embodiment, to determine the confidence level, various methods of classifying uncertainty can be used, including a variation ratio method, a mutual information method, or any other similar method (e.g., a predictive entropy method). For example, the variation ratio method can use a variation ratio R = 1 - M / N to the total number of all outcomes N to identify the number of successful identifications M of a target outcome (which can be a pattern of the overall distribution of outcomes) and a base confidence level. In at least one embodiment, the confidence level can be given by the determined variation ratio. In at least one embodiment, the confidence level can be a function (e.g., a non-linear function) of the determined variation ratio. In the mutual information method, the maximization can be a degree of correlation between the distribution of outcomes in the set and the distribution of correct outcomes. In at least one embodiment, the degree of correlation can be represented by the Kullback-Leibler divergence of the two distributions.

[0106] In at least one embodiment, classification of a pixel of an image (or a region of an image, or an entire image) can be performed on “illegal”, “not illegal”, and “uncertain” classes. In at least one embodiment, classification of a pixel of an image (or a region of an image, or an entire image) can be performed on “not illegal” with probability p0, “illegal” with probability pi, and “uncertain” with probability u, such that p0+ pi + u = 1. This uncertainty u can be determined using a Dirichlet distribution based multinomial opinion probability. In at least one embodiment, this can be done by defining a loss function and computing a Bayesian risk with respect to class predictors (e.g., “illegal”, “not illegal”, and “uncertain”). In at least one embodiment, a loss function can be a squared error loss function. In at least one embodiment, a loss function can be a Kullback-Leibler divergence. In at least one embodiment, a loss function can be a type II maximum likelihood function.

[0107] In at least one embodiment, during an inference phase of a neural network operation, an uncertainty probability u can be converted to a confidence level, e.g., a confidence level can be a function of uncertainty, CL = f(u), which can be a decreasing function of uncertainty, such that higher uncertainty corresponds to lower confidence level CL. In at least one embodiment, any appropriate function f(u) can be used to output a confidence level (e.g., within an interval of 0 to 1, 0% to 100%, or any other interval).

[0108] In at least one embodiment, during training, parameters of a neural network can be modified to maximize a confidence level CL. In at least one embodiment, parameters of a neural network can be modified to achieve a desired balance between increasing confidence level CL and increasing likelihood of cheat detection. For example, increasing confidence level can mean that some cheaters can not be detected. Thus, in at least one embodiment, some decrease in confidence level can be intentionally accepted to achieve more widespread detection of cheating. In at least one embodiment, training of a neural network can include optimizing a confidence level CL for a probability P of correct classification of images across a set (or subset) of training images. In at least one embodiment, optimization can involve finding a maximum of a function F(CL, P), such as F(CL, P) = a(CL) 2 + b(P) 2where the coefficients a and b are selected according to the target of the detection. More specifically, a higher value of a can favor detecting fewer instances of cheating with a higher confidence level, while a higher value of b can favor detecting more potential instances of cheating, even if some of these instances can be false positives (or false negatives). In at least one embodiment, other functions F(CL, P) can be used. In at least one embodiment, an optimization can be performed until the highest possible correct detection rate is achieved under the condition that the confidence level meets a minimum threshold. In at least one embodiment, an optimization can be performed until the highest possible confidence level is achieved under the condition that the correct detection rate meets a minimum threshold.

[0109] In at least one embodiment, the confidence level CL can be estimated during training using a method of interval bound propagation (IBP). For example, a training image can be modified by adding noise (as described above), and the output of the neural network - a classification of the image or image portion as fraudulent or genuine - can be generated. In at least one embodiment, an upper bound on the bias ±Δ of the output of the classification can then be estimated for a given predetermined uncertainty ±∈ of the input (e.g., caused by the noise). In at least one embodiment, the upper bound Δ can be estimated mathematically using one or more IBP models. In at least one embodiment, the upper bound Δ can be determined empirically as the maximum bias observed for a set of training images modified with noise. In at least one embodiment, the estimated or determined upper bound Δ can be used to determine or adjust the confidence level CL. In at least one embodiment, the confidence level for all images can take the upper bound into account. For example, if the estimated upper bound is Δ = 10%, then a confidence level CL = 80% can be reduced to 70%. In at least one embodiment, the parameters of the neural network that output the confidence level (which can include the upper bound adjustment) can be determined during training (where ground truth is known) and also used during the inference phase (where ground truth is not known).

[0110] In at least one embodiment, the first neural network 610 and the second neural network 620 can be trained to output a confidence level by adding noise to the parameters of the neural network, rather than to the training images, such as adding noise to the weight, bias, and / or activation parameters in the various layers (e.g., hidden layers) of the first network 610 and / or the second network 620. In at least one embodiment, adding noise to the parameters can be performed in addition to adding noise to the training images.

[0111] In at least one embodiment, a confidence level determined for one or a series of images on a user’s (player’s) computer during an inference phase can be used for different purposes, including but not limited to reporting instances of detected cheating, avoiding reporting some instances of suspected cheating, signaling when a retraining of a neural network should be performed, and the like. For example, if a confidence level is high, e.g., above a certain predetermined threshold CL1, then a detected instance of cheating can be reported to a publishing server 302, game server 304, or other resource. In at least one embodiment, if a confidence level is below CL1, then a detected instance of cheating can not be reported, in order to avoid false positive determinations. In at least one embodiment, reporting can be performed once a certain predetermined number of images are shown to a user (or a certain predetermined time of image exposure to a user is detected), and are determined to be illegal at least with a confidence level of CL1. In at least one embodiment, if a series of (triggering) images are encountered that are determined with a confidence level below a second threshold CL2 (which can be the same as or different from threshold CL1), then a report can be sent to a publishing server 302 or game server 304. In at least one embodiment, receipt of such a report can signal to a publishing server 302 or game server 304 that new cheating software (or a new version of existing software) can have become available, and that a retraining of a neural network using images generated with new cheating software should be performed. In at least one embodiment, a minimum number of different images (or a minimum number of different game scenarios received from different users) should exist before a report is sent, in order to avoid false positives that can be caused by various artifacts on a user’s computer, such as software glitches, power surges, and the like.

[0112] In at least one embodiment, during training of system 600, backpropagation 646 (shown in dashed line) can extend on both second network 630 and first network 610. Thus, in at least one embodiment, a target output 348 used to train system 650 can include a single value describing whether a training input is a genuine image or an illegal image. In such embodiments, a ground truth for feature map 620 describing characters of individual pixels is not included, as local scores 622 can be an internal value that is not explicitly accessible.

[0113] In at least one embodiment, system 600 is represented as a combination of first network 610 and second network 210 for ease of reader to show different functionality system 600 can have. In at least one embodiment, first network 610 and second network 620 represent a single network. In at least one embodiment, feature map 620 can be extracted from different intermediate output nodes, such as node 614, and used for training when applying local backpropagation 625. In at least one embodiment, feature map 620 can not be explicitly output (and can represent intermediate computational results that are not directly accessed), e.g., in case of using a single global backpropagation 646.

[0114] In at least one embodiment, system 600 can be trained against adversary attacks. Adversary attacks (e.g., attacks initiated by producers or distributors of cheat software) can be based on small modifications to cheat images that can be imperceptible to humans, but still sufficient to confuse a neural network into incorrectly classifying illegal images as non-illegal images. In at least one embodiment, system 600 provides inherent protection against adversary attacks. Specifically, because GPU 320 renders final images on display 329 at a rate of 30-60 frames / second or greater, potential attacks (e.g., initiated by cheat software operating on user machine 310) can have very limited time, preventing them from being successful attacks. Thus, in at least one embodiment, an attacker attempting to use adversary probe images to probe a response of a neural network (e.g., NN model 325) is likely to be detected before an image capable of successfully confusing the neural network can be developed. As such, in at least one embodiment, GPU 320 can detect patterns in which initial detection of illegal images ceases after a period of time. In at least one embodiment, GPU 320 can report such patterns, such as an indication that a user can be using cheat software capable of developing a successful adversary attack.

[0115] In at least one embodiment, parameters (weights, biases, parameters of activation functions, etc.) of a neural network (e.g., NN model 325) can be kept secret, e.g., encrypted protected with a key, a hash, and other protection mechanisms capable of preventing a potential attacker from developing a successful offline attack elsewhere that can later be transmitted to user machine 310 (e.g., a universal adversary image perturbation). Moreover, an attacker can lack access to an architecture of a neural network. In at least one embodiment, due to a short amount of time available for an adversary attack, an ongoing adversary attack can be a relatively simple type, such as a Madry attack, a Fast Gradient Sign Method (FGSM) attack, etc. In at least one embodiment, to protect against such attacks, during training of system 600, an adversary image can be prepared using an algorithm that can be deployed in such an immediate attack to train system 600 to ignore adversary perturbations that an attacker can exploit. In at least one embodiment, an IBP method can be used to train system 600 against an adversary attack. For example, during training, IBP can be determined for a selected uncertainty ±∈ of an adversary input (e.g., a perturbation) that can be used in an attack. In at least one embodiment, such prepared adversary input can be propagated through a neural network and a bound Δ can be determined for the selected uncertainty ∈. In at least one embodiment, training can be subsequently performed based on a confidence level that has been adjusted (based on a worst-case scenario) using the determined bound.

[0116] Figure 7 is a flowchart of an example method 700 of detecting illicit information in images rendered by a graphics processing unit on a display of a user device using one or more neural networks, in accordance with at least one embodiment. In at least one embodiment, method 700 can be performed by one or more circuits using one or more neural networks. In at least one embodiment, method 700 can be performed by processing logic of user machine 310. More specifically, method 700 can be performed by GPU 320 including one or more circuits (e.g., GPU cores 322) and one or more memory devices (e.g., GPU 324). In at least one embodiment, method 700 can be performed by a single processing thread. Alternatively, in at least one embodiment, method 700 can be performed by two or more processing threads, each thread performing one or more individual functions, routines, subroutines, or operations of method 700. In at least one embodiment, processing threads implementing method 700 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, in at least one embodiment, processing threads implementing method 700 can be performed asynchronously with respect to each other. In at least one embodiment, various operations of method 700 can be performed in parallel, in series, or in any combination of parallel and series. Figure 7The order shown is performed in contrast to different orders. In at least one embodiment, some operations of method 700 can be performed concurrently with other operations. In at least one embodiment, one or more operations shown in method 700 can not be performed. Although different operations of method 700 are described with respect to a game application and cheating software, it should be understood that, in at least one embodiment, similar operations can be used to detect illicit information in images associated with any other type of application. Figure 7

[0117] In at least one embodiment, processing logic performing method 700 can execute one or more neural networks (e.g., NN model 325) operating in an inferencing mode to determine whether to render and display images containing illicit (e.g., cheating) information to a user (e.g., a gamer). In at least one embodiment, the displayed images are rendered by the same processing device performing method 700. In at least one embodiment, the displayed images can be rendered by a different processing device. In at least one embodiment, at block 710, processing logic performing method 700 can receive, from a processing device (e.g., CPU 316), a representation of graphics associated with a computer game. For example, CPU 316 can output instructions to GPU to render graphics corresponding to a game scenario on a user’s display. In at least one embodiment, the graphics can include one or more images that are to be displayed on a user’s display simultaneously or sequentially (e.g., at successive instances in time). In at least one embodiment, the representation of graphics can be in any electronic (e.g., digital) format accessible to processing logic (e.g., GPU 324) performing method 700.

[0118] In at least one embodiment, at block 720, processing logic performing method 700 can generate one or more images based on the received representation. In at least one embodiment, the generated images can be stored in a buffer (e.g., buffer 326) before being sent to a user’s display (e.g., display 329). In at least one embodiment, at block 730, the processing device can use one or more trained neural networks to detect cheating by one or more users of the computer game. In at least one embodiment, the one or more images can be retrieved from the buffer before, concurrently with, or after the images are sent to the display. In at least one embodiment, the neural networks can be the neural networks described in connection with Figure 6 In at least one embodiment, the neural networks can output a determination (e.g., global score 640) indicating that one or more images include pixels for displaying cheating information.

[0119] ​In at least one embodiment, processing logic can generate, using one or more neural networks, a confidence level characterizing a confidence that one or more images include cheating information. In at least one embodiment, at block 750, processing logic can generate, using one or more neural networks, a report indicating that cheating was detected. In at least one embodiment, a report can include an indication that cheating was detected. In at least one embodiment, a report can also include a generated confidence level. In at least one embodiment, at block 760, a report can be communicated to a game server, which can be (or include) a publishing server, a server of a game developer, or any other server associated with a producer or distributor of game software.

[0120] In at least one embodiment, a game server that has received a report can initiate retraining of one or more neural networks (either automatically or upon command from a human developer / engineer). For example, in a case where a certain (e.g., predetermined) number of images are detected with a confidence level below a threshold (e.g., 20%, 50%, or any other predetermined value), a game server (and / or developer / engineer) can conclude that new cheating software (or a new version of existing cheating software) has become available. In at least one embodiment, new cheating software can have changed placement of cheating information, format (e.g., font, color, transparency of overlay, etc.) used to display cheating information, range and characters of cheating information, etc. As a result, in at least one embodiment, one or more neural networks can detect a change in performance of cheating images and signals via reduced confidence levels, i.e., retraining of one or more neural networks should be performed. In at least one embodiment, a game server can have access to game logs (e.g., game logs 306) of respective players that can be stored on the game server to help identify a context in which new cheating software is operating, a type of advantage that new cheating software is providing, etc. In at least one embodiment, game server 304, publishing server 302, and / or training server 340 can use game logs 306 to replay a game context, use game software to re-render images and add illicit information to produce a new set of training images. In at least one embodiment, upon having received a report containing reduced confidence levels, and when accessing game logs of one or more users, a developer / engineer can obtain new cheating software and install it on a game server and / or training server 304. In at least one embodiment, a game server and / or training server can include a same or similar set of neural networks as one or more neural networks installed on user devices. In at least one embodiment, neural networks can be retrained using a new set of retraining images, which in one embodiment can be obtained using method 800 described below. In at least one embodiment, after retraining using retraining images, updated parameters (biases, weights, filter parameters, activation parameters, etc.) of one or more neural networks can be obtained and communicated to user devices (e.g., during a downtime of user devices). In at least one embodiment, at block 770, processing logic on a user device can receive updated parameters and use received parameters to update one or more neural networks.

[0121] Figure 8This is a flowchart of an example method 800 for training one or more neural networks to detect illegal information in an image rendered by a graphics processing unit on a user device's display, according to at least one embodiment. In at least one embodiment, method 800 may be executed by the processing logic of a training server 340. In at least one embodiment, method 800 may be executed by a single processing thread. Alternatively, in at least one embodiment, method 800 may be executed by two or more processing threads, each thread executing one or more various functions, routines, subroutines, or operations of method 800. In at least one embodiment, the processing threads implementing method 800 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, in at least one embodiment, the processing threads implementing method 800 may execute asynchronously relative to each other. In at least one embodiment, various operations of method 800 may be performed in conjunction with... Figure 8 The order shown is compared to a different order of execution. In at least one embodiment, some operations of method 800 can be performed concurrently with other operations. In at least one embodiment, they may not be performed. Figure 8 One or more operations are shown. Although various operations of method 800 are described in relation to game applications and cheat software, it should be understood that similar operations can be used to detect illegitimate information in images associated with any other type of application. In at least one embodiment, method 800 can be used for training a neural network and for retraining a previously trained neural network when new cheat software becomes available.

[0122] In at least one embodiment, the processing logic of method 800 may execute one or more neural networks (e.g., NN model 325) operating in training mode to determine the parameters of one or more neural networks to detect cheating (or any other illegal) images rendered and displayed to a user (e.g., a gamer). In at least one embodiment, method 800 may be executed by processing logic capable of accessing game software (e.g., legitimate software) capable of generating genuine (non-cheating) images. In at least one embodiment, method 800 may be executed by processing logic capable of accessing cheating software, which may be used alone or in conjunction with game software to generate illegal images (cheating images), which may be any image displaying any information that provides an unfair advantage or contains any unauthorized information. Boxes 810, 820, and 830 refer to the generation of enhanced cheating training images, while boxes 812, 822, and 832 refer to the generation of enhanced non-cheating training images. In at least one embodiment, blocks 810, 820 and 830 may be executed independently (e.g., in parallel or sequentially) from blocks 812, 822 and 832.

[0123] In at least one embodiment, at block 810, processing logic executing method 800 can generate a non-cheating image using game software. In at least one embodiment, at block 820, processing logic executing method 800 can generate a base cheating image using cheating software. In at least one embodiment, at block 830, processing logic executing method 800 can replace portions of the base cheating image with portions of the non-cheating image. As a result, in at least one embodiment, an enhanced cheating training image can be generated that has a reduced amount of cheating information compared to the base cheating image. In at least one embodiment, blocks 810, 820, and 830 can be repeated multiple times to generate multiple batches of enhanced cheating training images (based on the same base cheating image) with different amounts of cheating information. Additionally, in at least one embodiment, blocks 810, 820, and 830 can be repeated using different base cheating images to obtain multiple batches (based on different base cheating images) with different amounts of cheating information.

[0124] In at least one embodiment, to impart paste artifacts to non-cheating images, similar cross-enhancement can be performed to generate non-cheating images, similar to imparting paste artifacts to cheating images. In at least one embodiment, at block 812, processing logic executing method 800 can generate a base non-cheating image using game software. In at least one embodiment, at block 822, processing logic executing method 800 can generate a cheating image using cheating software. In at least one embodiment, at block 832, processing logic executing method 800 can replace portions of the base non-cheating image with portions of the cheating image. As a result, in at least one embodiment, an enhanced non-cheating training image can be generated that has paste artifacts similar to the artifacts in the enhanced cheating image. In at least one embodiment, the portions of the cheating image pasted to the base non-cheating image contain only non-cheating information (non-cheating regions of the cheating image).

[0125] In at least one embodiment, blocks 812, 822, and 832 can be repeated multiple times to generate multiple batches of enhanced non-cheating training images (based on the same base non-cheating image) with different amounts of paste artifacts. Additionally, in at least one embodiment, blocks 812, 822, and 832 can be repeated using different base non-cheating images to obtain multiple batches (based on different base non-cheating images) with different amounts of paste artifacts.

[0126] In at least one embodiment, at block 840, the processing logic performing method 800 may use the enhanced cheating images obtained (via blocks 810, 820, and 830) to train one or more neural networks to detect cheating by one or more users in a computer game. In at least one embodiment, the processing logic performing method 800 may also use the enhanced non-cheating images obtained (via blocks 812, 822, and 832) to train one or more neural networks.

[0127] Data Center

[0128] Figure 9 An example data center 900 that can be used with at least one embodiment is shown. In at least one embodiment, the data center 900 includes a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0129] In at least one embodiment, such as Figure 9 As shown, the data center infrastructure layer 910 may include a resource coordinator 912, packet computing resources 914, and node computing resources (“nodes CR”) 916(1)-916(N), where “N” represents a positive integer (which may be an integer “N” different from the integers used in other diagrams). In at least one embodiment, nodes CR 916(1)-916(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 918(1)-918(N) (e.g., dynamic read-only memory, solid-state drives, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 916(1)-916(N) may be servers having one or more of the aforementioned computing resources.

[0130] In at least one embodiment, grouped computing resources 914 can include individual groupings of nodes C.R. housed within one or more racks (not shown), or housed within a number of racks (also not shown) within data centers at various geographic locations. In at least one embodiment, individual groupings of nodes C.R. within grouped computing resources 914 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several nodes C.R. including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches in any combination.

[0131] In at least one embodiment, resource orchestrator 912 can configure or otherwise control one or more nodes C.R. 916(1)-916(N) and / or grouped computing resources 914. In at least one embodiment, resource orchestrator 912 can include a software design infrastructure (“SDI”) management entity for data center 900. In at least one embodiment, resource orchestrator 112 can comprise hardware, software, or some combination thereof.

[0132] In at least one embodiment, as Figure 9As shown, the framework layer 920 includes a job scheduler 922, a configuration manager 924, a resource manager 926, and a distributed file system 928. In at least one embodiment, the framework layer 920 can include a framework that supports software 932 of a software layer 930 and / or one or more applications 942 of an application layer 940. In at least one embodiment, software 932 or applications 942 can include web-based service software or applications, respectively, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 920 can be, but is not limited to, a type of free and open-source software web application framework such as Apache Spark™ (hereinafter “Spark”) that can utilize the distributed file system 928 for large-scale data processing (e.g., “big data”). In at least one embodiment, the job scheduler 922 can include a Spark driver to facilitate scheduling workloads supported by various layers of the data center 900. In at least one embodiment, the configuration manager 924 can be capable of configuring different layers, such as the software layer 930 and the framework layer 920 including Spark and the distributed file system 928 for supporting large-scale data processing. In at least one embodiment, the resource manager 926 can be capable of managing clustered or grouped computing resources mapped to or allocated for supporting the distributed file system 928 and the job scheduler 922. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 914 on the data center infrastructure layer 910. In at least one embodiment, the resource manager 926 can coordinate with the resource orchestrator 912 to manage these mapped or allocated computing resources.

[0133] In at least one embodiment, software 932 included in the software layer 930 can include software used by at least a portion of the node C.R.s 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 928 of the framework layer 920. In at least one embodiment, one or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0134] In at least one embodiment, one or more applications 942 included in application layer 940 can include one or more types of applications used by at least portions of node C.R.s 916(1)-916(N), grouped computing resources 914, and / or distributed file system 928 of framework layer 920. In at least one embodiment, one or more types of applications can include, but are not limited to, any number and type of genomics applications, cognitive computing, applications, and machine learning applications including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0135] In at least one embodiment, any of configuration manager 924, resource manager 926, and resource orchestrator 912 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible fashion. In at least one embodiment, self-modification actions can relieve data center operators of data center 900 from making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a data center.

[0136] In at least one embodiment, data center 900 can include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by computing weight parameters according to a neural network architecture using software and computing resources described above with respect to data center 900. In at least one embodiment, using weight parameters computed by one or more training techniques described herein, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 900.

[0137] In at least one embodiment, a data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Furthermore, one or more software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference such as image recognition, speech recognition, or other artificial intelligence services.

[0138] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, IB, 2, 3, 4, 5, 6, and 7. Figure 1A and / or Figure 1BDetails regarding inference and / or training logic 115 are provided. In at least one embodiment, inference and / or training logic 115 can be used in the system Figure 9 for inferencing or predicting operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions, and / or architectures, or neural network use cases described herein.

[0139] Autonomous vehicle

[0140] Figure 10A An example of an autonomous vehicle 1000 is shown, in accordance with at least one embodiment. In at least one embodiment, autonomous vehicle 1000 (alternatively referred to herein as “vehicle 1000”) can be, but is not limited to, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle that can accommodate one or more passengers. In at least one embodiment, vehicle 1000 can be a semi-truck tractor-trailer used for hauling cargo. In at least one embodiment, vehicle 1000 can be an airplane, a robotic vehicle, or another type of vehicle.

[0141] Autonomous vehicles can be described in terms of automation levels defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). In at least one embodiment, vehicle 1000 can be capable of functioning according to one or more of Levels 1 through 5 of the automation levels. For example, in at least one embodiment, vehicle 1000 can be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on embodiment.

[0142] In at least one embodiment, vehicle 1000 can include, without limitation, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1000 can include, without limitation, a propulsion system 1050, such as a combustion engine, a hybrid electric device, a fully electric motor, and / or another propulsion system type. In at least one embodiment, propulsion system 1050 can be connected to a drivetrain of vehicle 1000, which can include, without limitation, a transmission to enable propulsion of vehicle 1000. In at least one embodiment, propulsion system 1050 can be controlled in response to receiving signals from throttle / accelerator 1052.

[0143] In at least one embodiment, when propulsion system 1050 is operating (e.g., when vehicle 1000 is in motion), steering system 1054 (which can include, without limitation, a steering wheel) is used to steer vehicle 1000 (e.g., along a desired path or course). In at least one embodiment, steering system 1054 can receive signals from steering actuator 1056. In at least one embodiment, a steering wheel can be optional for full automation (Level 5) functionality. In at least one embodiment, brake sensor system 1046 can be used to operate vehicle brakes in response to signals received from brake actuator 1048 and / or brake sensors.

[0144] In at least one embodiment, controller 1036 can include, without limitation, one or more system on a chip (“SoC”) (e.g., one or more processors) that can be configured to perform one or more operations described herein. In at least one embodiment, controller 1036 can include, without limitation, one or more processors that can be configured to perform one or more operations described herein. Figure 10Aand / or a graphics processing unit (“GPU”) to provide signals (e.g., representative of commands) to one or more components and / or systems of vehicle 1000. For example, in at least one embodiment, controller(s) 1036 can send signals to operate vehicle brakes by brake actuator(s) 1048, to operate steering system 1054 by steering actuator(s) 1056, to operate propulsion system 1050 by throttle / accelerator(s) 1052. In at least one embodiment, controller(s) 1036 can include one or more on-board (e.g., integrated) computing devices that process sensor signals and output operational commands (e.g., signals representative of commands) to enable autonomous driving and / or assist a human driver in driving vehicle 1000. In at least one embodiment, controller(s) 1036 can include a first controller for autonomous driving functionality, a second controller for functional safety functionality, a third controller for artificial intelligence functionality (e.g., computer vision), a fourth controller for infotainment functionality, a fifth controller for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller can handle two or more of above-described functionalities, two or more controllers can handle a single functionality, and / or any combination thereof.

[0145] In at least one embodiment, controller(s) 1036 provide signals for controlling one or more components and / or systems of vehicle 1000 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors such as, but not limited to, one or more global navigation satellite system (“GNSS”) sensors 1058 (e.g., one or more global positioning system sensors), one or more RADAR sensors 1060, one or more ultrasonic sensors 1062, one or more LIDAR sensors 1064, one or more inertial measurement unit (IMU) sensors 1066 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, one or more magnetorquers, etc.), one or more microphones 1096, one or more stereo cameras 1068, one or more wide-view cameras 1070 (e.g., fisheye cameras), one or more infrared cameras 1072, one or more surround cameras 1074 (e.g., 360-degree cameras), long-range cameras (not shown in FIG. 10), mid-range cameras (not shown in FIG. 10), and / or other sensors. Figure 10A In at least one embodiment, controller(s) 1036 provide signals for controlling one or more components and / or systems of vehicle 1000 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data can be received from sensors such as, but not limited to, one or more global navigation satellite system (“GNSS”) sensors 1058 (e.g., one or more global positioning system sensors), one or more RADAR sensors 1060, one or more ultrasonic sensors 1062, one or more LIDAR sensors 1064, one or more inertial measurement unit (IMU) sensors 1066 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, one or more magnetorquers, etc.), one or more microphones 1096, one or more stereo cameras 1068, one or more wide-view cameras 1070 (e.g., fisheye cameras), one or more infrared cameras 1072, one or more surround cameras 1074 (e.g., 360-degree cameras), long-range cameras (not shown in FIG. 10), mid-range cameras (not shown in FIG. 10), and / or other sensors. Figure 10Aone or more speed sensors 1044 (e.g., to measure a speed of the vehicle 1000), one or more vibration sensors 1042, one or more steering sensors 1040, one or more brake sensors (e.g., as part of a brake sensor system 1046), and / or other sensor types.

[0146] In at least one embodiment, the one or more controllers 1036 can receive inputs (e.g., represented by input data) from an instrument cluster 1032 of the vehicle 1000 and provide outputs (e.g., represented by output data, display data, etc.) through a human-machine interface (“HMI”) display 1034, audible annunciators, speakers, and / or other components of the vehicle 1000. In at least one embodiment, outputs can include information such as vehicle speed, velocity, time, map data (e.g., high-definition map Figure 10A In at least one embodiment, the HMI display 1034 can display information about the presence of one or more objects (e.g., a street sign, a warning sign, a traffic signal change, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., change lanes now, take exit 34B in two miles, etc.). In at least one embodiment, the HMI display 1034 can also display information about the vehicle’s surroundings (e.g., as shown in FIG. 13), such as a map showing the vehicle’s location (e.g., on a map), a location of other vehicles (e.g., an occupancy grid), information about objects, and a state of objects perceived by the one or more controllers 1036, etc.

[0147] In at least one embodiment, the vehicle 1000 also includes a network interface 1024 that can communicate over one or more networks using one or more wireless antennas 1026 and / or one or more modems. For example, in at least one embodiment, the network interface 1024 can be capable of communicating over a Long-Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”) network, etc. In at least one embodiment, the one or more wireless antennas 1026 can also use one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low power wide area networks (hereinafter “LPWANs”) (e.g., LoRaWAN, SigFox, etc. protocols) to enable communication between objects (e.g., vehicles, mobile devices) in an environment.

[0148] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Inference and / or training logic 115 are used to process electrical signals received from one or more sensors (e.g., image sensors, microphones, etc.), and provide other electronic signals in response to the electrical signals based, at least in part, on one or more trained models associated with the one or more embodiments. Figure 1A and / or Figure 1B Details are provided regarding the inference and / or training logic 115. In at least one embodiment, the inference and / or training logic 115 can be implemented in the system. Figure 10A The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0149] Figure 10B The illustration shows an embodiment according to at least one of the embodiments. Figure 10A Examples of camera positions and fields of view for an autonomous vehicle 1000. In at least one embodiment, the camera and its respective field of view are exemplary embodiments and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 1000.

[0150] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 1000. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with an array of RCCC, RCCB and / or RBGC color filters, may be used to improve photosensitivity.

[0151] In at least one embodiment, one or more cameras may be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundancy or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).

[0152] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (3D-printed) assembly, to cut out stray light and reflections within the vehicle 1000 (e.g., reflections from the dashboard in the windshield mirror), which may interfere with the camera's image data capture capabilities. Regarding the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly may be 3D-printed custom-made such that the camera mounting plate matches the shape of the rearview mirror. In at least one embodiment, one or more cameras may be integrated into the rearview mirror. In at least one embodiment, for side-view cameras, one or more cameras may also be integrated within four pillars at each corner of the cabin.

[0153] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view including a portion of the environment in front of the vehicle 1000 can be used for surround view and, with the assistance of one or more controllers 1036 and / or control SoCs, to help identify forward paths and obstacles, thereby providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many ADAS functions similar to LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions (e.g., traffic sign recognition).

[0154] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide-semiconductor”) color imager. In at least one embodiment, a wide-angle camera 1070 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the street, or bicycles). Although in Figure 10B Only one wide-angle camera 1070 is shown; however, in other embodiments, the vehicle 1000 may have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of remote cameras 1098 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote camera 1098 can also be used for object detection and classification, as well as basic object tracking.

[0155] In at least one embodiment, any number of stereo cameras 1068 can also be included in a forward-facing configuration. In at least one embodiment, one or more stereo cameras 1068 can include an integrated control unit that includes a scalable processing unit that can provide programmable logic (“FPGA”) and multi-core microprocessors with integrated controller area network (“CAN”) or Ethernet interfaces on a single chip. In at least one embodiment, such a unit can be used to generate a 3D map of an environment of vehicle 1000, including distance estimates for all points in an image. In at least one embodiment, one or more stereo cameras 1068 can include, without limitation, a compact stereo-vision sensor that can include, without limitation, two camera lenses (one each on the left and right) and an image processing chip that can measure distances from vehicle 1000 to target objects and use generated information (e.g., metadata) to activate autonomous emergency braking and lane-departure warning functionality. In at least one embodiment, other types of stereo cameras 1068 can be used in addition to or instead of those described herein.

[0156] In at least one embodiment, cameras with a field of view that includes a portion of an environment to the sides of vehicle 1000 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating an occupancy grid, as well as generating side collision warnings. For example, in at least one embodiment, surround cameras 1074 (e.g., four surround cameras as shown in Figure 10B In at least one embodiment, one or more surround cameras 1074 can include, without limitation, any number and combination of wide-view cameras, fisheye lenses, 360-degree cameras, and / or the like. For example, in at least one embodiment, four fisheye lens cameras can be located on front, back, and sides of vehicle 1000. In at least one embodiment, vehicle 1000 can use three surround cameras 1074 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., front-facing cameras) as a fourth surround view camera.

[0157] In at least one embodiment, cameras with a field of view that includes a portion of an environment to the rear of vehicle 1000 (e.g., rear-view cameras) can be used for parking assistance, surround view, rear collision warnings, and creating and updating an occupancy grid. In at least one embodiment, a wide variety of cameras can be used, including without limitation cameras that are also suitable as one or more forward-facing cameras (e.g., long-range cameras 1098 and / or one or more mid-range cameras 1076, one or more stereo cameras 1068, one or more infrared cameras 1072, etc.), as described herein.

[0158] Inference and / or training logic 115 is used to perform inference and / or training operations associated with one or more embodiments. Figure 1A and / or Figure 1B This document provides details regarding inference and / or training logic 115. In at least one embodiment, inference and / or training logic 115 may be... Figure 10B Used in systems for reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0159] Figure 10C The illustration shows an embodiment according to at least one of the embodiments. Figure 10A A block diagram of an example system architecture for an autonomous vehicle 1000. In at least one embodiment, Figure 10C Each of one or more components, one or more features, and one or more systems of vehicle 1000 is shown as connected via bus 1002. In at least one embodiment, bus 1002 may include, but is not limited to, a CAN data interface (which may alternatively be referred to herein as “CAN bus”). In at least one embodiment, CAN may be a network within vehicle 1000 used to help control various features and functions of vehicle 1000, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. In one embodiment, bus 1002 may be configured to have dozens or even hundreds of nodes, each node having its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 1002 can be read to find steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1002 may be an ASIL B compliant CAN bus.

[0160] In at least one embodiment, FlexRay and / or Ethernet protocols can also be used in addition to or instead of CAN. In at least one embodiment, there can be any number of buses forming bus 1002, which can include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses can be used to perform different functions, and / or can be used for redundancy. For example, a first bus can be used for collision avoidance functions, and a second bus can be used for actuation control. In at least one embodiment, each bus in bus 1002 can communicate with any component of vehicle 1000, and two or more of bus 1002 can communicate with respective components. In at least one embodiment, each of any number of system on a chip (“SoC”) 1004 (e.g., SoC 1004(A) and SoC 1004(B)), each of one or more controllers 1036, and / or every computer within a vehicle can have access to same input data (e.g., input from sensors of vehicle 1000), and can be connected to a common bus, such as a CAN bus.

[0161] In at least one embodiment, vehicle 1000 can include one or more controllers 1036, such as those described herein with respect to Figure 10A In at least one embodiment, controllers 1036 can be used for a variety of functions. In at least one embodiment, controllers 1036 can be coupled to any of various other components and systems of vehicle 1000, and can be used to control vehicle 1000, artificial intelligence of vehicle 1000, infotainment of vehicle 1000, and / or other functions.

[0162] In at least one embodiment, vehicle 1000 can include any number of SoCs 1004. In at least one embodiment, each of SoCs 1004 can include, without limitation, central processing units (“one or more CPUs”) 1006, graphics processing units (“one or more GPUs”) 1008, one or more processors 1010, one or more caches 1012, one or more accelerators 1014, one or more data stores 1016, and / or other non- shown components and features. In at least one embodiment, one or more SoCs 1004 can be used to control vehicle 1000 in a variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 1004 can be combined with a high definition (“HD”) map 1022 in a system (e.g., a system of vehicle 1000), which can be obtained from one or more servers via network interface 1024. In at least one embodiment, one or more SoCs 1004 can be used to control vehicle 1000 in a variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 1004 can be combined with a high definition (“HD”) map 1022 in a system (e.g., a system of vehicle 1000), which can be obtained from one or more servers via network interface 1024. Figure 10C map refresh and / or update is obtained (not shown).

[0163] In at least one embodiment, one or more CPU(s) 1006 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, one or more CPU(s) 1006 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, one or more CPU(s) 1006 can include eight cores in a multi-processor configuration coupled to one another. In at least one embodiment, one or more CPU(s) 1006 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., 2 MB L2 cache). In at least one embodiment, one or more CPU(s) 1006 (e.g., CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of one or more CPU(s) 1006 can be active at any given time.

[0164] In at least one embodiment, one or more CPU(s) 1006 can implement power management functionality including, but not limited to, one or more of the following features: individual hardware modules can be automatically clock-gated at idle to conserve dynamic power; each core clock can be gated when the core is not actively executing instructions due to execution of a wait for interrupt (“WFI”) / wait for event (“WFE”) instruction; each core can be independently powered; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. In at least one embodiment, one or more CPU(s) 1006 can further implement enhanced algorithms for managing power states where allowed power states and expected wake-up times are specified and hardware / microcode determines optimal power states for core, cluster, and CCPLEX inputs. In at least one embodiment, processing cores can support a simplified power state input sequence in software where work is offloaded to microcode.

[0165] In at least one embodiment, GPU(s) 1008 can include an integrated GPU (also referred to herein as an “iGPU”). In at least one embodiment, GPU(s) 1008 can be programmable and efficient for parallel workloads. In at least one embodiment, GPU(s) 1008 can use an enhanced tensor instruction set. In one embodiment, GPU(s) 1008 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“LI”) cache (e.g., an LI cache with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with 512 KB storage capacity). In at least one embodiment, GPU(s) 1008 can include at least eight streaming microprocessors. In at least one embodiment, GPU(s) 1008 can use a compute application programming interface (“API”). In at least one embodiment, GPU(s) 1008 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA model).

[0166] In at least one embodiment, GPU(s) 1008 can be power-optimized to achieve best performance in automotive and embedded use cases. For example, in at least one embodiment, GPU(s) 1008 can be fabricated on a finned field effect transistor (“FinFET”) circuit. In at least one embodiment, each streaming microprocessor can contain a plurality of mixed-precision processing cores divided into a plurality of blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a level zero (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, a streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads that mix compute and address operations. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.

[0167] In at least one embodiment, one or more GPU(s) 1008 can include high bandwidth memory (“HBM”) and / or 16 GB High Bandwidth Memory Second Generation (“HBM2”) memory subsystems to provide, in some examples, a peak memory bandwidth of about 900 GB / s. In at least one embodiment, in addition to, or alternatively to, HBM memory, synchronous graphics random access memory (“SGRAM”) can be used, such as graphics double data rate type five synchronous random-access memory (“GDDR5”).

[0168] In at least one embodiment, one or more GPU(s) 1008 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow one or more GPU(s) 1008 to directly access one or more CPU(s) 1006 page tables. In at least one embodiment, when a memory management unit (“MMU”) of a GPU of one or more GPU(s) 1008 experiences a miss, an address translation request can be sent to one or more CPU(s) 1006. In response, a CPU of one or more CPU(s) 1006 can look up a virtual-to-physical mapping for an address in its page tables and transmit a translation back to one or more GPU(s) 1008, in at least one embodiment. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both one or more CPU(s) 1006 and one or more GPU(s) 1008, simplifying programming of one or more GPU(s) 1008 and porting of applications to one or more GPU(s) 1008.

[0169] In at least one embodiment, one or more GPU(s) 1008 can include any number of access counters that can track how frequently one or more GPU(s) 1008 is accessing memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved into physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.

[0170] In at least one embodiment, one or more SoC(s) 1004 can include any number of caches 1012, including those described herein. For example, in at least one embodiment, one or more cache(s) 1012 can include a level three (“L3”) cache that can be available to and / or connected to one or more CPU(s) 1006 and one or more GPU(s) 1008. In at least one embodiment, one or more cache(s) 1012 can include a write-back cache that can track state of lines, for example, by using a cache coherence protocol (e.g., MESI, MSI, etc.). In at least one embodiment, although smaller cache sizes can be used, an L3 cache can include 4 MB of memory or more, according to embodiments.

[0171] In at least one embodiment, one or more SoC(s) 1004 can include one or more accelerator(s) 1014 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 1004 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement and offload some tasks of one or more GPU(s) 1008 (e.g., freeing up more cycles of one or more GPU(s) 1008 to perform other tasks). In at least one embodiment, one or more accelerator(s) 1014 can be used for target workloads (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.) that are stable enough to withstand the speedup test. In at least one embodiment, a CNN can include a region-based or region with convolutional neural network (“RCNN”) and a fast RCNN (e.g., as used for object detection) or other types of CNNs.

[0172] In at least one embodiment, one or more accelerators 1014 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). In at least one embodiment, one or more DLAs can include, without limitation, one or more Tensor Processing Units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can be further optimized for a particular set of neural network types and floating point operations and inferencing. In at least one embodiment, design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often significantly outperform CPUs. In at least one embodiment, one or more TPUs can perform several functions including support for INT8, INT16, and FP16 data types for features and weights, single instance convolution functionality, and post-processor functionality, for example. In at least one embodiment, one or more DLAs can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, as well as identification and detection, using data from microphones; CNNs for face recognition and vehicle owner identification using data from camera sensors; and / or CNNs for safety and / or safety related events.

[0173] In at least one embodiment, a DLA can perform any of functions of GPU(s) 1008, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or GPU(s) 1008 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations for CNNs on one or more DLAs, and leave other functions to GPU(s) 1008 and / or accelerator(s) 1014.

[0174] In at least one embodiment, one or more accelerators 1014 can include a programmable vision accelerator (“PVA”), which can be alternatively referred to herein as a computer vision accelerator. In at least one embodiment, one or more PVAs can be designed and configured to accelerate computer vision algorithms used for advanced driver assistance systems (“ADAS”) 1038, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, one or more PVAs can strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs can include, without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.

[0175] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any of cameras described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.

[0176] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 1006. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including, without limitation, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, without limitation, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0177] In at least one embodiment, vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, a PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core can include a processor subsystem, DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem can function as a primary processing engine for a PVA and can include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, a VPU core can include a digital signal processor, such as a single instruction multiple data (“SIMD”), very long instruction word (“VLIW”) digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.

[0178] In at least one embodiment, each vector processor can include an instruction cache and can be coupled to a dedicated memory. As a result, in at least one embodiment, each vector processor can be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA can be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA can execute a general purpose computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA can execute different computer vision algorithms on one image at a time, or even different algorithms on a sequence of images or portions of images. In at least one embodiment, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA, among other things. In at least one embodiment, a PVA can include additional error correcting code (“ECC”) memory to enhance overall system security.

[0179] In at least one embodiment, one or more accelerators 1014 can include on-chip computer vision networks and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 1014. In at least one embodiment, on-chip memory can include at least 4 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include on-chip computer vision networks that interconnect PVA and DLA to memory (e.g., using APB).

[0180] In at least one embodiment, on-chip computer vision networks can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards.

[0181] In at least one embodiment, one or more SoC 1004 can include real-time line-of-sight tracking hardware accelerators. In at least one embodiment, real-time line-of-sight tracking hardware accelerators can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.

[0182] In at least one embodiment, one or more accelerators 1014 have broad use for autonomous driving. In at least one embodiment, PVAs can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, capabilities of PVAs at low power and low latency match well with algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that can require predictable runtimes with low latency and low power. In at least one embodiment, PVAs can be designed to run classic computer vision algorithms, such as in vehicle 1000, as they can be efficient at object detection and integer math operations.

[0183] For example, in accordance with at least one embodiment of technology, a PVA is used to perform computer stereo vision. In at least one embodiment, semi-global matching based algorithms can be used in some examples, although this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use dynamic estimation / stereo matching in run-time (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0184] In at least one embodiment, a PVA can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, a PVA is used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.

[0185] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a confidence level for each object detection. In at least one embodiment, confidence level can be represented or interpreted as a probability, or as providing a relative “weight” of each detection relative to other detections. In at least one embodiment, a confidence measurement enables system to make further decisions as to which detections should be considered as true positive detections and not false positive detections. In at least one embodiment, system can set a threshold for confidence level, and only consider detections that exceed threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in vehicle automatically performing emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, neural network can take as its input at least some subset of parameters, such as bounding box size, ground plane estimates obtained (e.g., from another subsystem), outputs of one or more IMU sensors 1066 related to object’s vehicle 1000 direction, distance, 3D position estimates obtained from neural network and / or other sensors (e.g., one or more LIDAR sensors 1064 or one or more RADAR sensors 1060), etc.

[0186] In at least one embodiment, one or more SoC(s) 1004 can include one or more data storage(s) 1016 (e.g., memory). In at least one embodiment, one or more data storage(s) 1016 can be on-chip memory of one or more SoC(s) 1004 that can store neural networks to be executed on one or more GPU(s) 1008 and / or DLA. In at least one embodiment, one or more data storage(s) 1016 can have sufficient capacity to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data storage(s) 1016 can include L2 or L3 cache.

[0187] In at least one embodiment, one or more SoC(s) 1004 can include any number of processor(s) 1010 (e.g., embedded processors). In at least one embodiment, one or more processor(s) 1010 can include a boot and power management processor that can be a dedicated processor and subsystem to handle boot power and management functions and related security enforcement. In at least one embodiment, a boot and power management processor can be part of a one or more SoC(s) 1004 boot sequence and can provide run-time power management services. In at least one embodiment, a boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 1004 thermal and temperature sensor management, and / or one or more SoC(s) 1004 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 1004 can use ring oscillators to detect temperature of one or more CPU(s) 1006, one or more GPU(s) 1008, and / or one or more accelerator(s) 1014. In at least one embodiment, if a temperature is determined to exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoC(s) 1004 into a lower power state and / or put vehicle 1000 into a safe park pattern for the driver (e.g., cause vehicle 1000 to safely park).

[0188] In at least one embodiment, one or more processor(s) 1010 can also include a set of embedded processors that can function as an audio processing engine that can be an audio subsystem that can provide full hardware support for multi-channel audio to hardware through a number of interfaces as well as a broad and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0189] In at least one embodiment, one or more processor(s) 1010 can also include an always-on processor engine that can provide necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, a processor on an always-on processor engine can include, but is not limited to, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0190] In at least one embodiment, one or more processors 1010 can further include a safety cluster engine that includes, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 1010 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 1010 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.

[0191] In at least one embodiment, one or more processors 1010 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final video to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-angle cameras 1070, one or more surround cameras 1074, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, preferably, in-cabin monitoring camera sensors are monitored by a neural network running on another instance of SoC 1004 that is configured to identify cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change a vehicle’s infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.

[0192] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for simultaneous spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.

[0193] In at least one embodiment, video image compositor can also be configured to perform stereo correction on input stereoscopic lens frames. In at least one embodiment, when using an operating system desktop, video image compositor can also be used for user interface composition and one or more GPUs 1008 are not required to continuously render new surfaces. In at least one embodiment, when one or more GPUs 1008 are powered and active for 3D rendering, video image compositor can be used to offload one or more GPUs 1008 to improve performance and responsiveness.

[0194] In at least one embodiment, one or more SoCs in SoC(s) 1004 can also include mobile industry processor interface (“MIPI”) camera serial interfaces for receiving video and input from cameras, high-speed interfaces, and / or video input blocks that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoCs 1004 can also include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.

[0195] In at least one embodiment, one or more of SoCs 1004 can also include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. In at least one embodiment, one or more SoCs 1004 can be used to process data from cameras (e.g., connected over Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., one or more LIDAR sensors 1064, one or more RADAR sensors 1060, etc., which can be connected over Ethernet channels), data from bus 1002 (e.g., speed of vehicle 1000, steering wheel position, etc.), data from one or more GNSS sensors 1058 (e.g., connected over Ethernet bus or CAN bus), etc. In at least one embodiment, one or more of SoCs 1004 can also include dedicated high-performance mass storage controllers that can include their own DMA engines and can be used to free one or more CPUs 1006 from regular data management tasks.

[0196] In at least one embodiment, SoC(s) 1004 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technology for diversity and redundancy, which provides a platform that can provide a flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 1004 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 1014, when combined with CPU(s) 1006, GPU(s) 1008, and data storage device(s) 1016, can provide a fast, efficient platform for level 3-5 autonomous vehicles.

[0197] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C) to perform a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs typically cannot meet performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.

[0198] Embodiments described herein allow for simultaneous and / or sequential execution of multiple neural networks, and allow for combining results together to enable level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 1020) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that a neural network has not been specifically trained for. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing that semantic understanding to a path planning module running on a CPU Complex.

[0199] In at least one embodiment, multiple neural networks can be run simultaneously for a level 3, 4, or 5 drive. For example, in at least one embodiment, a warning sign consisting of a “Caution: flashing lights indicate icy conditions” sign with flashing lights can be interpreted independently or collectively by multiple neural networks. In at least one embodiment, the warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has already been trained), the text “flashing lights indicate icy conditions” can be interpreted by a second deployed neural network that informs vehicle’s path planning software (preferably executing on a CPU Complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be recognized by a third deployed neural network operating over multiple frames, informing vehicle’s path planning software of the existence (or non-existence) of flashing lights. In at least one embodiment, all three neural networks can be run simultaneously, e.g., within a DLA and / or on one or more GPU(s) 1008.

[0200] In at least one embodiment, a CNN for facial recognition and vehicle owner identification can use data from a camera sensor to identify presence of an authorized driver and / or owner of vehicle 1000. In at least one embodiment, when an owner approaches a driver door and opens a light, a normally open sensor processor engine can be used to unlock the vehicle, and, in a safe mode, when the owner leaves the vehicle, can be used to disable the vehicle. In this way, one or more SoC(s) 1004 provide safeguards against theft and / or carjacking.

[0201] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphones 1096 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1004 use a CNN to classify ambient and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify relative proximity of emergency vehicles (e.g., by using Doppler effect). In at least one embodiment, a CNN can also be trained to identify emergency vehicles for regions in which a vehicle is operating, as identified by one or more GNSS sensors 1058. In at least one embodiment, when operating in Europe, a CNN will seek to detect European sirens, while in North America, a CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, slow vehicle down, pull vehicle to side of road, stop, and / or idle vehicle until emergency vehicle passes, with assistance of one or more ultrasonic sensors 1062.

[0202] In at least one embodiment, vehicle 1000 can include one or more CPUs 1018 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 1004 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 1018 can include an X86 processor, such as one or more CPUs 1018 can be used to perform any of a variety of functions, such as including arbitrating inconsistent results between ADAS sensors and one or more SoCs 1004, and / or one or more supervisory controllers 1036 state and health and / or an on-chip information system (“information SoC”) 1030.

[0203] In at least one embodiment, vehicle 1000 can include one or more GPUs 1020 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 1004 via a high-speed interconnect (e.g., NVIDIA’s NVLINK channel). In at least one embodiment, one or more GPUs 1020 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 1000.

[0204] In at least one embodiment, vehicle 1000 can also include network interface 1024, which can include, without limitation, one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1024 can be used to enable wireless connectivity through an Internet cloud service (e.g., with servers and / or other network equipment) with other vehicles and / or computing devices (e.g., client devices of passengers). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 1000 and another vehicle and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a vehicle-to-vehicle communication link can be used to provide a direct link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 1000 with information about vehicles in a vicinity of vehicle 1000 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1000). In at least one embodiment, this aforementioned functionality can be part of a cooperative adaptive cruise control functionality of vehicle 1000.

[0205] In at least one embodiment, network interface 1024 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 1036 to communicate over wireless networks. In at least one embodiment, network interface 1024 can include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocol.

[0206] In at least one embodiment, vehicle 1000 can also include one or more data stores 1028, which can include, without limitation, off-chip (e.g., of SoC(s) 1004) storage. In at least one embodiment, one or more data stores 1028 can include, without limitation, one or more storage elements, including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.

[0207] In at least one embodiment, vehicle 1000 can also include one or more GNSS sensors 1058 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1058 can be used, including, for example and without limitation, a GPS using a universal serial bus (“USB”) connector with an Ethernet-to-serial (e.g., RS-232) bridge.

[0208] In at least one embodiment, vehicle 1000 can also include one or more RADAR sensors 1060. In at least one embodiment, one or more RADAR sensors 1060 can be used by vehicle 1000 for long-range vehicle detection, even in dark and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. In at least one embodiment, one or more RADAR sensors 1060 can use CAN bus and / or bus 1002 (e.g., to transmit data generated by one or more RADAR sensors 1060) for control and access to object tracking data, and in some examples an Ethernet channel can be accessed for raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more of RADAR sensors 1060 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 1060 are pulse Doppler RADAR sensors.

[0209] In at least one embodiment, RADAR sensor(s) 1060 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, long-range RADAR systems can provide a wide field of view enabled by two or more independent scans (e.g., out to 250 m). In at least one embodiment, RADAR sensor(s) 1060 can help distinguish between static and moving objects, and can be used by ADAS system 1038 for emergency brake assist and forward collision warning. In at least one embodiment, sensor(s) 1060 included in a long-range RADAR system can include, without limitation, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 1000 at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, other two antennas can expand the field of view, such that vehicles 1000 entering or leaving a lane can be quickly detected.

[0210] In at least one embodiment, as an example, a mid-range RADAR system can include, for example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 1060 designed to be mounted at either end of a rear bumper. When mounted at either end of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor the vehicle’s rearward direction and a blind spot near the vehicle. In at least one embodiment, a short-range RADAR system can be used in ADAS system 1038 for blind spot detection and / or lane change assist.

[0211] In at least one embodiment, vehicle 1000 can also include ultrasonic sensor(s) 1062. In at least one embodiment, ultrasonic sensor(s) 1062, which can be positioned in front, rear, and / or side locations of vehicle 1000, can be used for parking assist and / or to create and update occupancy grids. In at least one embodiment, a wide variety of ultrasonic sensors 1062 can be used, and different ultrasonic sensors 1062 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 1062 can operate at a functional safety level of ASIL B.

[0212] In at least one embodiment, vehicle 1000 can include one or more LIDAR sensors 1064. In at least one embodiment, one or more LIDAR sensors 1064 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 1064 can operate at a functional safety level of ASIL B. In at least one embodiment, vehicle 1000 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 1064 that can use Ethernet channels (e.g., provide data to a Gigabit Ethernet switch).

[0213] In at least one embodiment, one or more LIDAR sensors 1064 can be capable of providing a list of objects and their distances for a 360 degree field of view. In at least one embodiment, one or more LIDAR sensors 1064 that are commercially available can have an advertised range of approximately 100 m, have an accuracy of 2 cm - 3 cm, and support a 100 Mbps Ethernet connection, for example. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such embodiments, one or more LIDAR sensors 1064 can include small devices that can be embedded into front, rear, side, and / or corner locations of vehicle 1000. In at least one embodiment, one or more LIDAR sensors 1064, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200 m, even for low reflectivity objects. In at least one embodiment, a forward-facing one or more LIDAR sensors 1064 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0214] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. In at least one embodiment, 3D Flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200 m around vehicle 1000. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 1000 to an object. In at least one embodiment, flash LIDAR can allow for highly accurate and distortion-free images of surrounding environment to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 1000. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond class I (eye-safe) laser pulses per frame, and can capture reflected laser light as a 3D ranging point cloud and co-registered intensity data.

[0215] In at least one embodiment, vehicle 1000 can also include one or more IMU sensors 1066. In at least one embodiment, one or more IMU sensors 1066 can be located at a center of a rear axle of vehicle 1000. In at least one embodiment, one or more IMU sensors 1066 can include, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one magnetic compass, multiple magnetic compasses, and / or other sensor types. In at least one embodiment, such as in a six-axis application, one or more IMU sensors 1066 can include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as in a nine-axis application, one or more IMU sensors 1066 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer.

[0216] In at least one embodiment, one or more IMU sensors 1066 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude; in at least one embodiment, one or more IMU sensors 1066 can enable vehicle 1000 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from GPS to one or more IMU sensors 1066. In at least one embodiment, one or more IMU sensors 1066 and one or more GNSS sensors 1058 can be combined in a single integrated unit.

[0217] In at least one embodiment, vehicle 1000 can include one or more microphones 1096 placed within and / or around vehicle 1000. In at least one embodiment, additionally, microphone(s) 1096 can be used for emergency vehicle detection and identification.

[0218] In at least one embodiment, vehicle 1000 can also include any number of camera types, including one or more stereo cameras 1068, one or more wide-view cameras 1070, one or more infrared cameras 1072, one or more surround cameras 1074, one or more long-range cameras 1098, one or more mid-range cameras 1076, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around entire periphery of vehicle 1000. In at least one embodiment, type of cameras used depends on vehicle 1000. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 1000. In at least one embodiment, number of cameras deployed can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 1000 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communications, by way of example and without limitation. In at least one embodiment, cameras can be described in greater detail herein previously with reference to FIG. 6. Figure 10A and Figure 10B Each camera can be described in greater detail.

[0219] In at least one embodiment, vehicle 1000 can also include one or more vibration sensors 1042. In at least one embodiment, vibration sensor(s) 1042 can measure vibrations of components of vehicle 1000 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in road surface. In at least one embodiment, when two or more vibration sensors 1042 are used, differences between vibrations can be used to determine road surface friction or slippage (e.g., when there is a difference in vibration between a power driven axle and a free spinning axle).

[0220] In at least one embodiment, vehicle 1000 can include ADAS system 1038. In at least one embodiment, ADAS system 1038 can include, without limitation, an SoC. In at least one embodiment, ADAS system 1038 can include, without limitation, any number of adaptive / autonomous / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality, and combinations thereof.

[0221] In at least one embodiment, an ACC system can use one or more RADAR sensors 1060, one or more LIDAR sensors 1064, and / or any number of cameras. In at least one embodiment, an ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls distance to another vehicle immediately in front of vehicle 1000 and automatically adjusts speed of vehicle 1000 to maintain a safe distance from the vehicle in front. In at least one embodiment, a lateral ACC system performs distance keeping and suggests lane changes for vehicle 1000 when needed. In at least one embodiment, lateral ACC is relevant to other ADAS applications, such as LC and CW.

[0222] In at least one embodiment, a CACC system uses information from other vehicles that can be received from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet) via network interface 1024 and / or one or more wireless antennas 1026. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, V2V communications provide information about the vehicle immediately in front (e.g., the vehicle immediately in front of and in the same lane as vehicle 1000), while I2V communications provide information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V sources of information. In at least one embodiment, a CACC system can be more reliable with information about vehicles in front of vehicle 1000, and has potential to improve smoothness of traffic flow and reduce road congestion.

[0223] In at least one embodiment, an FCW system is designed to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system uses a forward-facing camera and / or one or more RADAR sensors 1060 coupled to a dedicated processor, a digital signal processor (“DSP”), an FPGA, and / or an ASIC that is electrically coupled to provide driver feedback such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an FCW system can provide a warning, for example, in the form of a sound, a visual warning, a vibration, and / or a quick brake pulse.

[0224] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, an AEB system can use one or more forward-facing cameras and / or one or more RADAR sensors 1060 coupled to a dedicated processor, a DSP, an FPGA, and / or an ASIC. In at least one embodiment, when an AEB system detects a hazard, it typically first warns a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, an AEB system can include techniques such as dynamic brake support and / or crash imminent braking.

[0225] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibration, to warn the driver when vehicle 1000 crosses lane markers. In at least one embodiment, an LDW system is not active when a driver indicates an intentional lane departure, such as by activating turn signals. In at least one embodiment, an LDW system can use a forward-facing camera coupled to a dedicated processor, a DSP, an FPGA, and / or an ASIC that is electrically coupled to provide driver feedback such as a display, a speaker, and / or a vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system. In at least one embodiment, if vehicle 1000 begins to deviate from a lane, an LKA system provides a steering input or braking to correct vehicle 1000.

[0226] In at least one embodiment, a BSW system detects and warns vehicle drivers of vehicles in a car’s blind spot. In at least one embodiment, a BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system can provide additional warnings when a driver uses turn signals. In at least one embodiment, a BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.

[0227] In at least one embodiment, a RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside of a rear camera range while vehicle 1000 is backing up. In at least one embodiment, a RCTW system includes an AEB system to ensure application of vehicle brakes to avoid a collision. In at least one embodiment, a RCTW system can use one or more rear-facing RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide driver feedback such as displays, speakers, and / or vibrating components.

[0228] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems warn the driver and allow that driver to decide whether a safety situation is truly present and take appropriate action. In at least one embodiment, in the event of a result conflict, vehicle 1000 itself decides whether to heed the results of a primary computer or a secondary computer (e.g., a first controller or a second controller of controller 1036). For example, in at least one embodiment, ADAS system 1038 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run redundant varieties of software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, outputs from ADAS system 1038 can be provided to a supervisory MCU. In at least one embodiment, if outputs from a primary computer and outputs from a secondary computer conflict, then a supervisory MCU decides how to reconcile the conflict to ensure safe operation.

[0229] In at least one embodiment, a host computer can be configured to provide a confidence score to a supervisory MCU to indicate a confidence of the host computer in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the host computer’s instructions regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, in cases where the confidence score does not satisfy a threshold, and in cases where the host computer and secondary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.

[0230] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which a secondary computer provides false alarms based at least in part on output from a host computer and output from a secondary computer. In at least one embodiment, a neural network in a supervisory MCU can learn when to trust output of a secondary computer, and when not to. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, a neural network in a supervisory MCU can learn when the FCW system identifies metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when there is a bicyclist or pedestrian present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 1004.

[0231] In at least one embodiment, ADAS system 1038 can include a secondary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, the secondary computer can use classic computer vision rules (if-then), and the presence of a neural network in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, a diversified implementation and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and non-identical software code running on a secondary computer provides consistent overall results, a supervisory MCU can be more confident that the overall results are correct, and that the bug in software or hardware on the host computer will not cause a significant error.

[0232] In at least one embodiment, output of ADAS system 1038 can be input into a perception module of host computer and / or a dynamic driving task module of host computer. For example, in at least one embodiment, if ADAS system 1038 indicates a forward collision warning due to an object directly in front, perception block can use this information when identifying the object. In at least one embodiment, as described herein, a secondary computer can have its own neural network that is trained to reduce risk of false positives.

[0233] In at least one embodiment, vehicle 1000 can also include infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, in at least one embodiment, infotainment system SoC 1030 can not be a SoC and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 1030 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, radio data system, vehicle related information such as fuel level, total covered distance, brake fluid level, oil level, doors open / close, air cleaner information, etc.) to vehicle 1000. For example, infotainment SoC 1030 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car, in-car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 1034, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 1030 can also be used to provide information (e.g., visual and / or audible) to a user of vehicle 1000, such as information from ADAS system 1038, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0234] In at least one embodiment, infotainment SoC 1030 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 1030 can communicate with other devices, systems, and / or components of vehicle 1000 over bus 1002. In at least one embodiment, infotainment SoC 1030 can be coupled to a supervisory MCU such that a GPU of an infotainment system can perform some autonomous driving functionality in the event of a failure of a host controller 1036 (e.g., a primary computer and / or a backup computer of vehicle 1000). In at least one embodiment, infotainment SoC 1030 can cause vehicle 1000 to enter a driver-to-safe-stop mode, as described herein.

[0235] In at least one embodiment, vehicle 1000 can also include an instrument cluster 1032 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument cluster, etc.). In at least one embodiment, instrument cluster 1032 can include, without limitation, a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 1032 can include, without limitation, any number and combination of gauges such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 1030 and instrument cluster 1032. In at least one embodiment, instrument cluster 1032 can be included as part of infotainment SoC 1030, and vice versa.

[0236] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 1 B, and 5. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 1 B, and 5. Figure 10C In at least one embodiment, inference and / or training logic 115 can be used in system

[0237] Figure 10D are cloud-based servers in accordance with at least one embodiment Figure 10AFIG. 10 illustrates a diagram of a system 1076 in communication between autonomous vehicles 1000. In at least one embodiment, system 1076 can include, without limitation, one or more servers 1078, one or more networks 1090, and any number and type of vehicles, including vehicles 1000. In at least one embodiment, one or more servers 1078 can include, without limitation, a plurality of GPUs 1084(A)-1084(H) (collectively referred to herein as GPUs 1084), PCIe switches 1082(A)-1082(D) (collectively referred to herein as PCIe switches 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPUs 1080). In at least one embodiment, GPUs 1084, CPUs 1080, and PCIe switches 1082 can be interconnected with high-speed connection lines such as, without limitation, NVLink interfaces 1088 developed by NVIDIA and / or PCIe connections 1086. In at least one embodiment, GPUs 1084 are connected by NVLink and / or NVSwitch SoC connections, and GPUs 1084 and PCIe switches 1082 are connected by PCIe interconnects. Although eight GPUs 1084, two CPUs 1080, and four PCIe switches 1082 are illustrated, this is not intended to be limiting. In at least one embodiment, each of one or more servers 1078 can include, without limitation, any number of GPUs 1084, CPUs 1080, and / or PCIe switches 1082 in any combination. For example, in at least one embodiment, one or more servers 1078 can each include eight, sixteen, thirty-two, and / or more GPUs 1084.

[0238] In at least one embodiment, one or more servers 1078 can receive, over one or more networks 1090 and from vehicles, image data representative of images showing unexpected or changing road conditions, such as road work that has recently begun. In at least one embodiment, one or more servers 1078 can transmit, over one or more networks 1090 and to vehicles, updated equalization neural networks 1092, and / or map information 1094 including, without limitation, information about traffic and road conditions. In at least one embodiment, updates to map information 1094 can include, without limitation, updates to HD map 1022, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural networks 1092 and / or map information 1094 can be the result of new training and / or experience represented in data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 1078 and / or other servers).

[0239] In at least one embodiment, one or more servers 1078 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., where associated neural networks benefit from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, no training data is labeled and / or pre-processed (e.g., where associated neural networks do not require supervised learning). In at least one embodiment, once a machine learning model is trained, it can be used by vehicles (e.g., transmitted to vehicles over one or more networks 1090, and / or machine learning models can be used by one or more servers 1078 to monitor vehicles remotely.

[0240] In at least one embodiment, one or more servers 1078 can receive data from vehicles and apply it to up-to-date, real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 1078 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 1084, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1078 can include deep learning infrastructure of data centers powered using CPUs.

[0241] In at least one embodiment, deep learning infrastructure of one or more servers 1078 can be capable of fast, real-time inference, and can use this capability to assess and validate health of processors, software, and / or related hardware in vehicles 1000. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicles 1000, such as sequences of images and / or objects that vehicles 1000 have located in that sequence of images (e.g., through computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural networks to identify objects and compare them to objects identified by vehicles 1000, and, if results do not match and deep learning infrastructure concludes that Al in vehicles 1000 is malfunctioning, one or more servers 1078 can send a signal to vehicles 1000 instructing a failsafe computer of vehicles 1000 to take control, notify passengers, and complete a safe parking operation.

[0242] In at least one embodiment, one or more servers 1078 can include one or more GPUs 1084 and one or more programmable inference accelerators (such as NVIDIA’s TensorRT 3 devices). In at least one embodiment, a combination of GPU-driven servers and inference-accelerated servers can make real-time responses possible. In at least one embodiment, CPU-, FPGA-, and other processor-driven servers can be used for inference, for example in cases where performance is less critical. In at least one embodiment, hardware structure 115 is used to perform one or more embodiments. Details regarding hardware structure 115 are provided herein. Figure 1A and / or Figure 1B Details regarding hardware structure 115 are provided herein.

[0243] Computer system

[0244] Figure 11 is a block diagram illustrating an exemplary computer system that can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units to execute an instruction, according to at least one embodiment. In at least one embodiment, according to the present disclosure, such as embodiments described herein, computer system 1100 can include, without limitation, components such as processor 1102 that includes execution units to perform logic to execute an algorithm for process data. In at least one embodiment, computer system 1100 can include a processor, such as a Pentium®, Core®, and / or Xeon® processor family, Intel® XScaleTM and / or StrongARM™, XScaleTM and / or StrongARMTM, Core TM or Nervana TM microprocessor, although other systems (including PCs, workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, computer system 1100 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems, embedded software, and / or graphical user interfaces can also be used.

[0245] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a DSP, a system on a chip ("SOC"), a networked system, a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can perform one or more instructions in accordance with at least one embodiment.

[0246] In at least one embodiment, computer system 1100 can include, but is not limited to, processor 1102, which can include, but is not limited to, one or more execution units 1108 to perform, e.g., machine learning model training and / or inferencing, in accordance with techniques described herein. In at least one embodiment, computer system 1100 is a single processor desktop or server system, but in another embodiment, computer system 1100 can be a multiprocessor system. In at least one embodiment, processor 1102 can include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combo of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1102 can be coupled to a processor bus 1110 that can transmit data signals between processor 1102 and other components in computer system 1100.

[0247] In at least one embodiment, processor 1102 can include, but is not limited to, level 1 ("Ll") internal cache memory ("cache") 1104. In at least one embodiment, processor 1102 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 1102's external. Other embodiments can include a combination of internal and external cache memory depending on the specific implementation and requirements. In at least one embodiment, register file 1106 can store different types of data such as integer, floating point, status, and instruction pointer registers in various registers within processor 1102.

[0248] In at least one embodiment, execution unit 1108 includes, without limitation, logic to perform integer and floating point operations, including, but not limited to, logic to execute integer and floating point instructions 1109. In at least one embodiment, processor 1102 can also include microcode ("ucode") read-only memory ("ROM"), which stores microcode for certain macroinstructions. In at least one embodiment, execution unit 1108 can also include logic to handle a packed data instruction set 1109. In at least one embodiment, by including this logic in execution unit 1108, processor 1102 can use packed data instructions for many multimedia applications that were originally coded using a different instruction set, such as MMX™ instructions set.

[0249] In at least one embodiment, execution unit 1108 can also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuits. In at least one embodiment, computer system 1100 can include, without limitation, memory 1120. In at least one embodiment, memory 1120 can be a Dynamic Random Access Memory ("DRAM") device, a Static Random Access Memory ("SRAM") device, a flash memory device, or another memory device. In at least one embodiment, memory 1120 can store instruction(s) 1119 and / or data 1121 represented by data signals that can be executed by processor 1102.

[0250] In at least one embodiment, a system logic chip can be coupled to processor bus 1110 and memory 1120. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 1116 and processor 1102 can communicate with MCH 1116 via processor bus 1110. In at least one embodiment, MCH 1116 can provide a high bandwidth memory path 1118 to memory 1120 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 1116 can direct data signals between processor 1102, memory 1120, and other components in computer system 1100, and

[0251] In at least one embodiment, computer system 1100 can use system I / O interface 1122 as a proprietary hub interface bus to couple MCH 1116 to I / O controller hub (“ICH”) 1130. In at least one embodiment, ICH 1130 can provide a direct connection to some I / O devices and indirectly through a high-speed I / O bus. In at least one embodiment, the high-speed I / O bus can include, without limitation, a high-speed I / O bus, for connecting peripheral devices to memory 1120, chipset, and processor 1102. Examples can include, without limitation, a data storage 1124, a legacy I / O controller 1123 containing user input and keyboard interfaces 1125, a serial expansion port 1127, such as a Universal Serial Bus (USB) port, and a network controller 1134. In at least one embodiment, data storage 1124 can include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0252] In at least one embodiment, Figure 11 A system including interconnected hardware devices or “chips” can be shown, while in other embodiments, Figure 11 A SoC can be shown. In at least one embodiment, Figure 11The devices illustrated in FIG. 12 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1200 are interconnected using a compute express link (CXL) interconnect.

[0253] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 1 B, and 3. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 1 B, and 3. Figure 11 In at least one embodiment, inference and / or training logic 115 can be used in a system to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0254] Figure 12 is a block diagram illustrating an electronic device 1200 for utilizing a processor 1210, in accordance with at least one embodiment. In at least one embodiment, electronic device 1200 can be, for example and without limitation, a laptop, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0255] In at least one embodiment, electronic device 1200 can include, without limitation, a processor 1210 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1210 is coupled using a bus or interface, such as an I2C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, processor 1210 is coupled to one or more input devices 1202 and one or more output devices 1204. Figure 12 A system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 12 An exemplary SoC can be shown. In at least one embodiment, Figure 12 The devices illustrated in FIG. 12 can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 12 one or more components of computer system 1200 are interconnected using a compute express link (CXL) interconnect.

[0256] In at least one embodiment, Figure 12The components can include a display 1224, a touch screen 1225, a touch pad 1230, a near field communication unit ("NFC") 1245, a sensor hub 1240, a thermal sensor 1246, an express chip set ("EC") 1235, a trusted platform module ("TPM") 1238, a BIOS / firmware / flash memory ("BIOS, FW Flash") 1222, a DSP 1260, a drive 1220 such as a solid state disk ("SSD") or a hard disk drive ("HDD"), a wireless local area network unit ("WLAN") 1250, a Bluetooth unit 1252, a wireless wide area network unit ("WWAN") 1256, a Global Positioning System ("GPS") unit 1255, a camera ("USB 3.0 camera") 1254 such as a USB 3.0 camera, and / or a low power double data rate ("LPDDR") memory unit ("LPDDR3") 1215 implemented in, for example, an LPDDR3 standard. These components can each be implemented in any suitable manner.

[0257] In at least one embodiment, other components can be communicatively coupled to processor 1210 by components described herein. In at least one embodiment, an accelerometer 1241, an ambient light sensor ("ALS") 1242, a compass 1243, and a gyroscope 1244 can be communicatively coupled to sensor hub 1240. In at least one embodiment, a thermal sensor 1239, a fan 1237, a keyboard 1236, and a touch pad 1230 can be communicatively coupled to EC 1235. In at least one embodiment, a speaker 1263, a headphone 1264, and a microphone ("mic") 1265 can be communicatively coupled to an audio unit ("audio codec and class D amplifier") 1262, which in turn can be communicatively coupled to DSP 1260. In at least one embodiment, audio unit 1262 can include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 1257 can be communicatively coupled to WWAN unit 1256. In at least one embodiment, components such as WLAN unit 1250 and Bluetooth unit 1252, as well as WWAN unit 1256, can be implemented as a next generation form factor ("NGFF").

[0258] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 8, 9, and 10. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 8, 9, and 10. Figure 12In use, for inferencing or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0259] Figure 13 A computer system 1300 according to at least one embodiment is shown. In at least one embodiment, computer system 1300 is configured to implement various processes and methods described throughout this disclosure.

[0260] In at least one embodiment, computer system 1300 includes, without limitation, at least one central processing unit (“CPU”) 1302 that is connected to a communication bus 1310 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), peripheral component interconnect express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1300 includes, without limitation, a main memory 1304 and control logic (e.g., implemented in hardware, software, or a combination thereof) and data can be stored in main memory 1304 in the form of random access memory (“RAM”).

[0261] In at least one embodiment, computer system 1300 includes, in at least one embodiment without limitation, an input device 1308, parallel processing system 1312, and display device 1306, which can be implemented using a conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from input device 1308 such as a keyboard, mouse, touchpad, microphone, or the like. In at least one embodiment, each of the modules described herein can be located on a single semiconductor platform.

[0262] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A and / or 1 B. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A and / or 1 B. In at least one embodiment, inference and / or training logic 115 can be used in system Figure 13 In use, for inferencing or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0263] Figure 14 A computer system 1400 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 1400 includes, without limitation, a computer 1410 and a USB stick 1420. In at least one embodiment, computer 1410 can include, without limitation, any number and type of processor (not shown) and memory (not shown). In at least one embodiment, computer 1410 includes, without limitation, a server, a cloud instance, a laptop computer, and a desktop computer.

[0264] In at least one embodiment, USB stick 1420 includes, without limitation, a processing unit 1430, a USB interface 1440, and USB interface logic 1450. In at least one embodiment, processing unit 1430 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1430 can include, without limitation, any number and type of processing core (not shown). In at least one embodiment, processing unit 1430 includes an application-specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, processing unit 1430 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, processing unit 1430 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.

[0265] In at least one embodiment, USB interface 1440 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1440 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1440 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1450 can include any number and type of logic that enables processing unit 1430 to interface with a device (e.g., computer 1410) via USB interface 1440.

[0266] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 7, 8A, 8B, 9, and 10. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided. In at least one embodiment, inference and / or training logic 115 can be used in system Figure 14 to infer or predict operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0267] Figure 15A An exemplary architecture is shown in which a plurality of GPUs 1510(1)- 1510(N) are communicatively coupled to a plurality of multi-core processors 1505(1)- 1505(M) over high-speed links 1540(1)-1540(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, high-speed links 1540(1)-1540(N) support a communication throughput of 4GB / s, 30GB / s, 80GB / s or higher. In at least one embodiment, various interconnect protocols can be used including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, “N” and “M” represent positive integers, which can vary from figure to figure.

[0268] In addition, in at least one embodiment, two or more GPUs 1510 are interconnected by high-speed links 1529(1)-1529(2), which can be implemented using similar or different protocols / links than those used for high-speed links 1540(1)-1540(N). Similarly, two or more multi-core processors 1505 can be connected by high-speed link 1528, which can be a symmetric multi-processor (SMP) bus that runs at 20GB / s, 30GB / s, 120GB / s or higher. Alternatively, similar protocols / links (e.g., over common interconnect fabric) can be used to Figure 15A all communication between various system components shown in FIG. 15.

[0269] In at least one embodiment, each multi-core processor 1505 is communicatively coupled to processor memory 1501(1)-1501(M) via memory interconnects 1526(1)- 1526(M), respectively, and each GPU 1510(1)-1510(N) is communicatively coupled to GPU memory 1520(1)-1520(N) by GPU memory interconnects 1550(1)-1550(N), respectively. In at least one embodiment, memory interconnects 1526 and 1550 can utilize similar or different memory access technologies. By way of example, but not limitation, processor memory 1501(1)-1501(M) and GPU memory 1520 can be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or can be non-volatile memory, such as 3D XPoint or Nano-Ram. In at least one embodiment, certain portions of processor memory 1501 can be volatile memory, while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0270] As described herein, although various multi-core processors 1505 and GPUs 1510 can be physically coupled to particular memories 1501, 1520, respectively, and / or can implement a unified memory architecture in which a virtual system address space (also referred to as an “effective address” space) is distributed among various physical memories. For example, processor memories 1501(1)-1501(M) can each contain 64 GB of system memory address space, and GPU memories 1520(1)-1520(N) can each contain 32 GB of system memory address space, resulting in a total of 256 GB of addressable memory size when M = 2 and N = 4. N and M can also be other values.

[0271] Figure 15B Additional details for interconnect between multi-core processor 1507 and graphics acceleration module 1546 are shown according to one exemplary embodiment. In at least one embodiment, graphics acceleration module 1546 can include one or more GPU chips integrated on a line card that is coupled via a high-speed link 1540 (e.g., a PCIe bus, NVLink, etc.) to processor 1507. In at least one embodiment, graphics acceleration module 1546 can alternatively be integrated on a package or chip with processor 1507.

[0272] In at least one embodiment, processor 1507 includes a number of cores 1560A-1560D, each having a translation lookaside buffer (“TLB”) 1561A-1561D and one or more caches 1562A-1562D. In at least one embodiment, cores 1560A-1560D can include various other components not shown for purposes of performing instructions and processing data. In at least one embodiment, caches 1562A-1562D can include level 1 (LI) and level 2 (L2) caches. Additionally, one or more shared caches 1556 can be included in caches 1562A-1562D and shared by groups of cores 1560A-1560D. For example, one embodiment of processor 1507 includes 24 cores, each with its own LI cache, twelve shared L2 caches, and twelve shared L3 caches. In that embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, processor 1507 and graphics acceleration module 1546 are connected with system memory 1514, which can include processor memories 1501(1)-1501(M) in Figure 15A

[0273] ​In at least one embodiment, consistency for data and instructions stored in respective caches 1562A-1562D, 1556, and system memory 1514 is maintained through inter-core communications over coherence bus 1564. In at least one embodiment, each cache can have cache coherence logic / circuitry associated therewith to communicate over coherence bus 1564 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snoop protocol is implemented over coherence bus 1564 to snoop cache accesses.

[0274] In at least one embodiment, agent circuit 1525 communicatively couples graphics acceleration module 1546 to coherence bus 1564, allowing graphics acceleration module 1546 to participate in a cache coherence protocol as a peer to cores 1560A-1560D. In particular, in at least one embodiment, interface 1535 provides connectivity from graphics acceleration module 1546 to agent circuit 1525 over high-speed link 1540, and interface 1537 connects graphics acceleration module 1546 to high-speed link 1540.

[0275] In at least one embodiment, accelerator integration circuit 1536 provides cache management, memory access, context management, and interrupt management services on behalf of graphics acceleration module’s plurality of graphics processing engines 1531(1)-1531(N). In at least one embodiment, graphics processing engines 1531(1)-1531(N) can each comprise a separate GPU. In at least one embodiment, graphics processing engines 1531(1)-1531(N) alternatively can comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, graphics acceleration module 1546 can be a GPU with a plurality of graphics processing engines 1531(1)-1531(N) or graphics processing engines 1531(1)-1531(N) can be individual GPUs integrated on a common package, line card, or chip.

[0276] In at least one embodiment, accelerator integration circuit 1536 includes a memory management unit (MMU) 1539 to provide for translation of virtual addresses into physical addresses, as is known to those of ordinary skill in the art. In at least one embodiment, MMU 1539 can include address translation lookaside buffers (TLBs) and instruction TLBs (ITLBs) to improve translation performance. In at least one embodiment, MMU 1539 can be used to implement a translation buffer cache of address translations used by graphics processing engine(s) 1531(1)-1531(N). In at least one embodiment, MMU 1539 includes a memory access protocol to access system memory 1514. In at least one embodiment, MMU 1539 can also include a translation lookaside buffer (“TLB”) (not shown) to cache virtual / valid to physical / real address translations. In at least one embodiment, cache 1538 can store commands and data for efficient access by graphics processing engines 1531(1)-1531(N). In at least one embodiment, data stored in cache 1538 and graphics memory 1533(1)-1533(M) can be kept coherent with core caches 1562A-1562D, 1556, and system memory 1514, possibly using fetch unit 1544. As described earlier, this task can be accomplished via proxy circuit 1525 on behalf of cache 1538 and graphics memory 1533(1)-1533(M) (e.g., sending updates related to modifications / accesses to a cache line on processor caches 1562A-1562D, 1556 to cache 1538 and receiving updates from cache 1538).

[0277] In at least one embodiment, a set of registers 1545 store context data for threads executed by graphics processing engines 1531(1)-1531(N), and context management circuit 1548 manages thread contexts. For example, context management circuit 1548 can perform save and restore operations to save and restore a context of an individual thread during a context switch (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by a graphics processing engine). For example, context management circuit 1548 can store current register values to a designated area in memory (e.g., identified by a context pointer) at a context switch. The register values can then be restored when returning to the context. In at least one embodiment, interrupt management circuit 1547 receives and processes interrupts received from system devices.

[0278] In at least one embodiment, MMU 1539 translates virtual / effective addresses from graphics processing engines 1531 into real / physical addresses in system memory 1514. In at least one embodiment, accelerator integration circuit 1536 supports a number of graphics accelerator modules 1546 and / or other accelerator devices (e.g., 4, 8, 16). In at least one embodiment, graphics accelerator modules 1546 can be dedicated to a single application executing on processor 1507 or can be shared between multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources of graphics processing engines 1531(1)-1531(N) are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into “slices” that are assigned to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.

[0279] In at least one embodiment, accelerator integration circuit 1536 performs as a bridge to system for a system of graphics acceleration modules 1546 and provides address translation and system memory cache services. Further, in at least one embodiment, accelerator integration circuit 1536 can provide virtualization facilities for a host processor to manage virtualization of graphics processing engines 1531(1)-1531(N), interrupts, and memory management.

[0280] In at least one embodiment, because hardware resources of graphics processing engines 1531(1)-1531(N) are explicitly mapped to real address space seen by host processor 1507, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integration circuit 1536 is to physically separate graphics processing engines 1531(1)-1531(N) so that they appear as independent units to a system.

[0281] In at least one embodiment, one or more graphics memories 1533(1)-1533(M) are coupled to each of graphics processing engines 1531(1)-1531(N), respectively, and N=M. In at least one embodiment, graphics memories 1533(1)-1533(M) store instructions and data for processing by each of graphics processing engines 1531(1)-1531(N). In at least one embodiment, graphics memories 1533(1)-1533(M) can be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3DXPoint or Nano-Ram.

[0282] In at least one embodiment, to reduce data traffic on high-speed link 1540, biasing techniques are used to ensure that data stored in graphics memory 1533(1)-1533(M) is that which is most frequently used by graphics processing engines 1531(1)-1531(N) and is preferably not used (at least not frequently) by cores 1560A-1560D. Similarly, in at least one embodiment, biasing mechanisms attempt to keep data needed by cores (and preferably not graphics processing engines 1531(1)-1531(N)) in caches 1562A-1562D, 1556, and system memory 1514.

[0283] Figure 15C Another exemplary embodiment is shown in which accelerator integration circuit 1536 is integrated within processor 1507. In this embodiment, graphics processing engines 1531(1)-1531(N) communicate directly over high-speed link 1540 to accelerator integration circuit 1536 via interface 1537 and interface 1535 (which can be any form of bus or interface protocol, as desired) in at least one embodiment, accelerator integration circuit 1536 can perform similar operations to those described with respect to Figure 15B accelerator integration circuit 1536. However, due to its close proximity to coherence bus 1564 and caches 1562A-1562D, 1556, it can have higher throughput. In at least one embodiment, accelerator integration circuit supports different programming models, including a dedicated process programming model (no graphics acceleration module virtualization) and a shared programming model (with virtualization), which can include programming models controlled by accelerator integration circuit 1536 and programming models controlled by graphics acceleration module 1546.

[0284] In at least one embodiment, graphics processing engines 1531(1)-1531(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1531(1)-1531(N), providing virtualization within a VM / partition.

[0285] In at least one embodiment, graphics processing engines 1531(1)-1531(N) can be shared by multiple VM / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize graphics processing engines 1531(1)-1531(N) to allow access by each operating system. In at least one embodiment, for a single-partition system without a hypervisor, an operating system owns graphics processing engines 1531(1)-1531(N). In at least one embodiment, an operating system can virtualize graphics processing engines 1531(1)-1531(N) to provide access to each process or application.

[0286] In at least one embodiment, graphics acceleration module 1546 or individual graphics processing engines 1531(1)-1531(N) use a process handle to select a process element. In at least one embodiment, a process element is stored in system memory 1514 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation-specific value provided to a host process when registering its context with a graphics processing engine 1531(1)-1531(N) (i.e., calling system software to add a process element to a process element linked list). In at least one embodiment, a lower 16 bits of a process handle can be an offset of a process element in a process element linked list.

[0287] Figure 15D An exemplary accelerator integration slice 1590 is shown. In at least one embodiment, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 1536. In at least one embodiment, an application is an effective address space 1582 in system memory 1514 that stores a process element 1583. In at least one embodiment, process element 1583 is stored in response to a GPU call 1581 from an application 1580 executing on processor 1507. In at least one embodiment, process element 1583 contains process state for a respective application 1580. In at least one embodiment, a work descriptor (WD) 1584 contained in process element 1583 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 1584 is a pointer to a job request queue in an application’s effective address space 1582.

[0288] In at least one embodiment, graphics acceleration module 1546 and / or individual graphics processing engines 1531(1)-1531(N) can be shared by all or a subset of processes in a system. In at least one embodiment, can include infrastructure for setting up process state and sending a WD 1584 to graphics acceleration module 1546 to start a job in a virtualized environment.

[0289] In at least one embodiment, a dedicated process programming model is implementation specific. In at least one embodiment, in this model, a single process owns a graphics acceleration module 1546 or individual graphics processing engines 1531. In at least one embodiment, when a graphics acceleration module 1546 is owned by a single process, a hypervisor initializes an accelerator integration circuit for the owned partition, and an operating system initializes an accelerator integration circuit 1536 for the owned process when a graphics acceleration module 1546 is assigned.

[0290] In at least one embodiment, in operation, a WD fetch unit 1591 in an accelerator integration slice 1590 fetches a next WD 1584 that includes an indication of work to be completed by one or more graphics processing engines of a graphics acceleration module 1546. In at least one embodiment, data from WD 1584 can be stored in registers 1545 and used by MMU 1539, interrupt management circuit 1547, and / or context management circuit 1548, as illustrated. For example, one embodiment of MMU 1539 includes segment / page walk circuitry for accessing segment / page tables 1586 within an OS virtual address space 1585. In at least one embodiment, interrupt management circuit 1547 can handle interrupt events 1592 received from a graphics acceleration module 1546. In at least one embodiment, effective addresses 1593 generated by graphics processing engines 1531(1)-1531(N) are translated to real addresses by MMU 1539 when performing graphics operations.

[0291] In at least one embodiment, registers 1545 are replicated for each graphics processing engine 1531(1)-1531(N) and / or graphics acceleration module 1546, and can be initialized by a hypervisor or operating system. In at least one embodiment, each of these replicated registers can be included in an accelerator integration slice 1590. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.

[0292]

[0293] Exemplary registers that can be initialized by an operating system are shown in Table 2.

[0294]

[0295]

[0296] In at least one embodiment, each WD 1584 is specific to a particular graphics acceleration module 1546 and / or graphics processing engine 1531(1)-1531(N). In at least one embodiment, it contains all information that graphics processing engine 1531(1)-1531(N) needs to complete the work assigned to WD 1584, or it can be a pointer to a memory location where application has set up a command queue of work to be completed.

[0297] Figure 15E Additional details of one exemplary embodiment of a shared model are shown. This embodiment includes a hypervisor real address space 1598 in which a list of process elements 1599 is stored. In at least one embodiment, hypervisor real address space 1598 is accessible via hypervisor 1596, which virtualizes graphics acceleration module engines for operating system 1595.

[0298] In at least one embodiment, a shared programming model allows all processes or a subset of processes from all partitions or a subset of partitions in a system to use a graphics acceleration module 1546. In at least one embodiment, there are two programming models in which a graphics acceleration module 1546 is shared by multiple processes and partitions, i.e., time-sliced sharing and graphics-directed sharing.

[0299] In at least one embodiment, in this model, system hypervisor 1596 owns graphics acceleration module 1546 and makes its functionality available to all operating systems 1595. In at least one embodiment, for graphics acceleration module 1546 to support virtualization by system hypervisor 1596, graphics acceleration module 1546 can adhere to certain requirements, such as (1) application’s job requests must be autonomous (i.e., no state needs to be maintained between jobs), or graphics acceleration module 1546 must provide a context save and restore mechanism, (2) graphics acceleration module 1546 guarantees that an application’s job request completes within a specified amount of time, including any translation faults, or graphics acceleration module 1546 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 1546 processes must be ensured when operating in a directed sharing programming model.

[0300] In at least one embodiment, application 1580 is required to use a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP) for an operating system 1595 system call. In at least one embodiment, the graphics acceleration module type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system specific value. In at least one embodiment, the WD is formatted specifically for a graphics acceleration module 1546 and can take the form of a graphics acceleration module 1546 command, a valid address pointer to a user defined structure, a valid address pointer to a queue of commands, or any other data structure describing work to be done by a graphics acceleration module 1546.

[0301] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to how an application program sets the AMR. In at least one embodiment, if an accelerator integration circuit 1536 (not shown) and graphics acceleration module 1546 implementation does not support a user authority mask override register (UAMOR), then the operating system can apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. In at least one embodiment, the hypervisor 1596 can selectively apply the current authority mask override register (AMOR) value before placing the AMR in the process element 1583. In at least one embodiment, the CSRP is one of registers 1545 that contains a valid address of an area in application’s effective address space 1582 for a graphics acceleration module 1546 to save and restore context state. In at least one embodiment, this pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be a fixed system memory.

[0302] Upon receiving the system call, operating system 1595 can verify that application 1580 has registered and is granted authority to use graphics acceleration module 1546. Operating system 1595 then uses the information shown in Table 3 to call hypervisor 1596, in at least one embodiment.

[0303]

[0304]

[0305] In at least one embodiment, upon receiving a hypervisor call, hypervisor 1596 verifies that operating system 1595 is registered and granted permission to use graphics acceleration module 1546. Hypervisor 1596 then, in at least one embodiment, places process element 1583 into a process element linked list of a corresponding graphics acceleration module 1546 type. In at least one embodiment, process elements can include information as illustrated in Table 4.

[0306]

[0307]

[0308] In at least one embodiment, hypervisor initializes a plurality of accelerator integration slice 1590 registers 1545.

[0309] As Figure 15F illustrated, in at least one embodiment, a unified memory is used that is addressable via a common virtual memory address space for accessing physical processor memory 1501(1)-1501(N) and GPU memory 1520(1)-1520(N). In this implementation, operations executing on GPU 1510(1)-1510(N) utilize the same virtual / effective memory address space to access processor memory 1501(1)-1501(M) and vice versa, simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1501(1), a second portion to a second processor memory 1501(N), a third portion to GPU memory 1520(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as an effective address space) is thus distributed among processor memories 1501 and GPU memories 1520 each, allowing any processor or GPU to access that memory with a virtual address that maps to any physical memory.

[0310] In at least one embodiment, bias / coherence management circuitry 1594A-1594E within one or more MMU 1539A-1539E ensures cache coherency between one or more host processors (e.g., 1505) and caches of GPU 1510, and implements bias techniques that dictate a bias of physical memory where certain types of data should be stored. In at least one embodiment, while multiple instances of bias / coherence management circuitry 1594A-1594E are shown in Figure 15F FIG. 15B, bias / coherence circuitry can be implemented within MMU(s) of one or more host processors 1505 and / or within accelerator integration circuit 1536.

[0311] One embodiment allows GPU memory 1520 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability to access GPU memory 1520 as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. In at least one embodiment, this arrangement allows software of host processor 1505 to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies include driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 1520 without cache coherency overhead can be critical to the execution time of offloaded computations. In at least one embodiment, for example, with heavy streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 1510. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can all play a role in determining the effectiveness of GPU offload.

[0312] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granularity structure (e.g., controlled at the granularity of a memory page) that includes a 1 or 2 bits per GPU-attached memory page. In at least one embodiment, with or without a bias cache in GPU 1510 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in the stolen memory range of one or more GPU memories 1520. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.

[0313] In at least one embodiment, prior to actually accessing GPU memory, the bias table entry associated with each access to GPU-attached memory 1520 is accessed, causing the following operations. In at least one embodiment, local requests from GPU 1510 that find their pages in GPU bias are forwarded directly to corresponding GPU memory 1520. In at least one embodiment, local requests from GPU that find their pages in host bias are forwarded to processor 1505 (e.g., over a high-speed link as described herein). In at least one embodiment, requests from processor 1505 that find requested pages in host processor bias complete the request similarly to a normal memory read. Alternatively, requests that point to GPU bias pages can be forwarded to GPU 1510. In at least one embodiment, if GPU is not currently using the page, GPU can then migrate the page to host processor bias. In at least one embodiment, the bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in limited cases purely hardware-based mechanisms.

[0314] In at least one embodiment, a mechanism for changing bias state employs an API call (e.g., OpenCL) that in turn invokes a device driver of a GPU that in turn sends a message (or causes a command descriptor to be enqueued) to the GPU, directing the GPU to change the bias state and, in certain migrations, to perform a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for a migration from host processor 1505 bias to GPU bias, but not for the reverse migration.

[0315] In at least one embodiment, cache coherency is maintained by temporarily rendering GPU bias pages that host processor 1505 cannot cache. In at least one embodiment, to access these pages, processor 1505 can request access from GPU 1510, which can or can not grant access immediately. Thus, in at least one embodiment, to reduce communication between processor 1505 and GPU 1510, it is beneficial to ensure that GPU bias pages are pages that are needed by the GPU and not by host processor 1505, and vice versa.

[0316] One or more hardware structures 115 are used to perform one or more embodiments. Details regarding one or more hardware structures 115 can be found in this document in connection with Figure 1A and / or Figure 1B Details regarding one or more hardware structures 115 are provided.

[0317] Figure 16Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrations, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.

[0318] Figure 16 is a block diagram illustrating an exemplary system on a chip integrated circuit 1600 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1600 includes one or more application processor(s) 1605 (e.g., CPUs), at least one graphics processor 1610, and can additionally include an image processor 1615 and / or a video processor 1620, any of which can be a modular IP core. In at least one embodiment, integrated circuit 1600 includes peripheral or bus logic including a USB controller 1625, a UART controller 1630, an SPI / SDIO controller 1635, and an I2S / I2C controller 1640. In at least one embodiment, integrated circuit 1600 can include a display device 1645 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1650 and a mobile industry processor interface (MIPI) display interface 1655. In at least one embodiment, storage can be provided by a flash memory subsystem 1660 including flash memory and a flash memory controller. Memory interface can be provided via a memory controller 1665 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1670.

[0319] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, IB, 2, 3, 4, 5, and 6. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, IB, 2, 3, 4, 5, and 6. In at least one embodiment, inference and / or training logic 115 can be used in integrated circuit 1600 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0320] Figures 17A-17B Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrations, other logic and circuitry can be included in the at least one embodiment, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.

[0321] Figures 17A-17Bis a block diagram illustrating an exemplary graphics processor used within a SoC, in accordance with the embodiments described herein. Figure 17A An exemplary graphics processor 1710 of a system on a chip integrated circuit is shown, which can be fabricated using one or more IP cores, in accordance with at least one embodiment. Figure 17B Another exemplary graphics processor 1740 of a system on a chip integrated circuit is shown, which can be fabricated using one or more IP cores, in accordance with at least one embodiment. In at least one embodiment, Figure 17A The graphics processor 1710 of FIG. 17A is a low power graphics processor core. In at least one embodiment, Figure 17B The graphics processor 1740 of FIG. 17B is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1710, 1740 can be a variant of the graphics processor 1610 of FIG. 16. Figure 16 Variants of the graphics processor 1610 of FIG. 16.

[0322] In at least one embodiment, the graphics processor 1710 includes a vertex processor 1705 and one or more fragment processors 1715A-1715N (e.g., 1715A, 1715B, 1715C, 1715D through 1715N-1, and 1715N). In at least one embodiment, the graphics processor 1710 can execute different shader programs via separate logical

[0323] In at least one embodiment, the graphics processor 1710 additionally includes one or more memory management units (MMUs) 1720A-1720B, one or more caches 1725A-1725B, and one or more circuit interconnects 1730A-1730B. In at least one embodiment, one or more MMUs 1720A-1720B provide virtual-to-physical address mappings for the graphics processor 1710, including for vertex processors 1705 and / or fragment processors 1715A-1715N, which can reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 1725A-1725B. In at least one embodiment, one or more MMUs 1720A-1720B can be synchronized with other MMUs within the system, including with... Figure 16 One or more application processors 1605, graphics processors 1615, and / or video processors 1620 are associated with one or more MMUs, enabling each processor 1605-1620 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1730A-1730B enable the graphics processor 1710 to connect to other IP cores within the SoC via the SoC's internal bus or via a direct connection.

[0324] In at least one embodiment, the graphics processor 1740 includes one or more shader cores 1755A-1755N (e.g., 1755A, 1755B, 1755C, 1755D, 1755E, 1755F to 1755N-1 and 1755N), such as Figure 17B As shown, it provides a unified shader core architecture, where a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1740 includes an inter-core task manager 1745, which acts as a thread dispatcher to assign execution threads to one or more shader cores 1755A-1755N and a tile unit 1758 to accelerate tile-based rendering operations, where scene rendering operations are subdivided in image space, for example, to take advantage of local spatial consistency within the scene or optimize the use of internal caches.

[0325] Inference and / or training logic 115 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 1A and / or Figure 1BDetails regarding the inference and / or training logic 115 are provided. In at least one embodiment, the inference and / or training logic 115 may be integrated into an integrated circuit. Figure 17A and / or Figure 17B The above is used for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions or architectures, or neural network use cases described herein.

[0326] Figures 18A-18B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. In at least one embodiment, Figure 18A It shows that it can be included in Figure 16 The graphics core 1800 within the graphics processor 1610, and in at least one embodiment, may be as follows: Figure 17B The unified shader cores shown are 1755A-1755N. Figure 18B A highly parallel general-purpose graphics processing unit (“GPGPU”) 1830 suitable for deployment on a multi-chip module is shown in at least one embodiment.

[0327] In at least one embodiment, the graphics core 1800 includes a shared instruction cache 1802, texture units 1818, and cache / shared memory 1820, which are common to the execution resources within the graphics core 1800. In at least one embodiment, the graphics core 1800 may include multiple slices 1801A-1801N or partitions of each core, and the graphics processor may include multiple instances of the graphics core 1800. In at least one embodiment, slices 1801A-1801N may include supporting logic, including local instruction caches 1804A-1804N, thread schedulers 1806A-1806N, thread dispatchers 1808A-1808N, and a set of registers 1810A-1810N. In at least one embodiment, slices 1801A-1801N may include a set of additional functional units (AFU 1812A-1812N), floating-point units (FPU 1814A-1814N), integer arithmetic logic units (ALU 1816A-1816N), address calculation units (ACU 1813A-1813N), double-precision floating-point units (DPFPU 1815A-1815N), and matrix processing units (MPU 1817A-1817N).

[0328] In at least one embodiment, FPUs 1814A-1814N can perform single-precision (32-bit) and half-precision (16-bit) floating point operations, while DPFPUs 1815A-1815N perform double-precision (64-bit) floating point operations. In at least one embodiment, ALUs 1816A-1816N can perform variable precision integer operations at 8-bit, 16-bit, and 32-bit precision, and can be configured to operate in mixed precision. In at least one embodiment, MPUs 1817A-1817N can also be configured for mixed precision matrix operations, including half-precision floating point and 8-bit integer operations. In at least one embodiment, MPUs 1817-1817N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated General Matrix to Matrix Multiplication (GEMM).

[0329] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A-1C, 2, 3, 6, 7, and 8. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A-1C, 2, 3, 6, 7, and 8. In at least one embodiment, inference and / or training logic 115 can be used in graphics processing unit 1800 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0330] Figure 18BA general-purpose processing unit (GPGPU) 1830 is illustrated in at least one embodiment, which can be configured to enable highly parallel computational operations to be performed by a set of graphics processing units. In at least one embodiment, the GPGPU 1830 can be directly linked to other instances of the GPGPU 1830 to create a multi-GPU cluster to improve the training speed for deep neural networks. In at least one embodiment, the GPGPU 1830 includes a host interface 1832 for connection to a host processor. In at least one embodiment, the host interface 1832 is a PCI Express interface. In at least one embodiment, the host interface 1832 may be a vendor-specific communication interface or communication structure. In at least one embodiment, the GPGPU 1830 receives commands from the host processor and uses a global scheduler 1834 to allocate execution threads associated with those commands to a set of compute clusters 1836A-1836H. In at least one embodiment, compute clusters 1836A-1836H share a cache memory 1838. In at least one embodiment, cache memory 1838 can be used as a higher-level cache within the cache memory of computing clusters 1836A-1836H.

[0331] In at least one embodiment, the GPGPU 1830 includes memories 1844A-1844B, which are coupled to the computing cluster 1836A-1836H via a set of memory controllers 1842A-1842B. In at least one embodiment, memories 1844A-1844B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), which includes graphics double data rate (GDDR) memory.

[0332] In at least one embodiment, each of the computing clusters 1836A-1836H includes a set of graphics cores, for example... Figure 18A The graphics core 1800 may include various types of integer and floating-point logic units that can perform computational operations across a range of precisions, including precisions suitable for machine learning computations. For example, in at least one embodiment, at least a subset of the floating-point units in each computing cluster 1836A-1836H may be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units may be configured to perform 64-bit floating-point operations.

[0333] In at least one embodiment, multiple instances of GPGPU 1830 can be configured to function as a compute cluster. In at least one embodiment, communication for synchronization and data exchange for compute clusters 1836A-1836H varies between embodiments. In at least one embodiment, multiple instances of GPGPU 1830 communicate via host interface 1832. In at least one embodiment, GPGPU 1830 includes an I / O hub 1839 that couples the GPGPU 1830 with a GPU link 1840 that enables a direct connection to other instances of GPGPU 1830. In at least one embodiment, GPU link 1840 is coupled to a specialized GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGP 1830. In at least one embodiment, GPU link 1840 is coupled with a high-speed interconnect to transmit and receive data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1830 are located in separate data processing systems and communicate via a network device accessible via host interface 1832. In at least one embodiment, GPU link 1840 can be configured to enable connection to a host processor in addition to or as an alternative to host interface 1832.

[0334] In at least one embodiment, GPGPU 1830 can be configured to train neural networks. In at least one embodiment, GPGPU 1830 can be used within an inferencing platform. In at least one embodiment, where GPGPU 1830 is used for inferencing, GPGPU 1830 can include fewer compute clusters 1836A-1836H relative to when GPGPU 1830 is used to train neural networks. In at least one embodiment, memory technology associated with memory 1844A-1844B can vary between inferencing and training configurations, with higher bandwidth memory technology dedicated to training configurations. In at least one embodiment, an inferencing configuration of GPGPU 1830 can support inferencing specific instructions. For example, in at least one embodiment, an inferencing configuration can provide support for one or more 8-bit integer dot product instructions that can be used during inferencing operations for deployed neural networks.

[0335] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, inference and / or training logic 115 include at least one of hardware logic elements. Figure 1A and / or Figure 1BDetails regarding the inference and / or training logic 115 are provided. In at least one embodiment, the inference and / or training logic 115 can be used in a GPGPU 1830 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0336] Figure 19 A block diagram of a computer system 1900 is shown, in accordance with at least one embodiment. In at least one embodiment, computer system 1900 includes a processing subsystem 1901, with one or more processor(s) 1902, and a system memory 1904, communicating via an interconnection path 1905, which can include a memory hub 1905. In at least one embodiment, the memory hub 1905 can be a separate component, or it can be integrated into one or more of the processor(s) 1902. In at least one embodiment, memory hub 1905 couples with processor(s) 1902 through communication links 1906. In at least one embodiment, processor(s) 1902 can include one or more of processing cores, which can be implemented as SHPs, as discussed above.

[0337] In at least one embodiment, processing subsystem 1901 includes one or more parallel processor(s) 1912, which can communicate with processor(s) 1902 stemming from the same root processing die. In at least one embodiment, communication can be made over the HyperTransport (HT) link 1913, which can be based on the InfiniBand (IB) protocol. One or more parallel processor(s) 1912 can be configured to receive commands from the processor(s) 1902 and communicate with a memory 1914, which can be shared between processor(s) 1902 and one or more parallel processor(s) 1912 under the control of memory hub 1905.

[0338] In at least one embodiment, system storage 1914 can connect to I / O hub 1907 to provide storage mechanisms for computing system 1900. In at least one embodiment, I / O switches 1916 can be used to provide an interface mechanism to enable connections between I / O hub 1907 and other components, such as network adapter 1918 and / or wireless network adapter 1919 that can be integrated into a platform, as well as various other devices that can be added via one or more add-in devices 1920. In at least one embodiment, network adapter 1918 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, wireless network adapter 1919 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more radio devices.

[0339] In at least one embodiment, computing system 1900 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which can also be connected to I / O hub 1907. In at least one embodiment, communication paths between various components shown in FIG. 19 can be implemented using any suitable protocols, such as PCI-based protocols (e.g., PCI-Express), or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed interconnect, or interconnect protocols. Figure 19

[0340] In at least one embodiment, parallel processor 1912 includes circuitry such as, for example, video circuitry, constituting a graphics processing unit (GPU). In at least one embodiment, parallel processor 1912 includes circuitry optimized for general use such as, for example, fixed function and programmable execution units. In at least one embodiment, one or more components of computer system 1900 can be integrated on one or more other components. In at least one embodiment, for example and without limitation, integrated circuitic of parallel processor 1912, memory hub 1905, processor 1902, and I / O hub 1907 can be integrated into a system on a chip (SoC) integrated circuit. In at least one embodiment, for example and without limitation, components of computer system 1900 can be integrated into a single package to form a system in a package (SIP) configuration. In at least one embodiment, for example and without limitation, components of computer system 1900 can be integrated into a multi-chip module (MCM), which can be interconnected with other modules to form a modular computing system.

[0341] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Inferences and / or training can be performed on a set of one or more processors, such as one or more of processors 112. In at least one embodiment, one or more other components of system 110 can perform inferencing and / or training, in conjunction with or independent of processor 112.​ Figure 1A and / or Figure 1B Details regarding inferencing and / or training logic 115 are provided. In at least one embodiment, inferencing and / or training logic 115 can be used in computing system 1900 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein. Figure 19

[0342] Processor

[0343] Figure 20A A parallel processor 2000, according to at least one embodiment, is shown. In at least one embodiment, various components of parallel processor 2000 can be implemented using one or more integrated circuits, which can be programmable integrated circuits, application-specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, parallel processor 2000 is a variant of one or more parallel processors 1912 shown and described herein. Figure 19 Variants of one or more parallel processors 1912 shown.

[0344] In at least one embodiment, parallel processor 2000 includes a parallel processing unit 2002. In at least one embodiment, parallel processing unit 2002 includes an I / O unit 2004 that enables communication with other devices, including other instances of parallel processing unit 2002. In at least one embodiment, I / O unit 2004 can be directly connected to other devices. In at least one embodiment, I / O unit 2004 connects with other devices using a hub or switch interface, such as memory hub 2105. In at least one embodiment, connections between memory hub 2005 and I / O unit 2004 form a communication link 2013. In at least one embodiment, I / O unit 2004 connects with a host interface 2006 and a memory crossbar 2016, where host interface 2006 receives commands directed to processing operations and memory crossbar 2016 receives commands directed to memory operations.

[0345] ​In at least one embodiment, when host interface 2006 receives a command buffer via I / O unit 2004, host interface 2006 can direct a work operation to execute those commands to front end 2008. In at least one embodiment, front end 2008 is coupled with scheduler 2010, which is configured to assign commands or other work items to processing cluster array 2012. In at least one embodiment, scheduler 2010 ensures that processing cluster array 2012 is properly configured and in an active state before assigning tasks to processing cluster array 2012. In at least one embodiment, scheduler 2010 is implemented by firmware logic executing on a microcontroller. In at least one embodiment, microcontroller- implemented scheduler 2010 is configurable to perform complex scheduling and work distribution operations with both coarse and fine grain, enabling fast preemption and context switching of threads executing on processing array 2012. In at least one embodiment, host software can prove a workload for scheduling on processing array 2012 through one of multiple graphics processing paths. In at least one embodiment, workload can then be automatically distributed by scheduler 2010 logic within a microcontroller including scheduler 2010 on processing array 2012.

[0346] In at least one embodiment, processing cluster array 2012 can include up to “N” processing clusters (e.g., cluster 2014A, cluster 2014B, through cluster 2014N), where “N” represents a positive integer (which can be a different integer than integer “N” used in other Figures). In at least one embodiment, each cluster 2014A-2014N of processing cluster array 2012 can execute a large number of concurrent threads. In at least one embodiment, scheduler 2010 can use various scheduling and / or work distribution algorithms to assign work to clusters 2014A-2014N of processing cluster array 2012, which can vary depending on workload produced by each program or type of computation. In at least one embodiment, scheduling can be handled dynamically by scheduler 2010, or can be assisted in part by compiler logic during compilation of program logic configured for execution by processing cluster array 2012. In at least one embodiment, different clusters 2014A-2014N of processing cluster array 2012 can be allocated for processing different types of programs or for performing different types of computations.

[0347] In at least one embodiment, processing cluster array 2012 can be configured to perform a variety of types of parallel processing operations. In at least one embodiment, processing cluster array 2012 is configured to perform general-purpose parallel compute operations. For example, in at least one embodiment, processing cluster array 2012 can include logic to perform processing tasks including filtering of video and / or audio data, performing modeling operations, including physics operations, and performing data transformations.

[0348] In at least one embodiment, processing cluster array 2012 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 2012 can include additional logic to support the performance of such graphics processing operations including, but not limited to, texture sampling logic to perform texture operations for three-dimensional (3D) graphics, tessellation logic, and other vertex processing logic. In at least one embodiment, processing cluster array 2012 can be configured to execute a graphics-processing-related shader program, such as, for example and without limitation, a Vertex Shader, a Geometry Shader, and / or a Pixel Shader. In at least one embodiment, parallel processing unit 2002 can transfer data to be processed by processing cluster array 2012 from memory, such as system memory, via I / O unit 2004. In at least one embodiment, during processing, results can be written to on-chip memory (e.g., parallel processor memory 2022) for processing on processing cluster array 2012 before eventually being written to system memory.

[0349] In at least one embodiment, when parallel processing unit 2002 is used to perform graphics processing, scheduler 2010 can be configured to divide the processing workload into approximately equal sized tasks, to better enable distribution of the graphics processing operations across set of clusters 2014A-2014N. In at least one embodiment, portions of processing cluster array 2012 can be configured to perform different types of processing. For example, in at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations, to produce a rendered image for display. In at least one embodiment, intermediate data produced by one or more of clusters 2014A-2014N can be stored in buffers to allow transmission of the intermediate data between clusters 2014A-2014N for further processing.

[0350] In at least one embodiment, processing cluster array 2012 can receive processing tasks to be executed via a scheduler 2010 that receives commands defining processing tasks from front end 2008. In at least one embodiment, processing tasks can include indices of data to be processed, e.g., surface (patch) data, raw data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., what program is to be executed). In at least one embodiment, scheduler 2010 can be configured to fetch indices corresponding to a task, or can receive indices from front end 2008. In at least one embodiment, front end 2008 can be configured to ensure that processing cluster array 2012 is configured in an effective state before launching a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).

[0351] In at least one embodiment, each of one or more instances of parallel processing unit 2002 can be coupled to parallel processor memory 2022 via memory crossbar 2016. In at least one embodiment, memory crossbar 2016 can receive memory requests from processing cluster array 2012, as well as I / O unit 2004. In at least one embodiment, memory crossbar 2016 can be configured to have a wide memory bus to provide high bandwidth required for memory operations on vertex data. In at least one embodiment, memory crossbar 2016 can be configured to include one or more memory crossbars.

[0352] In at least one embodiment, memory units 2024A-2024N can include various types of memory devices including dynamic random access memory (DRAM) or graphics random access memory, e.g., synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 2024A-2024N can also include 3D stacked memory including, but not limited to, high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps can be stored across memory units 2024A-2024N, allowing partition units 2020A-2020N to write portions of each rendering target in parallel to efficiently use available bandwidth of parallel processor memory 2022. In at least one embodiment, local instances of parallel processor memory 2022 can be excluded from being utilized in favor of a unified memory design that utilizes system memory in combination with local cache memory.

[0353] In at least one embodiment, any of clusters 2014A-2014N of processing cluster array 2012 can process data that is to be written to any of memory units 2024A-2024N within parallel processor memory 2022. In at least one embodiment, memory crossbar 2016 can be configured to transmit outputs of each cluster 2014A-2014N to any partition unit 2020A-2020N or another cluster 2014A-2014N, which can perform additional processing operations on the outputs. In at least one embodiment, each cluster 2014A-2014N can communicate with memory interface 2018 through memory crossbar 2016 to read from or write to various external memory devices. In at least one embodiment, memory crossbar 2016 has a connection to memory interface 2018 to communicate with I / O unit 2004, and a connection to a local instance of parallel processor memory 2022, enabling processing elements within different processing clusters 2014A-2014N to communicate with system memory or other memories not local to parallel processor 2002. In at least one embodiment, memory crossbar 2016 can separate transactions according to software addresses or according to hardware addresses.

[0354] In at least one embodiment, multiple instances of parallel processing unit 2002 can be provided on a single add-in card, or multiple add-in cards can be interconnected. In at least one embodiment, different instances of parallel processing unit 2002 can be configured to operate in coordination with one another (e.g., in a clustered computing environment). In at least one embodiment, parallel processing unit 2002 of different instances can be configured to operate in a lockstep fashion to provide increased processing capability for certain workloads. For example, in at least one embodiment, some instances of parallel processing unit 2002 can include a higher precision floating point unit relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 2002 or parallel processor 2000 can be implemented in a variety of form factors, including without limitation a desktop, laptop, or hand-held personal computer, server, work station, game console, and / or embedded systems.

[0355] Figure 20B is a block diagram of a partition unit 2020 in accordance with at least one embodiment. In at least one embodiment, partition unit 2020 is an instance of one of partition units 2020A-2020N of Figure 20A In at least one embodiment, partition unit 2020 includes an L2 cache 2021, a frame buffer interface 2025, and a ROP 2026 (raster operations unit). In at least one embodiment, L2 cache 2021 is a read / write cache that is configured to perform load and store operations received from memory crossbar 2016 and ROP 2026. In at least one embodiment, L2 cache 2011 outputs read misses and urgent write-back requests to frame buffer interface 2025 for processing. In at least one embodiment, updates can also be sent to a frame buffer via frame buffer interface 2025 for processing. In at least one embodiment, frame buffer interface 2025 interacts with a memory unit of parallel processor memory 2022, such as memory units 2024A-2024N (e.g., within parallel processor memory 2022), to store or retrieve data. Figure 20A In at least one embodiment, frame buffer interface 2025 is configured to send data to an associated frame buffer, which can be configured to process pixel data. In at least one embodiment, frame buffer interface 2025 can be configured to transmit pixel data for display on an electronic display device.

[0356] In at least one embodiment, ROP 2026 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, ROP 2026 then outputs processed graphics data into graphics memory. In at least one embodiment, ROP 2026 includes compression logic to compress depth or color data being written into memory and decompress depth or color data being read from memory. In at least one embodiment, compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. In at least one embodiment, a type of compression performed by ROP 2026 can vary based on statistical properties of data to be compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis on depth and color data.

[0357] In at least one embodiment, ROP 2026 is included within each processing cluster (e.g., processing clusters 2014A-2014N) rather than in a segment unit 2020. In at least one embodiment, read and write requests to pixel data are transmitted by memory crossbar 2016 rather than pixel fragment data. In at least one embodiment, processed graphics data can be displayed on one of one or more display devices 1910, routed by processor 1302 for further processing, or routed by one of processing entities within parallel processor 2000 for further processing. Figure 20A In at least one embodiment, ROP 2026 is included within each processing cluster (e.g., processing clusters 2014A-2014N) rather than in a segment unit 2020. In at least one embodiment, read and write requests to pixel data are transmitted by memory crossbar 2016 rather than pixel fragment data. In at least one embodiment, processed graphics data can be displayed on one of one or more display devices 1910, routed by processor 1302 for further processing, or routed by one of processing entities within parallel processor 2000 for further processing. Figure 19 In at least one embodiment, ROP 2026 is included within each processing cluster (e.g., processing clusters 2014A-2014N) rather than in a segment unit 2020. In at least one embodiment, read and write requests to pixel data are transmitted by memory crossbar 2016 rather than pixel fragment data. In at least one embodiment, processed graphics data can be displayed on one of one or more display devices 1910, routed by processor 1302 for further processing, or routed by one of processing entities within parallel processor 2000 for further processing. Figure 20A In at least one embodiment, ROP 2026 is included within each processing cluster (e.g., processing clusters 2014A-2014N) rather than in a segment unit 2020. In at least one embodiment, read and write requests to pixel data are transmitted by memory crossbar 2016 rather than pixel fragment data. In at least one embodiment, processed graphics data can be displayed on one of one or more display devices 1910, routed by processor 1302 for further processing, or routed by one of processing entities within parallel processor 2000 for further processing.

[0358] Figure 20C is a block diagram of a processing cluster 2014 within a parallel processing unit in accordance with at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of processing clusters 2014A-2014N. In at least one embodiment, processing cluster 2014 can be configured to execute a large number of threads concurrently, where a thread Figure 20A is a block diagram of a processing cluster 2014 within a parallel processing unit in accordance with at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of processing clusters 2014A-2014N. In at least one embodiment, processing cluster 2014 can be configured to execute a large number of threads concurrently, where a thread

[0359] In at least one embodiment, operation of processing cluster 2014 can be controlled via a pipeline manager 2032 that assigns processing tasks to SIMT parallel processor. In at least one embodiment, pipeline manager 2032 receives instructions from Figure 20AThe scheduler 2010 receives instructions, and manages execution of those instructions by the graphics multiprocessor 2034 and / or the texture unit 2036. In at least one embodiment, the graphics multiprocessor 2034 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of differing architectures can be included within the processing cluster 2014. In at least one embodiment, one or more instances of the graphics multiprocessor 2034 can be included within a processing cluster 2014. In at least one embodiment, a graphics multiprocessor 2034 can process data, and a data crossbar 2040 can be used to distribute the processed data to one of multiple possible destinations, including other shader units. In at least one embodiment, a pipeline manager 2032 can facilitate distribution by, for example, specifying destinations of processed data.

[0360] In at least one embodiment, each graphics multiprocessor 2034 within the processing cluster 2014 can include an identical set of functional execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, functional execution logic can be configured in a pipelined manner, where new instructions can be issued before previous instructions are complete. In at least one embodiment, functional execution logic supports a variety of operations including integer and floating point arithmetic, comparison operations, Boolean operations, shift operations, and the like. In at least one embodiment, same functional -unit hardware can be leveraged to perform different operations in response to different instruction sets being issued to the functional -unit hardware. Any combination of

[0361] In at least one embodiment, instructions transmitted to the processing cluster 2014 form a thread for execution. In at least one embodiment, a set of threads executing across a group of parallel processing engines forms a warp. In at least one embodiment, a thread group is scheduled to execute on the processing cluster 2014. In at least one embodiment, each thread within a thread group can be assigned to a different processing engine within the graphics multiprocessor 2034. In at least one embodiment, a thread group can include fewer threads than are available processing engines within the graphics multiprocessor 2034. In at least one embodiment, when a thread group includes fewer threads than the number of processing engines within the graphics multiprocessor 2034, one or more of the processing engines can be idle during the execution of the warp. In at least one embodiment, a thread group can also include more threads than are available processing engines within the graphics multiprocessor 2034. In at least one embodiment, when a thread group includes more threads than the number of processing engines within the graphics multiprocessor 2034, multiple threads can be executed per each clock cycle by virtue of having multiple threads assigned to each processing engine. In at least one embodiment, a plurality of thread groups can be in process on the graphics multiprocessor 2034.

[0362] In at least one embodiment, the graphics multiprocessor 2034 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multiprocessor 2034 may forgo the internal cache and use a cache memory within the processing cluster 2014 (e.g., L1 cache 2048). In at least one embodiment, each graphics multiprocessor 2034 may also access partition units (e.g., Figure 20A The L2 cache is located within partition units 2020A-2020N, ​​which are shared among all processing clusters 2014 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 2034 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory outside of the parallel processing unit 2002 can be used as global memory. In at least one embodiment, the processing cluster 2014 includes multiple instances of the graphics multiprocessor 2034, which can share common instructions and data that can be stored in the L1 cache 2048.

[0363] In at least one embodiment, each processing cluster 2014 may include a memory management unit (“MMU”) 2045 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2045 may reside in Figure 20A The memory interface 2018 is located within the MMU 2045. In at least one embodiment, the MMU 2045 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles and optionally to cache line indices. In at least one embodiment, the MMU 2045 may include an address translation lookup buffer (TLB) or a cache that may reside within the graphics multiprocessor 2034, the L1 cache 2048, or the processing cluster 2014. In at least one embodiment, physical addresses are processed to allocate surface data access locality for efficient request interleaving between partition units. In at least one embodiment, cache line indices may be used to determine whether a request for a cache line is a hit or a miss.

[0364] In at least one embodiment, processing cluster 2014 can be configured such that each graphics multiprocessor 2034 is coupled to a texture unit 2036 for performing texture mapping operations in accordance with texture coordinate values. In at least one embodiment, texture data is read from an internal texture Ll cache (not shown) or from an L2 cache (not shown) as needed. In at least one embodiment, texture data is also fetched from a graphics processor memory, such as a shared memory 2070, an L2 cache, or a system memory, as needed. In at least one embodiment, each graphics multiprocessor 2034 outputs processed tasks to a data crossbar 2040 in processing cluster 2014. In at least one embodiment, data crossbar 2040 performs a write operation for the task results to shared memory 2070 via a memory crossbar 2016. In at least one embodiment, data crossbar 2040 performs read for task data from shared memory 2070 via memory crossbar 2016. In at least one embodiment, data crossbar 2040 also performs a read operation for task data from a shared memory 2070 via memory crossbar 2016 and performs a write operation for a task result to shared memory 2070 via memory crossbar 2016. In at least one embodiment, data crossbar 2040 is configured to perform the read and write operations as specific instructions of a vertex or geometry processing pipeline. Figure 20A In at least one embodiment, preROP 2042 is configured to receive data from graphics multiprocessors 2034, direct the data to a ROP unit which can be located within preROP 2042 or which can be part of the main graphics processor. In at least one embodiment, preROP 2042 can perform optimizations and

[0365] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 1 and 12. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 1 and 12. In at least one embodiment, inference and / or training logic 115 can be used in graphics processing cluster 2014 for inferencing or predicting operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0366] Figure 20D A graphics multiprocessor 2034 according to at least one embodiment is shown. In at least one embodiment, graphics multiprocessor 2034 is coupled to a pipeline manager 2032 of processing cluster 2014. In at least one embodiment, graphics multiprocessor 2034 has a graphics processing pipeline that includes, in at least one embodiment, without limitation, an instruction cache 2052, an instruction unit 2054, an address mapping unit 2056, a register file 2058, one or more general-purpose graphics processing unit(s) (GPGPU) core(s) 2062, and one or more load / store units 2066. In at least one embodiment, GPGPU core(s) 2062 and load / store units 2066 are coupled by a memory and cache interconnect 2068 with cache memory 2072 and shared memory 2070.

[0367] In at least one embodiment, instruction cache 2052 receives a stream of instructions 2050 to be executed by graphics processing engine 2030 from pipeline manager 2032. In at least one embodiment, instructions 2050 are cached in instruction cache 2052 and dispatched for execution by instruction unit 2054. In at least one embodiment, instruction unit 2054 can dispatch instructions to threads assigned to different execution units within GPGPU cores 2062 as thread groups, for example, thread warps. In at least one embodiment, instructions can access any of local, shared, or global address spaces using addresses translated by address mapping unit 2056. In at least one embodiment, address mapping unit 2056 can be used to access the different memory address spaces by the different execution units.

[0368] In at least one embodiment, register file 2058 provides a set of registers for functional units of graphics processing engine 2034. In at least one embodiment, register file 2058 provides temporary storage for operands of the data

[0369] In at least one embodiment, GPGPU cores 2062 can each include floating point

[0370] In at least one embodiment, GPGPU cores 2062 include SIMD logic capable of performing a single instruction on multiple sets of data. In at least one embodiment, GPGPU cores 2062 can physically execute SIMD 4, SIMD 8, and SIMD 16 instructions and logically execute a SIMD 1, SIMD 2, and SIMD 32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated by a shader compiler during compilation of code written by a programmer. In at least one embodiment, a programmer writing code for the programmable processing unit 2000 can write SIMD code in a high level programming language, which is then compiled into multiple instruction packets that can include one or more SIMD instructions and / or one or more SIMD control instructions. In at least one embodiment, a single

[0371] In at least one embodiment, memory and cache interconnect 2068 is an interconnect network that connects each functional unit of graphics multiprocessor 2034 to register file 2058 and shared memory 2070. In at least one embodiment, memory and cache interconnect 2068 is a crossbar interconnect that allows load / store units 2066 to effect load and store operations between shared memory 2070 and register file 2058. In at least one embodiment, register file 2058 can operate at the same frequency as GPGPU cores 2062, making the latency in transferring data between GPGPU cores 2062 and register file 2058 very low. In at least one embodiment, shared memory 2070 can be used for communications between threads executing on functional units within graphics multiprocessor 2034. In at least one embodiment, cache memory 2072 can be used to cache data, such as texture data, communicated between the functional units and texture unit 2036. In at least one embodiment, shared memory 2070 can also be used for program managed caching. In at least one embodiment, in addition to automatically cached data stored in cache memory 2072, a thread executing on GPGPU cores 2062 can store data in shared memory in a programmatic manner.

[0372] In at least one embodiment, parallel processor(s) or GPGPUs as described herein are communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, GPU can be communicatively coupled to host processor / cores by a bus or other interconnect (e.g., a high speed

[0373] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 7, 8A, 8B, 9, and / or 10. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1 A, 7, 8A, 8B, 9, and / or 10.

[0374] Figure 21A multi-GPU computing system 2100 is shown in accordance with at least one embodiment. In at least one embodiment, multi-GPU computing system 2100 can include a processor 2102 coupled to a plurality of general purpose graphics processing units (GPGPUs) 2106A-D via a host interface switch 2104. In at least one embodiment, host interface switch 2104 is a PCI Express switch device that couples processor 2102 to a PCI Express bus over which processor 2102 can communicate with GPGPUs 2106A-D. In at least one embodiment, GPGPUs 2106A-D can be interconnected via a set of high-speed P2P GPU-to-GPU links 2116. In at least one embodiment, GPU-to-GPU links 2116 connect to each of GPGPUs 2106A-D via a dedicated GPU link. In at least one embodiment, P2P GPU links 2116 enable direct communication between each GPGPU 2106A-D without having to communicate through host interface bus 2104 to which processor 2102 is connected. In at least one embodiment, where GPU-to-GPU traffic is directed to P2P GPU links 2116, host interface bus 2104 remains available for system memory access or communication with other instances of multi-GPU computing system 2100, e.g., via one or more network devices. While in at least one embodiment GPGPUs 2106A-D are connected to processor 2102 via host interface switch 2104, in at least one embodiment processor 2102 includes direct support for P2P GPU links 2116 and can be directly connected to GPGPUs 2106A-D.

[0375] Inference and / or training logic 115 are used to perform inferencing and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1A-1C, 2, 3, 5, 6, 7, 8, and 9. In embodiments in which inference and / or training logic 115 is used for inferencing, without training, inference and / or training logic 115 can be considered to be part of processing circuitry 120. Figure 1A and / or Figure 1B Details regarding inference and / or training logic 115 are provided below in conjunction with FIGS. 1A-1C, 2, 3, 5, 6, 7, 8, and 9. In embodiments in which inference and / or training logic 115 is used for inferencing, without training, inference and / or training logic 115 can be considered to be part of processing circuitry 120.

[0376] Figure 22is a block diagram of a graphics processor 2200 according to at least one embodiment. In at least one embodiment, graphics processor 2200 includes an ring interconnect 2202, a front-end pipeline 2204, a media engine 2237, and graphics cores 2280A-2280N. In at least one embodiment, ring interconnect 2202 couples graphics processor 2200 to other processing units including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 2200 is one of a number of processors integrated within a multi-core processing system.

[0377] In at least one embodiment, graphics processor 2200 receives batches of commands via ring interconnect 2202. In at least one embodiment, incoming commands are interpreted by a command streamer 2203 in pipeline front-end 2204. In at least one embodiment, graphics processor 2200 includes scalable execution logic to perform 3D geometry processing and media processing via the graphics cores 2280A-2280N. In at least one embodiment, for 3D geometry processing commands, command streamer 2203 supplies commands to geometry pipe...

Claims

1. A processor comprising: one or more circuits to detect cheating by one or more users of a computer game using one or more neural networks based at least in part on one or more images generated by the computer game, wherein training of the one or more neural networks is based at least in part on one or more cheating images and one or more non-cheating images, at least a subset of the one or more cheating images comprising images augmented with information from non-cheating images.

2. The processor of claim 1, wherein the one or more circuits are further to: store the one or more images in a buffer, wherein the stored one or more images are to be provided as input to the one or more neural networks and to be rendered on a display unit.

3. The processor of claim 2, wherein the one or more circuits are further to: generate a report indicating that cheating was detected; and communicate the report to a game server.

4. The processor of claim 3, wherein the one or more circuits are further to: generate, using the one or more neural networks, a confidence level characterizing a confidence that the one or more images include cheating information.

5. The processor of claim 4, wherein the report is communicated to the game server in response to determining that the confidence level is at or above a threshold value.

6. The processor of claim 3, wherein the one or more circuits are further to: receive, from the game server, updated parameters of the one or more neural networks, wherein the updated parameters are generated based on a set of retraining images.

7. The processor of claim 1, wherein the one or more circuits are to generate an authentication signal to a game server, wherein the authentication signal is to prove to the game server that the one or more circuits are capable of detecting cheating associated with the computer game.

8. A processor comprising: one or more circuits to perform training of one or more neural networks to detect cheating by one or more users of a computer game, wherein the training is based at least in part on one or more cheating images and one or more non-cheating images, at least a subset of the one or more cheating images comprising images augmented with information from non-cheating images.

9. The processor of claim 8, wherein the one or more cheating images are generated using cheating software associated with the computer game.

10. The processor of claim 8, wherein each of the one or more cheating images includes cheating information and each of the one or more non-cheating images is free of cheating information.

11. The processor of claim 8, wherein the non-cheating images are generated by game software associated with the computer game.

12. The processor of claim 11, wherein each of the subset of one or more cheat images includes a portion replaced with a portion of a non-cheat image.

13. The processor of claim 10, wherein at least a subset of the one or more non-cheat images includes a non-cheat image augmented with information from a cheat image generated by game software associated with the computer game.

14. The processor of claim 13, wherein each of the subset of one or more non-cheat images includes a portion replaced with a portion of a cheat image containing non-cheat information.

15. The processor of claim 8, wherein the one or more circuits are further to perform training of the one or more neural networks against adversarial attacks, the training of the one or more neural networks against adversarial attacks based at least in part on a subset of one or more cheat images modified with adversarial perturbations.

16. A system comprising: one or more processors to detect cheating by one or more users of a computer game based at least in part on one or more images generated by the computer game using one or more neural networks; and one or more memories to store parameters associated with the one or more neural networks; wherein training of the one or more neural networks is based at least in part on one or more cheat images and one or more non-cheat images, at least a subset of the one or more cheat images including images augmented with information from non-cheat images.

17. The system of claim 16, wherein the one or more processors are further to: store the one or more images in a buffer, wherein the stored one or more images are to be provided as input to the one or more neural networks and to be rendered on a display unit.

18. The system of claim 17, wherein the one or more processors are further to: generate a report indicating detection of cheating; and communicate the report to a game server.

19. The system of claim 18, wherein to communicate the report to the game server, the one or more processors are to: generate, using the one or more neural networks, a confidence level characterizing a confidence that the one or more images include cheating information; and determine that the confidence level is at or above a threshold.

20. A system comprising: one or more processors to perform training of one or more neural networks to detect cheating by one or more users of a computer game based at least in part on one or more cheat images generated by cheating software associated with the computer game; and one or more memories to store parameters associated with the one or more neural networks. wherein the training of the one or more neural networks is further based at least in part on one or more non-cheating images, and at least a subset of the one or more cheating images includes images augmented with information from non-cheating images.

21. The system of claim 20, wherein the one or more cheating images are generated using cheating software associated with the computer game, each of the one or more cheating images includes cheating information, and each of the one or more non-cheating images is free of cheating information.

22. The system of claim 21, wherein, The non-cheating images are generated by game software associated with the computer game.

23. The system of claim 22, wherein each of the subset of the one or more cheating images includes a portion replaced with a portion of a non-cheating image.

24. The system of claim 21, wherein at least a subset of the one or more non-cheating images includes a non-cheating image augmented with information from a cheating image generated by game software associated with the computer game.

25. The system of claim 24, wherein each of the subset of the one or more non-cheating images includes a portion replaced with a portion of the cheating image that includes non-cheating information.

26. The system of claim 20, wherein the one or more processors are further to perform training of the one or more neural networks against an adversary attack, the training of the one or more neural networks against an adversary attack is based at least in part on a subset of one or more cheating images modified with adversary perturbations.

27. A method comprising: receiving, from a computing device, a representation of graphics associated with a computer game; generating, by one or more circuits, one or more images based on the received representation; and processing, by the one or more circuits, the one or more images using one or more neural networks to detect cheating by one or more users of the computer game, wherein training of the one or more neural networks is based at least in part on one or more cheating images and one or more non-cheating images, at least a subset of the one or more cheating images includes images augmented with information from non-cheating images.

28. The method of claim 27, further comprising: generating a report indicating that cheating was detected; and communicating the report to a game server.

29. The method of claim 28, wherein communicating the report to a game server is in response to: generating, by the one or more circuits, using the one or more neural networks, a confidence level characterizing a confidence that the one or more images include cheating information; and determining that the confidence level is at or above a threshold.

30. The method of claim 29, further comprising: receiving updated parameters of the one or more neural networks, wherein the updated parameters are generated based on retraining images.

31. The method of claim 27, wherein the one or more neural networks are trained using: the one or more cheat images generated using cheat software associated with the computer game; and the one or more non-cheat images that are free of cheat information.

32. The method of claim 31, wherein the one or more neural networks are further trained against an adversarial attack, wherein the training against an adversarial attack is based at least in part on a subset of the one or more cheat images modified with adversarial perturbations.

Citation Information

Patent Citations

  • Games plug-in detection method and device thereof

    CN110102051A

  • Game tag-on service behavior determination method and device, electronic equipment and storage medium

    CN111803956A

  • Automatically reducing use of cheat software in online game environment

    CN111886059A