IDENTIFYING STORAGE FOR RAY TRACING

By estimating ray tracing workload before simulation and distributing rays across multiple processing units, the method addresses memory and computational limitations in ray tracing, ensuring the completion of complex simulations.

DE102025113545A1Pending Publication Date: 2025-10-09NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025113545
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-08
Filing Date
2025-04-07
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Ray tracing simulations face memory and computational resource limitations due to the exponential scaling of diffuse and diffraction interactions, particularly in high-resolution or large-scale applications, leading to incomplete simulations.

Method used

A method is introduced to estimate ray tracing workload prior to simulation, allowing for the distribution of rays across multiple processing units and memory spaces, enabling completion of simulations with a large number of rays and interactions.

Benefits of technology

This approach ensures the completion of ray tracing simulations with a significant number of rays and interactions, overcoming memory constraints by distributing the workload across multiple GPUs, thus enhancing computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Apparatus, systems, and methods for performing ray tracing are disclosed. In at least one embodiment, a ray tracing workload is estimated prior to a ray tracing simulation, for example, based on less than all of the data associated with the rays to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] At least one embodiment relates to processing resources used to perform and enable ray tracing. For example, at least one embodiment relates to processors or computer systems used to estimate the ray tracing workload according to various novel techniques described herein. background

[0002] Ray tracing can require significant memory, time, or computational resources. The amount of memory, time, or computational resources used to perform ray tracing can be improved. Short description of the drawings Fig. 1 shows an example of a system for performing ray tracing according to at least one embodiment; Fig. 2 shows an example of a method for performing ray tracing according to at least one embodiment; Fig. 3 shows an example of a method for performing ray tracing according to at least one embodiment; Fig. 4 illustrates an example of a processor with modules for performing ray tracing according to at least one embodiment; Fig. 5 illustrates an example block diagram illustrating a driver and / or runtime environment including one or more libraries to provide one or more application programming interfaces (APIs), according to at least one embodiment; Fig. 6 illustrates an example data center system according to at least one embodiment; Fig. 7A shows an example of an autonomous vehicle according to at least one embodiment; Fig. Figure 7B shows an example of camera positions and fields of view for the autonomous vehicle of Fig. 7A according to at least one embodiment; Fig. Figure 7C is a block diagram illustrating an example system architecture for the autonomous vehicle of Fig. 7A according to at least one embodiment; Fig. Figure 7D is a diagram illustrating a system for communication between one or more cloud-based servers and the autonomous vehicle of Fig. 7A according to at least one embodiment; Fig. 8 is a block diagram illustrating a computer system according to at least one embodiment; Fig. 9 is a block diagram illustrating a computer system according to at least one embodiment; Fig. 10 illustrates a computer system according to at least one embodiment; Fig. 11 illustrates a computer system according to at least one embodiment; Fig. 12A illustrates a computer system according to at least one embodiment; Fig. 12B illustrates a computer system according to at least one embodiment; Fig. 12C illustrates a computer system according to at least one embodiment; Fig. 12D illustrates a computer system according to at least one embodiment; Fig. 12E and Fig. 12F illustrate a common programming model according to at least one embodiment; Fig. 13 illustrates example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 14A and Fig. 14B illustrate example integrated circuits and associated graphics processors according to at least one embodiment; Fig. 15A and Fig. 15B illustrate additional example graphics processor logic according to at least one embodiment; Fig. 16 illustrates a computer system according to at least one embodiment; Fig. 17A illustrates a parallel processor according to at least one embodiment; Fig. 17B illustrates a partition unit according to at least one embodiment; Fig. 17C illustrates a processing cluster according to at least one embodiment; Fig. 17D illustrates a graphics multiprocessor according to at least one embodiment; Fig. 18 illustrates a multi-graphics processing unit (GPU) system according to at least one embodiment; Fig. 19 illustrates a graphics processor according to at least one embodiment; Fig. 20 is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment; Fig. 21 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 22 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 23 illustrates at least portions of a graphics processor according to one or more embodiments; Fig. 24 is a block diagram of a graphics processing engine of a graphics processor according to at least one embodiment; Fig. 25 is a block diagram of at least portions of a graphics processor core according to at least one embodiment; Fig. 26A and Fig. 26B illustrate thread execution logic comprising an array of processor elements of a graphics processor core, according to at least one embodiment; Fig. 27 illustrates a parallel processing unit ("PPU") according to at least one embodiment; Fig. 28 illustrates a general processing cluster (“GPC”) according to at least one embodiment; Fig. 29 illustrates a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment; Fig. 30 illustrates a streaming multiprocessor according to at least one embodiment; Fig. 31 illustrates a network for communicating data within a 5G wireless communication network according to at least one embodiment; Fig. 32 illustrates a network architecture for a 5G LTE wireless network according to at least one embodiment; Fig. 33 is a diagram illustrating some basic functions of a mobile telecommunications network / system operating according to LTE and 5G principles, according to at least one embodiment; Fig. 34 illustrates a radio access network that may be part of a 5G network architecture, according to at least one embodiment; Fig. 35 provides an exemplary illustration of a 5G mobile communications system using a variety of different types of devices, according to at least one embodiment; Fig. 36 illustrates a high-level example of a system according to at least one embodiment; Fig. 37 illustrates a system architecture of a network according to at least one embodiment; Fig. 38 illustrates exemplary components of a device according to at least one embodiment; Fig. 39 illustrates exemplary interfaces of baseband circuits according to at least one embodiment; Fig. 40 illustrates an example uplink channel according to at least one embodiment; Fig. 41 illustrates a system architecture of a network according to at least one embodiment; Fig. 42 illustrates a control plane protocol stack according to at least one embodiment; Fig. 43 illustrates a user-plane protocol stack according to at least one embodiment; Fig. 44 illustrates components of a core network according to at least one embodiment; Fig. 45 illustrates components of a system for supporting network function virtualization (NFV) according to at least one embodiment; and Fig. 46 illustrates components of a system for accessing a large language model according to at least one embodiment. Detailed description

[0003] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one skilled in the art that the inventive concepts may be practiced without one or more of these specific details.

[0004] In ray tracing applications where diffuse and diffraction interactions are relevant, such as modeling the propagation channel of radio frequency (RF) waves in wireless communications or modeling sound propagation, each ray that interacts with a diffuse surface or with a diffracting wall edge in a virtual scene results in the generation of multiple rays that are further traced in the virtual scene. For high-resolution targets or more accurate results, the number of secondary or generated rays can be very large, which can lead to incomplete ray tracing simulations due to insufficient memory or time budget. In at least one embodiment, generating a ray tracing workload prior to ray tracing can mitigate or eliminate the problems described above.To illustrate the above computational challenge, consider an urban-scale ray tracing application for RF wave propagation, where a typical map size is several square kilometers, with a deployment of 100 radio transmitters and 10,000 radio receivers with 64 antenna pairs per transmitter-receiver pair. Ray tracing in such an application requires deterministically constructing all possible paths between each of the transmitter-receiver antenna pairs with exact vertex positions at each interaction with a building surface and each interaction with a building edge to calculate phase-coherent and polarization-dependent electromagnetic fields, which may require uniform sampling for each interaction. Due to this, the ray tracing workload scales exponentially with the number of diffuse and diffraction interactions.For example, if 5,000 rays are generated for each diffuse interaction, the worst-case scenario might require 25,000 to 125,000,000 new rays to be generated, considering single- and double-reflection diffuse interactions for each launched ray. If 1,000,000 rays are emitted from each transmit antenna, and half of them are assumed to hit a diffuse surface, a total of 12.5 billion or 62.5 trillion rays might need to be traced for single-reflection diffuse interactions and two-reflection diffuse interactions, respectively. A similar complexity scaling applies for diffraction interactions. In these cases, it might not be possible to complete the ray tracing workload with a single run or with a single GPU.

[0005] In at least one embodiment, one or more circuits may cause an amount of memory used by one or more ray tracing software programs to be identified to a user. In at least one embodiment, this amount of memory may relate to the ray tracing workload. In at least one embodiment, causing the amount of memory to be identified to a user is to indicate the ray tracing workload to the user before the one or more ray tracing software programs perform the ray tracing. In at least one embodiment, the amount of memory is to enable selecting one or more rays resulting from the execution of the one or more ray tracing software programs.In at least one embodiment, the one or more rays are selected based on the amount of memory for single-run, single-processing unit, and / or single-session execution. In at least one embodiment, the amount of memory required to perform ray tracing may exceed the available memory to perform ray tracing within a processing unit, and the one or more rays may be selected to be performed in a different processing unit and / or at a different time such that the available memory is sufficient to perform ray tracing of the remaining rays. In at least one embodiment, the user may be a user of a video game, an application, and / or variations thereof.

[0006] In at least one embodiment, one or more circuits or processors may compute the ray tracing workload prior to ray tracing based at least in part on less than all of the data associated with the rays to be traced. In at least one embodiment, the ray tracing workload includes information related to ray intersections in a virtual environment, where the ray intersections include one or more of: specular reflection, diffraction, refraction, and diffuse reflection. In at least one embodiment, the ray tracing workload specifies an amount of memory required to perform the ray tracing.In at least one embodiment, less than all of the data associated with the rays to be traced is a minimum amount of data required to calculate one or more numbers of one or more types of secondary rays of a primary ray. In at least one embodiment, rays to be traced are divided into two or more ray groups based at least in part on the ray tracing workload and available memory of one or more processing units, and ray tracing of the two or more ray groups is performed using the one or more processing units. In at least one embodiment, ray tracing is to be performed by two or more graphics processing units (GPUs) in parallel.In at least one embodiment, generating a ray tracing workload is referred to as lightweight ray tracing, in which only numbers and types of secondary rays are traced. In at least one embodiment, lightweight ray tracing is performed before ray tracing primary rays. In at least one embodiment, a lightweight ray is a ray without all data associated with the ray or a ray that contains less than all data. In at least one embodiment, a lightweight ray contains less information than a ray to be traced (via path tracing). In at least one embodiment, a lightweight ray is extracted or computed from a ray to be traced by a ray tracing engine, module, and / or software.

[0007] The techniques presented here represent an improvement over previous solutions at least because knowledge of the ray tracing workload prior to ray tracing allows rays to be distributed across multiple ray tracing runs or distributed across multiple GPU devices to ensure completion of the ray tracing simulation and / or accelerate ray tracing. The techniques presented reduce or, in at least one embodiment, eliminate the need to voluntarily limit the size of the virtual scene, the number of rays to be traced, or the number of interactions in the virtual environment to fully perform computationally intensive ray tracing simulations.The presented techniques enable, in at least one embodiment, ray tracing simulations with a large number of rays and high-order diffuse scattering and diffraction interactions within a single GPU or well-scaled with many GPU devices. To further describe the present technology, examples are now given with reference to the figures.

[0008] Fig. 1 shows an example of a system 100 for performing ray tracing according to at least one embodiment. In at least one embodiment, the system 100 is implemented as shown in Fig. 1, using one or more systems, processors, or devices for communication. In at least one embodiment, ray tracing refers to a technique for simulating the propagation of light, sound, radio waves (RF), and / or variations thereof and simulating the effects of their encounters with virtual objects. In at least one embodiment, performing ray tracing consists of finding mathematical solutions to calculate intersection or interaction of a ray with various types of geometry. In at least one embodiment, ray tracing can be used in a computer graphics technique for generating images and / or a network technique for transmitting and receiving signals such as 5G (Fifth Generation), 6G (Sixth Generation), 3GPP (3rd Generation Partnership Project) standards, and / or other forms or formats of signals.

[0009] In at least one embodiment, system 100 may include primary rays 102. In at least one embodiment, primary rays 102 are rays to be traced in a ray tracing simulation within a virtual environment. In at least one embodiment, primary rays 102 may be any type of wave, including, but not limited to, light, sound, radio frequency, and signals. In at least one embodiment, primary rays 102 may be light waves from a light source, such as the sun, and the primary rays are to be traced in a virtual environment of a video game scene to generate a representation of the scene. In at least one embodiment, primary rays 102 may be cellular signals transmitted from a signal station, and the primary rays are to be traced in a virtual environment of a city model to determine the positions of receivers or signal strength at receivers.In at least one embodiment, primary beams 102 refer to beams emanating directly from a source, prior to interactions with virtual objects. In at least one embodiment, primary beams 102 include various data associated with the primary beams, such as radiance, sound intensity, electromagnetic field strength, direction, speed, frequency, magnitude, and / or variations thereof.

[0010] In at least one embodiment, lightweight primary rays 104 may be derived from primary rays 102. In at least one embodiment, lightweight primary rays 104 contain less than all of the data associated with primary rays 102. In at least one embodiment, lightweight primary rays 104 contain only scalars, such as the direction and frequency of primary rays 102. In at least one embodiment, lightweight primary rays 104 contain a minimum amount of data necessary to calculate one or more numbers of one or more types of secondary rays of primary rays 102. In at least one embodiment, secondary rays refer to rays after intersection or interaction with virtual objects. In at least one embodiment, intersection may include, but is not limited to, specular reflection, diffraction, refraction, and diffuse reflection.In at least one embodiment, types of secondary rays include mirrored secondary rays, refractive secondary rays, diffractive secondary rays, diffuse secondary rays, and / or variations thereof. In at least one embodiment, the number of secondary rays is categorized by the type of virtual surfaces or objects hit, for example, a number of hits on diffuse surfaces and a number of hits on diffractive edges. In at least one embodiment, the lightweight primary rays 104 may be divided into groups based on the workload estimated during a previous run, for example, a previous ray tracing workload 110.

[0011] In at least one embodiment, system 100 may include one or more processing units 106 having a workload estimation module 108, memory 112, and a ray tracing module 116. In at least one embodiment, processing units 106 may include one or more CPUs, GPUs, or other processors for executing modules containing software executed by the processors. In at least one embodiment, processing units 106 may be one or more graphics processing units ("GPUs"), central processing units ("CPUs"), or other parallel processing units ("PPUs"). In at least one embodiment, lightweight primary rays 104 are processed by or input to workload estimation module 108. In at least one embodiment, workload estimation module 108 performs lightweight ray tracing using lightweight primary rays 104.In at least one embodiment, lightweight ray tracing refers to ray tracing without all data associated with primary rays. In at least one embodiment, lightweight ray tracing may be partial ray tracing. For example, ray tracing with only geometric data of the primary rays. In at least one embodiment, workload estimation module 108 may be integrated with, combined with, or identical to ray tracing module 116. In at least one embodiment, both workload estimation module 108 and ray tracing module 116 perform ray tracing with input data of rays and generate ray tracing results that can be computed from the input data. In at least one embodiment, workload estimation module 108 and path tracing module 116 may share the same ray launching and shading pipeline.In at least one embodiment, the workload estimation module 108 may perform functions in a ray exploration phase. In at least one embodiment, the workload estimation module 108 requires less runtime and / or less memory than the ray tracing module 116.

[0012] In at least one embodiment, the workload estimation module 108 outputs the ray tracing workload 110. In at least one embodiment, the ray tracing workload 110 includes information related to ray intersecting in a virtual environment, where the ray intersecting includes one or more of: specular reflection, refraction, diffraction, and diffuse reflection. In at least one embodiment, the ray tracing workload 110 may store data in a scalar format. In at least one embodiment, the ray tracing workload 110 indicates the number of different types of secondary rays of the primary rays 102. For example, the ray tracing workload 110 may include a number of diffuse surface hits and / or a number of diffraction edge hits for the primary rays 102 in a virtual environment.In at least one embodiment, the ray tracing workload 110 specifies an amount of memory required to perform ray tracing of the primary rays 102 and / or lightweight primary rays 104. In at least one embodiment, the ray tracing workload 110 specifies the amount of memory required to execute the workload estimation module 108 and / or the ray tracing module 116. In at least one embodiment, the ray tracing workload 110 specifies an upper limit on the required memory space.In at least one embodiment, the ray tracing workload 110 provides numbers of various types of secondary rays after the primary rays 102 have interacted with virtual objects in a virtual environment, and a total required memory to perform ray tracing of the primary rays 102 can be calculated using a customizable table of required memory for each type of secondary rays. In at least one embodiment, in a given run, the ray tracing workload 110 is calculated before the ray tracing results 118 are obtained. In at least one embodiment, in a given run, lightweight ray tracing is performed prior to ray tracing, where lightweight ray tracing can only trace numbers and types of secondary rays.

[0013] In at least one embodiment, the required memory specified in the ray tracing workload 110 is compared to the available memory 112 in processing units 106. In at least one embodiment, if the memory 112 is less than the required memory, ray splitting may be performed on the primary rays 102 and / or the lightweight primary rays 104. In at least one embodiment, the processing units 106 may be a single GPU, and the primary rays may be split to be executed in one or more passes by the single GPU.In at least one embodiment, processing units 106 may be multiple GPUs with a master device and one or more worker devices, and the primary beams may be split into groups to be executed in parallel by the master device and / or the one or more worker devices. In at least one embodiment, beam splitting may be unnecessary if memory 112 is larger than the required memory specified by ray tracing workload 110, and primary beams 102 and / or lightweight primary beams 104 may be processed directly by ray tracing module 116 and / or workload estimation module 108.

[0014] In at least one embodiment, partitioned primary rays 114 are obtained by dividing the primary rays 102 into groups depending on the memory 112. In at least one embodiment, partitioned primary rays 114 are two or more groups of primary rays 102, where each group can be processed in one run or simulation session by one or more processing units with sufficient memory. In at least one embodiment, the partitioned primary rays 114 are partitioned based on the ray tracing workload 110 and the available memory 112 of the processing units 106. In at least one embodiment, the methods used to obtain the partitioned primary rays 114 are not limited. In at least one embodiment, the partitioned primary rays 114 may include further optimization of the primary rays 102, such as reordering the rays to be traced.In at least one embodiment, the partitioned primary rays 114 are partitioned in such a way that optimal ray tracing utilization is achieved using one or more runs and / or one or more processing units.

[0015] In at least one embodiment, the partitioned primary rays 114 are processed by or input to a ray tracing module 116. In at least one embodiment, the ray tracing module 116 performs ray tracing using the partitioned primary rays 114 and / or the primary rays 102. In at least one embodiment, the ray tracing module 116 performs ray tracing with complete data associated with primary rays. In at least one embodiment, the ray tracing module 116 may be integrated with, combined with, or identical to the workload estimation module 108. In at least one embodiment, both the workload estimation module 108 and the ray tracing module 116 perform ray tracing with input data of the rays and generate ray tracing results that can be calculated from the input data.In at least one embodiment, the workload estimation module 108 and the ray tracing module 116 may share the same ray launching and shading pipeline. In at least one embodiment, the ray tracing module 116 may perform functions in a ray tracing phase. In at least one embodiment, the ray tracing module 116 requires more runtime and / or more memory than the workload estimation module 108.

[0016] In at least one embodiment, the ray tracing module 116 generates ray tracing results 118. In at least one embodiment, the ray tracing results 118 include information in addition to the information displayed by the ray tracing workload 110. For example, the ray tracing results 118 may store data associated with a radiance, a sound intensity, an electromagnetic field strength, and / or variations thereof. In at least one embodiment, the ray tracing results 118 store data in a non-scalar format.

[0017] Fig. 2 shows an example method for performing ray tracing according to at least one embodiment. In at least one embodiment, part or all of the method 200 (or any other methods or processes described herein, or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable computer instructions and is implemented as code (e.g., executable computer instructions, one or more computer programs, or one or more applications) that is collectively executed on one or more processors by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions that may be used to perform method 200 are not stored exclusively using transient signals (e.g., a propagating transient electrical or electromagnetic transmission). In at least one embodiment, a non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) in transient signal transceivers. In at least one embodiment, method 200 is performed at least in part on a computer system as described elsewhere in this disclosure.In at least one embodiment, method 200 may be performed by a processor using neural networks. In at least one embodiment, one or more operations performed as part of method 200 may be performed in different orders and combinations than in [the original text]. Fig. 2, including parallel.

[0018] In at least one embodiment, lightweight ray tracing is performed at step 202 to obtain a number of secondary rays. In at least one embodiment, the lightweight ray tracing is performed by the workload estimation module 108, as described in Fig. 1. In at least one embodiment, the lightweight ray tracing processes lightweight primary rays, such as those described in Fig. 1, or less than all of the data associated with the primary rays. In at least one embodiment, lightweight ray tracing uses a minimum amount of data associated with the primary rays necessary to obtain the number of secondary rays. In at least one embodiment, lightweight ray tracing outputs numbers of different types of secondary rays. For example, based on a given set of lightweight primary rays generated in a virtual environment, lightweight ray tracing may generate a first number of secondary rays caused by specular reflection, a second number of secondary rays caused by diffraction, a third number of diffuse secondary rays, and a fourth number of secondary rays caused by refraction.

[0019] In at least one embodiment, in step 204, the ray tracing workload is obtained based on the number of secondary rays obtained in step 202. In at least one embodiment, the ray tracing workload indicates the required computational resources, such as memory, to perform the ray tracing. In at least one embodiment, the number of secondary rays is grouped, divided, or sorted according to the types of secondary rays. In at least one embodiment, each type of secondary ray is associated with a required memory to perform the tracing of the type of secondary ray, and a total amount of required memory space can be calculated when the number of each type of secondary rays is obtained in step 202. In at least one embodiment, the ray tracing workload is the ray tracing workload 110, as described in Fig. 1 described.

[0020] In at least one embodiment, the primary rays are partitioned in step 206 based on the ray tracing workload obtained in step 204 and the available memory of processing units performing ray tracing. In at least one embodiment, partitioning the primary rays means grouping the primary rays into two or more groups, where each group of primary rays can be processed in a single run. For example, to perform ray tracing for 100 primary rays, 2 GB of memory is required, as indicated by the ray tracing workload obtained in step 204, and a single GPU may only have 1 GB of available memory. In this case, the primary rays can be split into two groups of 50 primary rays, and the single GPU can perform ray tracing of the two groups in two runs or sessions with sufficient memory.

[0021] In at least one embodiment, in step 208, ray tracing is performed using the partitioned primary rays obtained in step 206. In at least one embodiment, the partitioned primary rays may comprise the Fig. 1 partitioned primary beams 114. In at least one embodiment, the ray tracing may be performed by the ray tracing module 116 according to Fig. 1. In at least one embodiment, ray tracing of partitions of the primary rays may be performed sequentially by a single GPU until all primary rays are processed. In at least one embodiment, ray tracing of partitions of the primary rays may be performed by multiple GPUs in parallel in one run or session of ray tracing.

[0022] In at least one embodiment, step 210 determines whether additional rays need to be processed. In at least one embodiment, step 210 may be performed at the end of a current run and / or at the beginning of a next run. In at least one embodiment, steps 202-208 are repeated for additional rays that need to be processed.

[0023] Fig. 3 shows an example of a method 300 for performing ray tracing according to at least one embodiment. In at least one embodiment, part or all of the method 300 (or other processes or methods described herein, or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable computer instructions and is implemented as code (e.g., executable computer instructions, one or more computer programs, or one or more applications) that executes collectively on one or more processors by hardware, software, or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions executable by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions that may be used to perform method 300 are not stored exclusively using transient signals (e.g., a propagating transient electrical or electromagnetic transmission). In at least one embodiment, a non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) in transient signal transceivers. In at least one embodiment, method 300 is performed at least in part on a computer system as described elsewhere in this disclosure.In at least one embodiment, method 300 may be performed by a processor using neural networks. In at least one embodiment, one or more operations performed as part of method 300 may be performed in different orders and combinations than in [the original text]. Fig. 3 shown, also in parallel.

[0024] In at least one embodiment, at step 302, the required memory space to perform ray tracing using less than all of the data associated with the rays to be processed is determined. In at least one embodiment, the less than all of the data associated with the rays may relate to properties, values, and / or parameters of the rays required to count secondary rays in a virtual environment. In at least one embodiment, the required memory space to perform ray tracing may be specified by the ray tracing workload, as described in Fig. 1 and Fig. 2. In at least one embodiment, the required storage space is calculated by summing products of the required storage space of each type of secondary beams and the number of each type of secondary beams.

[0025] In at least one embodiment, the available memory space on processing units is determined in step 304. In at least one embodiment, the available memory space may be memory 112, as shown in Fig. 1. In at least one embodiment, the available memory may be the total memory on one or more processing units that can be allocated to perform ray tracing. In at least one embodiment, the available memory may indicate memory locations of individual processing units in a set of processing units, for example, multiple GPUs with a master device and worker devices.

[0026] In at least one embodiment, a comparison is performed to determine if the required memory space is less than the available memory space. In at least one embodiment, if the required memory space is greater than the available memory space, additional memory management steps are performed, such as those described in step 308. In at least one embodiment, if the required memory space is less than the available memory space, additional memory management steps may not be required, and step 310 may be performed for all primary beams to be processed.In at least one embodiment, a GPU performing ray tracing may have less available memory than the required memory, in which case step 308 is performed to divide the rays into groups, where each group of rays requires less memory to process than the available memory on the GPU, and where the GPU can process a group of rays at a time or in each run or simulation. In at least one embodiment, multiple GPUs performing ray tracing may have less available memory than the required memory. Then step 308 is performed to divide rays into groups, where each group of rays can be processed by one GPU without memory issues, and the multiple GPUs can process groups of rays at the same time in a run or simulation.

[0027] In at least one embodiment, the beams are divided into beam groups in step 308. In at least one embodiment, the primary beams 114 may be partitioned or divided for the beam groups, as described in Fig. 1. In at least one embodiment, a group of rays is processed by a processing unit in one run.

[0028] In at least one embodiment, in step 310, ray tracing is performed using the processing units discussed with step 304. In at least one embodiment, ray tracing may be performed by the ray tracing module 116, as described in Fig. 1. In at least one embodiment, ray tracing may be performed for all rays in a run, session, or simulation, and may also be performed for groups of rays simultaneously or sequentially, depending on the comparison result of step 306 and whether step 308 is performed. In at least one embodiment, ray tracing in step 310 eliminates or reduces the out-of-memory issues.

[0029] Fig. 4 shows an example of a processor 400 with modules for performing ray tracing according to at least one embodiment. In at least one embodiment, the processor 400 performs one or more methods as described with reference to Fig. 1-3 for performing ray tracing.

[0030] In at least one embodiment, the processor 400 includes one or more processors as described in connection with the Fig. 17 to 30. In at least one embodiment, the processor 400 is any suitable processing unit or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, or PPUs. In at least one embodiment, the processor 400 includes a workload estimation module 402, a partition module 404, and a ray tracing module 406. In at least one embodiment, the workload estimation module 402, the partition module 404, and the ray tracing module 406 are part of the processor 400, as in the example of Fig. 4, or may be part of one or more other processors. In at least one embodiment, the workload estimation module 402, the partition module 404, and the ray tracing module 406 are distributed across multiple processors that are communicated via a bus, a network, by writing to a shared memory, or any suitable communication process, such as the methods described with respect to the Fig. 17-30 described, communicate.

[0031] In at least one embodiment, the workload estimation module 402 includes circuitry for estimating the ray tracing workload prior to performing the ray tracing. In at least one embodiment, the workload estimation module 402 may be the workload estimation module 108, as described in Fig. 1. In at least one embodiment, the workload estimation module 402 may perform operations to determine the Fig. 2 and / or the steps 202-204 shown in Fig. 3. In at least one embodiment, the workload estimation module 402 may be integrated with, combined with, or identical to the ray tracing module 406. In at least one embodiment, the workload estimation module 402 is a module for performing ray tracing when less than all ray data or lightweight rays are given and outputs less than all ray tracing results, for example, it may only output the number and types of secondary rays.

[0032] In at least one embodiment, the partition module 404 includes circuitry for dividing, subdividing, or partitioning the workload and / or primary beams to be processed. In at least one embodiment, the partition module 404 allocates available memory of processing units and / or generates a memory allocation plan for the available memory, such as memory 112, as shown in Fig. 1. In at least one embodiment, the partition module 404 may perform operations to Fig. 2 and / or the step 206 shown in Fig. 3 to implement steps 304-308.

[0033] In at least one embodiment, the ray tracing module 406 includes circuitry for performing a ray tracing simulation. In at least one embodiment, the ray tracing module 406 may be the ray tracing module 116, as described in Fig. 1. In at least one embodiment, the ray tracing module 406 may perform operations to Fig. 2 and / or the step 208 shown in Fig. 3. In at least one embodiment, the ray tracing module 406 may be integrated with, combined with, or identical to the workload estimation module 402.

[0034] Fig. 5 shows an example of a block diagram illustrating a driver and / or runtime environment including one or more libraries to provide one or more application programming interfaces (APIs), according to at least one embodiment. In at least one embodiment, a software program 502 is a software module. In at least one embodiment, a software program 502 includes one or more software modules. In at least one embodiment, one or more software modules are furthermore not exclusively in Fig. 4. In at least one embodiment, one or more APIs 510 are sets of software instructions that, when executed, cause one or more processors to perform one or more computational operations. In at least one embodiment, one or more APIs 510 are distributed or otherwise provided as part of one or more libraries 506, runtime environments 504, drivers 504, and / or any other grouping of software and / or executable code as described herein. In at least one embodiment, one or more APIs 510 perform one or more computational operations in response to being invoked by software programs 502. In at least one embodiment, a software program 502 is a collection of software code, commands, instructions, or other text sequences to instruct a computing device to perform one or more operations and / or to execute one or more other sets of instructions, such asAPIs 510 or API functions 512 for execution. In at least one embodiment, the functions provided by one or more APIs 510 include software functions 512, such as those that can be used to accelerate one or more portions of software programs 502 using one or more parallel processing units (PPUs), such as graphics processing units (GPUs).

[0035] In at least one embodiment, APIs 510 are hardware interfaces to one or more circuits to perform one or more computational operations. In at least one embodiment, one or more software APIs 510 as described herein are implemented as one or more circuits to perform one or more techniques or methods associated with the Fig. 1-3. In at least one embodiment, one or more software programs 502 include instructions that, when executed, cause one or more devices and / or circuits to perform one or more techniques or methods further associated with the Fig. 1-3 are described.

[0036] In at least one embodiment, software programs 502, such as user-implemented software programs, utilize one or more application programming interfaces (APIs) 510 to perform various computational operations, such as memory allocation, matrix multiplication, arithmetic operations, or any computational operations performed by parallel processing units (PPUs), such as graphics processing units (GPUs), as described herein. In at least one embodiment, one or more APIs 510 provide a set of callable functions 512, referred to herein as APIs, API functions, and / or functions, that individually perform one or more computational operations, such as computational operations related to parallel processing.

[0037] In at least one embodiment, one or more software programs 502 interact with or otherwise communicate with one or more APIs 510 to perform one or more computational operations using one or more PPUs, such as GPUs. In at least one embodiment, one or more computational operations using one or more PPUs include at least one or more groups of computational operations to be accelerated by execution at least in part by the one or more PPUs. In at least one embodiment, one or more software programs 502 interact with one or more APIs 510 to enable parallel computing using a remote or local interface.

[0038] In at least one embodiment, an interface is software instructions that, when executed, provide access to one or more functions 512 provided by one or more APIs 510. In at least one embodiment, a software program 502 uses a local interface when a software developer compiles one or more software programs 502 in conjunction with one or more libraries 506 that include or otherwise provide access to one or more APIs 510. In at least one embodiment, one or more software programs 502 are statically compiled in conjunction with precompiled libraries 506 or uncompiled source code that includes instructions for executing one or more APIs 510.In at least one embodiment, one or more software programs 502 are dynamically compiled, and the one or more software programs use a linker to link one or more precompiled libraries 506 that include one or more APIs 510.

[0039] In at least one embodiment, a software program 502 uses a remote interface when a software developer executes a software program that uses or otherwise communicates with a library 506 comprising one or more APIs 510 over a network or other remote communication medium. In at least one embodiment, one or more libraries 506 comprising one or more APIs 510 are to be executed by a remote computing service, such as a computing resource service provider. In another embodiment, one or more libraries 506 comprising one or more APIs 510 are to be executed by any other computer host that provides the one or more APIs 510 to one or more software programs 502.

[0040] In at least one embodiment, a processor executing or using one or more software programs 502 invokes, uses, executes, or otherwise implements one or more APIs 510 to allocate and otherwise manage memory to be used by the software programs 502. In at least one embodiment, one or more software programs 502 utilize one or more APIs 510 to allocate and otherwise manage memory to be used by one or more portions of the software programs 502 to be accelerated using one or more PPUs, such as GPUs, or another accelerator or processor, as further described herein.These software programs 502 may be executed by one or more processors, depending at least in part on the latency of connections coupled to the one or more processors, using functions 512 provided, in one embodiment, by one or more APIs 510.

[0041] In at least one embodiment, an API 510 is an API for facilitating parallel processing. In at least one embodiment, an API 510 is any other API as further described herein. In at least one embodiment, an API 510 is provided by a driver and / or runtime 504. In at least one embodiment, an API 510 is provided by a CUDA user-mode driver. In at least one embodiment, an API 510 is provided by a CUDA runtime. In at least one embodiment, a driver 504 corresponds to data values ​​and software instructions that, when executed, perform or otherwise enable the operation of one or more functions 512 of an API 510 during the loading and execution of one or more portions of a software program 502.In at least one embodiment, a runtime environment 504 consists of data values ​​and software instructions that, when executed, perform or otherwise enable the operation of one or more functions 512 of an API 510 during execution of a software program 502. In at least one embodiment, one or more software programs 502 utilize one or more APIs 510 implemented or otherwise provided by a driver and / or runtime environment 504 to perform combined arithmetic operations by the one or more software programs 502 during execution by one or more PPUs, such as GPUs.

[0042] In at least one embodiment, one or more software programs 502 utilize one or more APIs 510 provided by a driver and / or runtime environment 504 to perform combined arithmetic operations of one or more PPUs, such as GPUs. In at least one embodiment, one or more APIs 510 provide combined arithmetic operations via a driver and / or runtime environment 504, as described above. In at least one embodiment, one or more software programs 502 utilize one or more APIs 510 provided by a driver and / or runtime environment 504 to allocate or otherwise reserve one or more memory blocks 514 to one or more PPUs, such as GPUs.In at least one embodiment, one or more software programs 502 use one or more APIs 510 provided by a driver and / or runtime environment 504 to allocate or otherwise reserve memory blocks. In at least one embodiment, one or more APIs 510 are designed to perform combined arithmetic operations, as described in connection with the . Fig. 1-3 described.

[0043] To improve the usability of software programs 502 and / or to accelerate the optimization of one or more portions of the software programs 502 by one or more PPUs, such as GPUs, in one embodiment, one or more APIs 510 provide one or more API functions 512 to execute a system implemented by one or more computing devices as described above and further in conjunction with the Fig. 1-3. In at least one embodiment, an example block diagram 500 illustrates a processor including one or more circuits to execute one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, an example block diagram 500 illustrates a system including one or more processors to execute one or more software programs to combine two or more application programming interfaces (APIs) into a single API.

[0044] In at least one embodiment, parts, methods and / or a system used in conjunction with Fig. 5, furthermore in the Fig. 1-4 not exclusively shown. DATA CENTER

[0045] Fig. 6 shows an example of a data center 600 in which at least one embodiment may be used. In at least one embodiment, the data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640.

[0046] In at least one embodiment, as described in Fig. 6, the data center infrastructure layer 610 may include a resource orchestrator 612, clustered compute resources 614, and node compute resources ("Node CRs") 616(1)-616(N), where "N" represents any integer positive number. In at least one embodiment, the Node CRs 616(1)-616(N) may include any number of central processing units ("CPUs") or other processors (including accelerators, Field Programmable Gate Arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or hard disk drives), network input / output devices ("NW I / O"), network switches, virtual machines ("VMs"), power modules and cooling modules, etc. In at least one embodiment, one or more of the Node CRs may bes 616(1)-616(N) may be a server that has one or more of the computing resources mentioned above.

[0047] In at least one embodiment, the grouped computing resources 614 may include separate groupings of node CRs housed in one or more racks (not shown) or multiple racks housed in data centers in different geographic locations (also not shown). In at least one embodiment, separate groupings of node CRs within the grouped computing resources 614 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs including CPUs or processors may be grouped in one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0048] In at least one embodiment, resource orchestrator 612 may design or otherwise control one or more node CRs 616(1)-616(N) and / or clustered computing resources 614. In at least one embodiment, resource orchestrator 612 may include a software design infrastructure ("SDI") management entity for data center 600. In at least one embodiment, resource orchestrator may include hardware, software, or a combination thereof.

[0049] In at least one embodiment, as described in Fig. 6, the framework layer 620 includes a job scheduler 632, a configuration manager 634, a resource manager 636, and a distributed file system 638. In at least one embodiment, the framework layer 620 may include a framework for supporting the software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application(s) 642 may each include web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may be some type of free and open source software web application framework such as Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 638 for processing large amounts of data (e.g., "Big Data").In at least one embodiment, the job scheduler 632 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 600. In at least one embodiment, the configuration manager 634 may be capable of configuring various layers, such as the software layer 630 and the framework layer 620, which includes Spark and the distributed file system 638, to support processing large amounts of data. In at least one embodiment, the resource manager 636 may be capable of managing clustered or grouped computing resources allocated or assigned to support the distributed file system 638 and the job scheduler 632. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 614 in the infrastructure layer 610 of the data center.In at least one embodiment, the resource manager 636 may be coordinated with the resource orchestrator 612 to manage these allocated or assigned computing resources.

[0050] In at least one embodiment, the software 632 included in software layer 630 may include software used by at least portions of node CRs 616(1)-616(N), clustered computing resources 614, and / or distributed file system 638 of framework layer 620. In at least one embodiment, one or more types of software may include, but are not limited to, internet search software, email virus scanning software, database software, and streaming video content software.

[0051] In at least one embodiment, the application(s) 642 included in the application layer 640 may include one or more types of applications used by at least portions of the node CRs 616(1)-616(N), clustered compute resources 614, and / or the distributed file system 638 of the framework layer 620. In at least one embodiment, one or more types of applications may include any number of genomic applications, cognitive computation, and machine learning applications, including, but not limited to, training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in connection with one or more embodiments.

[0052] In at least one embodiment, each of configuration manager 634, resource manager 636, and resource orchestrator 612 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may relieve a data center operator of data center 600 from potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of a data center.

[0053] In at least one embodiment, data center 600 may include tools, services, software, or other resources for training one or more machine learning models or predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weighting parameters according to a neural network architecture using software and computing resources described above with respect to data center 600.In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 600 by using weighting parameters calculated by one or more training techniques described herein.

[0054] In at least one embodiment, the data center may use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inferencing using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or perform information inferencing, such as image recognition, speech recognition, or other artificial intelligence services.

[0055] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0056] Fig. 7A shows an example of an autonomous vehicle 700 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 700 (alternatively referred to herein as "vehicle 700") may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, the vehicle 700 may be a semi-trailer truck used to transport goods. In at least one embodiment, the vehicle 700 may be an aircraft, a robotic vehicle, or another type of vehicle.

[0057] Autonomous vehicles may be described in terms of automation levels defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of that standard). In one or more embodiments, vehicle 700 may be capable of performing functionality according to one or more of Levels 1 through 5 of the levels of autonomous driving. For example, in at least one embodiment, vehicle 700 may be capable of conditionally automated (Level 3), highly automated (Level 4), and / or fully automated (Level 5), depending on the embodiment.

[0058] In at least one embodiment, vehicle 700 may include, without limitation, components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 700 may include, without limitation, a propulsion system 750, such as an internal combustion engine, a hybrid electric drive, a pure electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 750 may be connected to a drivetrain of vehicle 700, which may include, among other things, a transmission, to enable propulsion of vehicle 700. In at least one embodiment, propulsion system 750 may be controlled in response to receiving signals from a throttle / accelerator pedal(s) 752.

[0059] In at least one embodiment, a steering system 754, which may include, without limitation, a steering wheel, is used to steer a vehicle 700 (e.g., along a desired path or route) when a propulsion system 750 is operating (e.g., when the vehicle is in motion). In at least one embodiment, a steering system 754 may receive signals from one or more steering actuators 756. In at least one embodiment, the steering wheel may optionally be employed for full automation (Level 5). In at least one embodiment, a brake sensor system 746 may be used to apply the vehicle brakes in response to receiving signals from one or more brake actuators 748 and / or brake sensors.

[0060] In at least one embodiment, the controller(s) 736, which may include, without limitation, one or more system-on-chips (“SoCs”) (in Fig. 7A not shown) and / or graphics processing units (“GPUs”), send signals (e.g., representative of commands) to one or more components and / or systems of the vehicle 700. In at least one embodiment, the controller(s) 736 may, for example, send signals to actuate the vehicle brakes via the brake actuators 748, to actuate the steering system 754 via the steering actuator(s) 756, and to actuate the propulsion system 750 via a throttle / accelerator pedal(s) 752. In at least one embodiment, the controller(s) 736 may include one or more on-vehicle (e.g., on-board) computing devices (e.g., supercomputers) that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in operating the vehicle 700.In at least one embodiment, the controller(s) 736 may include a first controller 736 for autonomous driving functions, a second controller 736 for functional safety functions, a third controller 736 for artificial intelligence functions (e.g., computer vision), a fourth controller 736 for infotainment functions, a fifth controller 736 for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller 736 may perform two or more of the above functions, two or more controllers 736 may perform a single function, and / or any combination thereof.

[0061] In at least one embodiment, the controller(s) 736 provides signals to control one or more components and / or systems of the vehicle 700 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data may be obtained, for example and without limitation, from Global Navigation Satellite Systems (“GNSS”) sensor(s) 758 (e.g., Global Positioning System sensor(s)), RADAR sensor(s) 760, ultrasonic sensor(s) 762, LIDAR sensor(s) 764, Inertial Measurement Unit (“IMU”) sensor(s) 766 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 796, stereo camera(s) 768, wide-angle camera(s) 770 (e.g., fisheye cameras), infrared camera(s) 772, ambient camera(s) 774 (e.g., 360-degree cameras), remote cameras (not included in Fig. 7A), mid-range camera(s) (not shown in Fig. 7A), speed sensor(s) 744 (e.g., for measuring the speed of the vehicle 700), vibration sensor(s) 742, steering sensor(s) 740, brake sensor(s) (e.g., as part of the brake sensor system 746), and / or other types of sensors.

[0062] In at least one embodiment, one or more of the controllers 736 may receive inputs (e.g., in the form of input data) from an instrument cluster 732 of the vehicle 700 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (“HMI”) display 734, an audible annunciator, a speaker, and / or via other components of the vehicle 700. In at least one embodiment, the outputs may include information such as vehicle speed, RPM, time, map data (e.g., a high-resolution map (in Fig. 7A not shown)), position data (e.g., the position of the vehicle 700, as on a map), direction, position of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 736, etc. In at least one embodiment, the HMI display 734 may, for example, display information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about maneuvers the vehicle has performed, is currently performing, or will perform (e.g., lane change now, exit 34B in two miles, etc.).

[0063] In at least one embodiment, the vehicle 700 further includes a network interface 724 that may utilize wireless antenna(s) 726 and / or modem(s) to communicate over one or more networks. For example, in at least one embodiment, the network interface 724 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. In at least one embodiment, the wireless antenna(s) 726 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low-power wide area networks ("LPWANs") such as LoRaWAN, SigFox, etc.

[0064] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0065] Fig. Figure 7B shows an example of camera positions and fields of view for the autonomous vehicle 700 of Fig. 7A, according to at least one embodiment. In at least one embodiment, the cameras and the respective fields of view represent an example embodiment and are not to be considered limiting. For example, in at least one embodiment, additional and / or alternative cameras may be present and / or the cameras may be located at other locations on the vehicle 700.

[0066] In at least one embodiment, the camera types may include, but are not limited to, digital cameras that may be adapted for use with components and / or systems of the vehicle 700. In at least one embodiment, the camera(s) may operate at Automotive Safety Integrity Level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera types may achieve any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the cameras may use rolling shutter, global shutter, another shutter type, or a combination thereof.In at least one embodiment, the color filter array may include a red-clear-clear-clear color filter array ("RCCC"), a red-clear-clear-blue color filter array ("RCCB"), a red-blue-green-clear color filter array ("RBGC"), a Foveon X3 color filter array, a Bayer sensor color filter array ("RGGB"), a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, clear-pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.

[0067] In at least one embodiment, one or more cameras may be used to implement advanced driver assistance systems ("ADAS") (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multifunction mono camera may be installed, providing functions such as lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).

[0068] In at least one embodiment, one or more cameras may be mounted in a mounting arrangement, such as a custom-designed (three-dimensional ("3D") printed) arrangement, to eliminate stray light and reflections from the vehicle interior (e.g., reflections from the dashboard reflected in the windshield mirrors) that may impair the camera's ability to collect image data. In at least one embodiment, the assemblies for the outside mirrors may be custom 3D printed so that the camera mounting plate conforms to the shape of the outside mirror. In at least one embodiment, the camera(s) may be integrated into the outside mirror. In at least one embodiment, for side-mounted cameras, the camera(s) may also be integrated into four pillars at each corner of the vehicle.

[0069] In at least one embodiment, cameras with a field of view including portions of the environment in front of the vehicle 700 (e.g., forward-facing cameras) may be used for surround vision to assist in detecting forward paths and obstacles, as well as, with the assistance of one or more controllers 736 and / or control SoCs, to provide information critical to establishing an occupancy grid and / or determining preferred vehicle paths. In at least one embodiment, forward-facing cameras may be used to perform many of the same ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance.In at least one embodiment, forward-facing cameras may also be used for ADAS features and systems, including but not limited to lane departure warning (“LDW”), autonomous cruise control (“ACC”), and / or other features such as traffic sign recognition.

[0070] In at least one embodiment, a plurality of cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform having a complementary metal oxide semiconductor (CMOS) color image sensor. In at least one embodiment, the wide-angle camera 770 may be used to detect objects entering the field of view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. 7B, in other implementations, any number (including zero) of wide-angle cameras 770 may be present on the vehicle 700. In at least one embodiment, any number of wide-angle cameras 798 (e.g., a wide-angle stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the wide-angle camera(s) 798 may also be used for object detection and classification, as well as basic object tracking.

[0071] In at least one embodiment, any number of stereo cameras 768 may also be present in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 768 may have an integrated control unit including a scalable processing unit that may provide a programmable logic ("FPGA") and a multi-core microprocessor with an integrated Controller Area Network ("CAN") or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to create a 3D map of the surroundings of the vehicle 700 that includes a distance estimate for all points in the image.In at least one embodiment, one or more of the stereo cameras 768 may comprise, without limitation, compact stereo vision sensors, which may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance between the vehicle 700 and the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo cameras 768 may also be used in addition to or alternatively to those described herein.

[0072] In at least one embodiment, cameras with a field of view including portions of the environment to the side of the vehicle 700 (e.g., side cameras) may be used for the environment view and provide information used to create and update the occupancy grid and to generate side impact warnings. In at least one embodiment, the environment camera(s) 774 (e.g., four environment cameras 774, as shown in Fig. 7B) may be positioned on the vehicle 700. In at least one embodiment, the surround camera(s) 774 may include, without limitation, any number and combination of wide-angle camera(s) 770, fisheye camera(s), 360-degree camera(s), and / or the like. For example, in at least one embodiment, four fisheye cameras may be positioned on the front, rear, and sides of the vehicle 700. In at least one embodiment, the vehicle 700 may utilize three surround camera(s) 774 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.

[0073] In at least one embodiment, cameras with a field of view including portions of the environment behind the vehicle 700 (e.g., rearview cameras) may be used for parking assistance, surround view, rear collision warnings, and occupancy grid creation and updating. In at least one embodiment, a variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., wide-range cameras 798 and / or mid-range camera(s) 776, stereo camera(s) 768), infrared camera(s) 772, etc.), as described herein.

[0074] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0075] Fig. 7C is a block diagram illustrating an example system architecture for the autonomous vehicle 700 of Fig. 7A according to at least one embodiment. In at least one embodiment, each component, feature, and system of the vehicle 700 is Fig. 7C as connected via a bus 702. In at least one embodiment, bus 702 may include, without limitation, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, a CAN may be a network within vehicle 700 used to support the control of various features and functions of vehicle 700, such as brake application, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 702 may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 702 may be read to determine steering wheel angle, vehicle speed, engine revolutions per minute ("RPMs"), button positions, and / or other vehicle status indicators.In at least one embodiment, bus 702 may be a CAN bus that is ASIL B compliant.

[0076] In at least one embodiment, FlexRay and / or Ethernet may be used in addition to or alternatively to CAN. In at least one embodiment, any number of buses 702 may be present, including, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using a different protocol. In at least one embodiment, two or more buses 702 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 702 may be used for collision avoidance functionality and a second bus 702 may be used for actuation control. In at least one embodiment, each bus 702 may communicate with any components of the vehicle 700, and two or more buses 702 may communicate with the same components.In at least one embodiment, any number of system(s) on chip(s) ("SoC(s)") 704, each controller 736, and / or each computer in the vehicle may have access to the same input data (e.g., inputs from sensors of the vehicle 700) and be connected to a common bus, such as the CAN bus.

[0077] In at least one embodiment, the vehicle 700 may include one or more controllers 736 as described herein with respect to Fig. 7A. In at least one embodiment, the controller(s) 736 may be used for a variety of functions. In at least one embodiment, the controller(s) 736 may be coupled to various other components and systems of the vehicle 700 and may be used for control of the vehicle 700, artificial intelligence of the vehicle 700, infotainment for the vehicle 700, and / or the like.

[0078] In at least one embodiment, the vehicle 700 may include any number of SoCs 704. Each of the SoCs 704 may include, without limitation, central processing units ("CPU(s)") 706, graphics processing units ("GPU(s)") 708, processor(s) 710, cache(s) 712, accelerators 714, data storage 716, and / or other components and features not shown. In at least one embodiment, SoC(s) 704 may be used to control the vehicle 700 in a variety of platforms and systems. For example, in at least one embodiment, SoC(s) 704 may be combined in a system (e.g., the system of the vehicle 700) with a high-definition ("HD") card 722 that may be accessed via a network interface 724 from one or more servers (in Fig. 7C not shown) can receive map refreshes and / or updates.

[0079] In at least one embodiment, the CPU(s) 706 may comprise a CPU cluster or CPU complex (alternatively referred to herein as a "CCPLEX"). In at least one embodiment, the CPU(s) 706 may comprise multiple cores and / or Level Two ("L2") caches. In at least one embodiment, the CPU(s) 706 may comprise, for example, eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU(s) 706 may comprise four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). In at least one embodiment, the CPU(s) 706 (e.g., CCPLEX) may be configured to support concurrent cluster operation such that any combination of clusters of the CPU(s) 706 may be active at any time.

[0080] In at least one embodiment, one or more of the CPU(s) 706 may implement power management features including, without limitation, one or more of the following: individual hardware blocks may be automatically clocked when idle to conserve dynamic power; each core clock may be clocked when the core is not actively executing instructions due to the execution of Wait for Interrupt ("WFI") / Wait for Event ("WFE") instructions; each core may be independently power-driven; each core cluster may be independently clock-driven if all cores are clock-driven or power-driven; and / or each core cluster may be independently power-driven if all cores are power-driven.In at least one embodiment, the CPU(s) 706 may further implement an advanced power state management algorithm, where allowable power states and expected wake-up times are determined, and the hardware / microcode determines the best power state to enter for the core, cluster, and CCPLEX. In at least one embodiment, the processor cores may support simplified sequences for entering the power state in software, offloading the work to the microcode. In at least one embodiment, processor cores are referred to as compute units or computation units.

[0081] In at least one embodiment, the GPU(s) 708 may include an integrated GPU (alternatively referred to herein as an "iGPU"). In at least one embodiment, the GPU(s) 708 may be programmable and efficient for parallel workloads. In at least one embodiment, the GPU(s) 708 may use an extended Tensor instruction set. In at least one embodiment, the GPU(s) 708 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache with a memory capacity of at least 96 KB) and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with a memory capacity of 512 KB). In at least one embodiment, the GPU(s) 708 may include at least eight streaming microprocessors.In at least one embodiment, the GPU(s) 708 may use one or more application programming interfaces (APIs) for computation. In at least one embodiment, the GPU(s) 708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0082] In at least one embodiment, one or more of the GPU(s) 708 may be power-optimized for best performance in automotive and embedded use cases. For example, in one embodiment, the GPU(s) 708 may be fabricated with fin field-effect transistors ("FinFETs"). In at least one embodiment, each streaming microprocessor may include a number of mixed-precision compute cores divided into multiple blocks. For example, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In at least one embodiment, each processing block may be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA TENSOR COREs with mixed precision for deep learning matrix arithmetic, a level zero ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file.In at least one embodiment, streaming microprocessors may include independent parallel integer and floating-point datapaths to enable efficient execution of workloads with a mix of computations and addressing calculations. In at least one embodiment, streaming microprocessors may include an independent thread scheduling function to enable finer-grained synchronization and collaboration between parallel threads. In at least one embodiment, streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0083] In at least one embodiment, one or more of the GPU(s) 708 may include high-bandwidth memory ("HBM") and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, synchronous graphics random-access memory ("SGRAM"), such as synchronous graphics double-data-rate random-access memory type 5 ("GDDR5"), may be used in addition to or alternatively to the HBM memory.

[0084] In at least one embodiment, the GPU(s) 708 may include unified memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU(s) 708 to directly access page tables of the CPU(s) 706. In at least one embodiment, an address translation request may be communicated to the CPU(s) 706 when the memory management unit ("MMU") of the GPU(s) 708 detects a fault. In response, the CPU(s) 706 may look up a virtual-to-physical mapping of the address in its page tables and, in at least one embodiment, transmit the translation back to the GPU(s) 708.In at least one embodiment, unified memory technology may enable a single, unified virtual address space for the memory of both the CPU(s) 706 and the GPU(s) 708, thereby simplifying programming of the GPU(s) 708 and connecting applications to the GPU(s) 708.

[0085] In at least one embodiment, the GPU(s) 708 may include any number of access counters that can track the frequency of access by the GPU(s) 708 to the memory of other processors. In at least one embodiment, access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses pages most frequently, thereby improving the efficiency of memory regions shared between processors.

[0086] In at least one embodiment, one or more of the SoC(s) 704 may include any number of cache(s) 712, including those described herein. For example, in at least one embodiment, the cache(s) 712 may include a Level 3 ("L3") cache available to both the CPU(s) 706 and the GPU(s) 708 (e.g., connected to both the CPU(s) 706 and the GPU(s) 708). In at least one embodiment, the cache(s) 712 may include a write-back cache that can track line states, e.g., by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may be 4 MB or larger, depending on the embodiment, although smaller cache sizes may be used.

[0087] In at least one embodiment, one or more of the SoC(s) 704 may include one or more accelerators 714 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC(s) 704 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement the GPU(s) 708 and offload some tasks of the GPU(s) 708 (e.g., to free up more cycles of the GPU(s) 708 to perform other tasks).In at least one embodiment, the accelerator(s) 714 may be used for targeted workloads (e.g., perception, convolutional neural networks ("CNNs"), feedback neural networks ("RNNs"), etc.) that are robust enough to be suitable for acceleration. In at least one embodiment, a CNN may include a region-based or regional convolutional neural network ("RCNNs") and a fast RCNN (e.g., as used for object detection), or another type of CNN.

[0088] In at least one embodiment, the accelerator(s) 714 (e.g., hardware acceleration cluster) may include a deep learning accelerator ("DLA"). A DLA(s) may include, without limitation, one or more tensor processing units ("TPUs"), which may be configured to provide an additional tens of trillion operations per second for deep learning applications and inferencing. In at least one embodiment, the TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA(s) may further be optimized for a specific set of neural network types and floating-point operations, as well as for inferencing. In at least one embodiment, the design of the DLA(s) may provide more performance per millimeter than a typical general-purpose GPU and typically far exceeds the performance of a CPU.In at least one embodiment, the TPU(s) may perform multiple functions, including a single-instance convolution function that supports, for example, INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.In at least one embodiment, DLA(s) can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for a variety of functions, including, for example and without limitation: a CNN for object identification and recognition using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for emergency vehicle detection and identification and recognition using data from microphones 796; a CNN for facial recognition and vehicle owner identification using data from camera sensors; and / or a CNN for safety-relevant and / or security-related events.

[0089] In at least one embodiment, DLA(s) may perform any function of GPU(s) 708, and by using an inference accelerator, a developer may, for example, dedicate either DLA(s) or GPU(s) 708 to any function. For example, in at least one embodiment, the developer may focus the processing of CNNs and floating-point operations on DLA(s) and leave other functions to GPU(s) 708 and / or one or more other accelerators 714.

[0090] In at least one embodiment, the accelerator(s) 714 (e.g., hardware acceleration cluster) may include a programmable image processing accelerator ("PVA"), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA(s) may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems ("ADAS") 738, autonomous driving, augmented reality ("AR"), and / or virtual reality ("VR") applications. PVA(s) may provide a balance between performance and flexibility. In at least one embodiment, each PVA may include, for example and without limitation, any number of reduced instruction set ("RISC") cores, direct memory access ("DMA") cores, and / or any number of vector processors.

[0091] In at least one embodiment, the RISC cores may interact with image sensors (e.g., image sensors of one of the cameras described herein), image signal processors, and / or the like. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC cores may use one of several protocols, depending on the embodiment. In at least one embodiment, RISC cores may execute a real-time operating system ("RTOS"). In at least one embodiment, RISC cores may be implemented with one or more integrated circuit devices, application-specific integrated circuits ("ASICs"), and / or memory devices. In at least one embodiment, RISC cores may include, for example, an instruction cache and / or tightly coupled RAM.

[0092] In at least one embodiment, a DMA may enable components of the PVA(s) to access system memory independently of the CPU(s) 706. In at least one embodiment, a DMA may support any number of features used to optimize the PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, a DMA may support up to six or more dimensions of addressing, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0093] In at least one embodiment, vector processors may be programmable processors that may be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may act as the primary processing unit of the PVA and may include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, the VPU core may include a digital signal processor, such asA digital signal processor with multiple data for one instruction ("SIMD") and very long instruction words ("VLIW"). In at least one embodiment, a combination of SIMD and VLIW can increase throughput and speed.

[0094] In at least one embodiment, each of the vector processors may include an instruction cache and be connected to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to operate independently of other vector processors. In at least one embodiment, vector processors included in a given PVA may be configured to use data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but for different image regions. In at least one embodiment, vector processors included in a given PVA may concurrently execute different image processing algorithms for the same image, or even different algorithms for consecutive images or portions of an image.In at least one embodiment, among other things, there may be any number of PVAs in a hardware acceleration cluster and any number of vector processors in each PVA. In at least one embodiment, the PVA(s) may include additional error correction code ("ECC") memory to increase overall system security.

[0095] In at least one embodiment, the accelerator(s) 714 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory ("SRAM") to provide high-bandwidth, low-latency SRAM to the accelerator(s) 714. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, consisting of, for example, and without limitation, eight field-configurable memory blocks accessible by both the PVA and the DLA. In at least one embodiment, each pair of memory blocks may include an extended peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used.In at least one embodiment, PVA and DLA may access the memory via a backbone that provides PVA and DLA high-speed access to the memory. In at least one embodiment, the backbone may include an on-chip computer vision network connecting PVA and DLA to the memory (e.g., using an APB).

[0096] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and the DLA are providing ready and valid signals before transmitting control signals / addresses / data. In at least one embodiment, an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. In at least one embodiment, an interface may conform to International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("IEC") 61508 standards, although other standards and protocols may be used.

[0097] In at least one embodiment, one or more of the SoC(s) 704 may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine positions and extents of objects (e.g., within a world model) to generate real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for the simulation of sonar systems, for general wave propagation simulation, for comparison with lidar data for localization, and / or for other functions, and / or for other purposes.

[0098] In at least one embodiment, the accelerator(s) 714 (e.g., hardware accelerator clusters) has / have a wide range of applications for autonomous driving. In at least one embodiment, a PVA may be a programmable image processing accelerator that can be used for critical processing steps in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of a PVA are well suited for algorithmic areas that require predictable processing with low power and low latency. In other words, a PVA is well suited for semi-dense or dense regular computations, even on small datasets, that require predictable runtimes with low latency and low power consumption. In at least one embodiment, for autonomous vehicles, such asVehicle 700, PVAs developed to run classical computer vision algorithms because they are efficient in object detection and work with integer mathematical methods.

[0099] For example, in at least one embodiment of a technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, motion estimation / stereo matching while driving is used in Level 3-5 autonomous driving applications (e.g., structure from motion, pedestrian detection, lane detection, etc.). In at least one embodiment, the PVA may perform a computer stereo vision function on inputs from two monocular cameras.

[0100] In at least one embodiment, a PVA may be used to implement dense optical flow. For example, in at least one embodiment, a PVA may process raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In at least one embodiment, a PVA is used for depth-of-flight processing, processing raw time-of-flight data to provide, for example, processed time-of-flight data.

[0101] In at least one embodiment, a DLA may be used to power any type of network to improve control and driving safety, including, for example and without limitation, a neural network that outputs a confidence measure for each object detection. In at least one embodiment, confidence may be represented or interpreted as a probability, or as the relative "weight" of each detection compared to other detections. In at least one embodiment, confidence further allows the system to make decisions about which detections should be considered true positives and which should be considered false positives. In at least one embodiment, a system may set a confidence threshold and consider only detections that exceed the threshold to be true positives.In an embodiment using an automatic emergency braking ("AEB") system, false positive detections would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. In at least one embodiment, highly confident detections may be considered triggers for AEB. In at least one embodiment, a DLA may employ a neural network to regress the confidence score. In at least one embodiment, the neural network may take as input at least a subset of parameters, such as the bounding box dimensions, the footprint estimate obtained (e.g., from another subsystem), the output of the IMU sensor(s) 766 correlated with the orientation of the vehicle 700, the distance, the object's 3D position estimates obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 764 or RADAR sensor(s) 760), and others.

[0102] In at least one embodiment, one or more SoC(s) 704 may include one or more data stores 716 (e.g., a memory). In at least one embodiment, the data store(s) 716 may be on-chip memory of the SoC(s) 704, which may store neural networks to be executed on GPU(s) 708 and / or a DLA. In at least one embodiment, the capacity of the data store(s) 716 may be large enough to store multiple instances of neural networks for redundancy and security. In at least one embodiment, the data store(s) 712 may include L2 or L3 cache(s).

[0103] In at least one embodiment, one or more of the SoC(s) 704 may include any number of processors 710 (e.g., embedded processors). In at least one embodiment, the processor(s) 710 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle the boot power and management functions and associated security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC(s) 704 and provide runtime power management services.In at least one embodiment, the processor may provide for boot power and management, clock and voltage programming, low-power system transition support, management of SoC(s) 704 temperatures and temperature sensors, and / or management of SoC(s) 704 power states. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and SoC(s) 704 may use ring oscillators to sense temperatures of CPU(s) 706, GPU(s) 708, and / or accelerator(s) 714.In at least one embodiment, if temperatures are determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine and place the SoC(s) 704 into a lower power state and / or place the vehicle 700 into a chauffeur-safe stop mode (e.g., bring the vehicle 700 to a safe stop).

[0104] In at least one embodiment, the processor(s) 710 may further comprise a set of embedded processors that can serve as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0105] In at least one embodiment, the processor(s) 710 may further include an "always on" processor engine that may provide the necessary hardware functions to support low-power sensor management and wake-up use cases. In at least one embodiment, the "always on" processor engine may include, without limitation, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0106] In at least one embodiment, the processor(s) 710 may further comprise a security cluster engine, including, without limitation, a dedicated processor subsystem for handling security management for automotive applications. In at least one embodiment, the security cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a security mode, in at least one embodiment, two or more cores may operate in a lockstep mode, functioning as a single core with comparison logic to detect any differences between their operations.In at least one embodiment, processor(s) 710 may further comprise a real-time camera engine, which may, without limitation, comprise a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, processor(s) 710 may further comprise a high dynamic range signal processor, which may, without limitation, comprise an image signal processor that is a hardware engine that is part of the camera processing pipeline.

[0107] In at least one embodiment, the processor(s) 710 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. In at least one embodiment, the video image compositor may perform lens distortion correction on the wide-angle camera(s) 770, the surround camera(s) 774, and / or the in-booth surveillance camera sensor(s). In at least one embodiment, the in-booth surveillance camera sensor(s) is / are preferably monitored by a neural network running on another instance of the SoC 704 and configured to detect and respond to events in the booth.In at least one embodiment, a system within the vehicle can perform lip reading without limitation to activate cellular service and place a call, dictate emails, change the destination, activate or change the vehicle's infotainment system and settings, or enable voice-activated web browsing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in an autonomous mode and are disabled otherwise.

[0108] In at least one embodiment, the video image compositor may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in a video, the noise reduction appropriately weights the spatial information and reduces the weight of information provided by neighboring frames. In at least one embodiment, where an image or a portion of an image has no motion, the temporal noise reduction performed by the video image compositor may use information from the previous frame to reduce noise in the current frame.

[0109] In at least one embodiment, the video image compositor may also be configured to perform stereo rectification on input stereo lens frames. In at least one embodiment, the video image compositor may also be used for user interface design when the operating system desktop is in use and the GPU(s) 708 are not needed to continuously render new surfaces. In at least one embodiment, when the GPU(s) 708 are turned on and actively performing 3D rendering, the video image compositor may be used to offload the GPU(s) 708 to improve performance and responsiveness.

[0110] In at least one embodiment, one or more of the SoC(s) 704 may further include a MIPI serial camera interface for receiving video and inputs from cameras, a high-speed interface, and / or a video input block that may be used for camera and related pixel input functions. In at least one embodiment, one or more of the SoC(s) 704 may further include one or more input / output controllers that may be controlled by software and may be used to receive I / O signals that are not associated with a specific role.

[0111] In at least one embodiment, one or more SoC(s) 704 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders ("codecs"), power management, and / or other devices. SoC(s) 704 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., LIDAR sensor(s) 764, RADAR sensor(s) 760, etc., which may be connected via Ethernet), data from bus 702 (e.g., speed of vehicle 700, steering wheel position, etc.), data from GNSS sensor(s) 758 (e.g., connected via Ethernet or CAN bus), etc.In at least one embodiment, one or more of the SoC(s) 704 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to offload routine data management tasks from the CPU(s) 706.

[0112] In at least one embodiment, the SoC(s) 704 may be an end-to-end platform with a flexible architecture spanning automation levels 3 through 5, thereby providing a comprehensive functional safety architecture, leveraging computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack along with deep learning tools. In at least one embodiment, the SoC(s) 704 may be faster, more reliable, and even more power and space efficient than conventional systems. For example, in at least one embodiment, the accelerator(s) 714, in combination with the CPU(s) 706, the GPU(s) 708, and the memory(s) 716, may form a fast, efficient platform for Level 3-5 autonomous vehicles.

[0113] In at least one embodiment, computer vision algorithms may be executed on CPUs, which may be configured using high-level programming language, such as C, to execute a variety of processing algorithms on a variety of visual data. However, in at least one embodiment, CPUs are often unable to meet the performance requirements of many image processing applications, such as execution time and power consumption requirements. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, which are used in in-vehicle ADAS applications and in practical Level 3-5 autonomous vehicles.

[0114] Embodiments as described herein enable the concurrent and / or sequential execution of multiple neural networks and the combination of the results to enable Level 3-5 autonomous driving capabilities. For example, in at least one embodiment, a CNN running on a DLA or a discrete GPU (e.g., GPU(s) 720) may include text and word recognition, enabling the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. In at least one embodiment, a DLA may further include a neural network capable of identifying, interpreting, and semantically understanding traffic signs, and passing this semantic understanding to path planning modules running on a CPU complex.

[0115] In at least one embodiment, multiple neural networks may be executed simultaneously, such as in Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign reading "Caution: Flashing lights indicate icing" along with an electric light may be interpreted independently or jointly by multiple neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate black ice" may be interpreted by a second deployed neural network, which informs the vehicle's path planning software (preferably running on a CPU complex) that, when flashing lights are detected, black ice is present.In at least one embodiment, the turn signal may be identified by operating a third neural network across multiple images, which informs the vehicle's path planning software of the presence (or absence) of turn signals. In at least one embodiment, all three neural networks may run concurrently, for example, within a DLA and / or on GPU(s) 708.

[0116] In at least one embodiment, a facial recognition and vehicle owner identification CNN may use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 700. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle when the owner approaches the driver's door and turn on the lights, and, in security mode, to disable the vehicle when the owner exits the vehicle. In this way, the SoC(s) 704 provide security against theft and / or carjacking.

[0117] In at least one embodiment, a CNN for detecting and identifying emergency vehicles may use data from microphones 796 to detect and identify emergency vehicle sirens. In at least one embodiment, the SoC(s) 704 use a CNN to classify environmental and urban noise, as well as visual data. In at least one embodiment, a CNN running on a DLA is trained to detect the relative approach speed of emergency vehicles (e.g., using the Doppler effect). In at least one embodiment, a CNN may also be trained to identify emergency vehicles specific to the local area in which the vehicle is traveling, as identified by GNSS sensor(s) 758.In at least one embodiment, when deployed in Europe, a CNN will attempt to detect European sirens, and when deployed in the United States, the CNN will attempt to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program may be used to execute an emergency vehicle safety routine, slow the vehicle, pull over to the side of the road, park the vehicle, and / or idle the vehicle using ultrasonic sensor(s) 762 until the emergency vehicle(s) pass by.

[0118] In at least one embodiment, the vehicle 700 may include one or more CPU(s) 718 (e.g., discrete CPU(s) or dCPU(s)) that may be connected to the SoC(s) 704 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU(s) 718 may include, for example, an x86 processor. CPU(s) 718 may be used to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and SoC(s) 704 and / or monitoring the status and health of the controller(s) 736 and / or an infotainment system on a chip (“infotainment SoC”) 730, for example.

[0119] In at least one embodiment, the vehicle 700 may include GPU(s) 720 (e.g., discrete GPU(s) or dGPU(s)) that may be coupled to the SoC(s) 704 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, GPU(s) 720 may provide additional artificial intelligence functionality, such as by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from sensors of the vehicle 700.

[0120] In at least one embodiment, the vehicle 700 may further include a network interface 724, which may include, without limitation, one or more wireless antennas 726 (e.g., one or more wireless antennas 726 for various communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, the network interface 724 may be used to enable a wireless connection via the internet to a cloud (e.g., to one or more servers and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). In at least one embodiment, a direct connection between the vehicle 80 and another vehicle and / or an indirect connection (e.g., via networks and the internet) may be established to communicate with other vehicles.In at least one embodiment, direct connections may be established via a vehicle-to-vehicle communication link. In at least one embodiment, the vehicle-to-vehicle communication link may provide information to the vehicle 700 about vehicles in the vicinity of the vehicle 700 (e.g., vehicles in front of, beside, and / or behind the vehicle 700). In at least one embodiment, the aforementioned functionality may be part of a cooperative adaptive cruise control function of the vehicle 700.

[0121] In at least one embodiment, the network interface 724 may include an SoC that provides modulation and demodulation functions and enables the controller(s) 736 to communicate over wireless networks. In at least one embodiment, the network interface 724 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. In at least one embodiment, the frequency conversions may be performed in any technically feasible manner. For example, frequency conversions may be performed by known methods and / or using superheterodyne techniques. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip.In at least one embodiment, the network interface may include wireless functionality for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0122] In at least one embodiment, the vehicle 700 may further include one or more data stores 728, which may include, without limitation, off-chip memory (e.g., off-SoC(s) 704). In at least one embodiment, the data store(s) 728 may include, without limitation, one or more memory elements, including RAM, SRAM, dynamic random access memory ("DRAM"), video random access memory ("VRAM"), flash, hard drives, and / or other components and / or devices capable of storing at least one bit of data.

[0123] In at least one embodiment, the vehicle 700 may further include GNSS sensor(s) 758 (e.g., GPS and / or assisted GPS sensors) to assist with mapping, sensing, occupancy grid generation, and / or path planning. In at least one embodiment, any number of GNSS sensor(s) 758 may be used, including, for example and without limitation, a GPS using a USB port with an Ethernet-to-serial bridge (e.g., RS-232).

[0124] In at least one embodiment, vehicle 700 may further include RADAR sensor(s) 760. The RADAR sensor(s) 760 may be used by a vehicle 700 for vehicle detection over long distances, even in darkness and / or adverse weather conditions. In at least one embodiment, the RADAR functional safety levels may be ASIL B. The RADAR sensor(s) 760 may use CAN and / or bus 702 (e.g., to transmit the data generated by the RADAR sensor(s) 760) for control and access to object tracking data, with some examples accessing raw data over an Ethernet. In at least one embodiment, a wide range of RADAR sensor types may be used. For example, and without limitation, RADAR sensor(s) 760 may be suitable for use with front, rear, and side RADAR.In at least one embodiment, one or more of the RADAR sensors 760 are pulse Doppler RADAR sensors.

[0125] In at least one embodiment, the RADAR sensor(s) 760 may have various configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, etc. In at least one embodiment, the long range RADAR may be used for adaptive cruise control. In at least one embodiment, long range RADAR systems may provide a wide field of view, realized by two or more independent scans, e.g., within a 250 m range. In at least one embodiment, the RADAR sensor(s) 760 may help distinguish between stationary and moving objects and may be used by the ADAS system 738 for emergency braking assistance and forward collision warning.In at least one embodiment, the sensor(s) 760 included in a long-range radar system may comprise, without limitation, a monostatic multimodal radar with multiple (e.g., six or more) fixed radar antennas and a high-speed CAN and FlexRay interface. In at least one embodiment with six antennas, four antennas in the center may create a focused beam pattern used to detect the vehicle's surroundings at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, the other two antennas may expand the field of view so that vehicles entering or exiting the lane of vehicle 700 can be quickly detected.

[0126] For example, in at least one embodiment, medium-range radar systems may have a range of up to 160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, short-range radar systems may include, without limitation, any number of radar sensors 760, which may be installed at either end of the rear bumper. In at least one embodiment, a radar sensor system, when installed at either end of the rear bumper, may create two beams that continuously monitor the blind spot to the rear and side of the vehicle. In at least one embodiment, short-range radar systems may be used in the ADAS system 738 for blind spot detection and / or lane change assistance.

[0127] In at least one embodiment, the vehicle 700 may further include ultrasonic sensor(s) 762. In at least one embodiment, the ultrasonic sensor(s) 762, which may be located at the front, rear, and / or sides of the vehicle 700, may be used for parking assistance and / or for creating and updating an occupancy grid. In at least one embodiment, a plurality of ultrasonic sensors 762 may be used, and different ultrasonic sensors 762 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor(s) 762 may operate at functional safety levels of ASIL B.

[0128] In at least one embodiment, the vehicle 700 may include LIDAR sensor(s) 764. The LIDAR sensor(s) 764 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor(s) 764 may have an ASIL B functional safety rating. In at least one embodiment, the vehicle 700 may include multiple LIDAR sensors 764 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0129] In at least one embodiment, the LIDAR sensor(s) 764 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, the commercially available LIDAR sensor(s) 764 may have an indicated range of approximately 100 m, with an accuracy of 2 cm to 3 cm, and with support for a 100 Mbps Ethernet connection, for example. In at least one embodiment, one or more non-prominent LIDAR sensors 764 may be used. In such an embodiment, the LIDAR sensor(s) 764 may be implemented as a small device that may be embedded in the front, rear, sides, and / or corners of the vehicle 700.In at least one embodiment, the LIDAR sensor(s) 764 in such an embodiment may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees with a range of 200 m, even for low-reflectivity objects. In at least one embodiment, the front-mounted LIDAR sensor(s) 764 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0130] In at least one embodiment, LIDAR technologies, such as 3D Flash LIDAR, may also be used. 3D Flash LIDAR uses a laser flash as the transmission source to illuminate the surroundings of the vehicle 700 up to a distance of approximately 200 m. In at least one embodiment, a Flash LIDAR unit includes, without limitation, a receptor that records the time of flight of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance of the vehicle 700 from objects. In at least one embodiment, Flash LIDAR may enable highly accurate and distortion-free images of the surroundings to be generated with each laser flash. In at least one embodiment, four Flash LIDAR sensors may be deployed, one on each side of the vehicle 700.In at least one embodiment, 3D flash LIDAR systems include, without limitation, a solid-state 3D star array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, the flash LIDAR device may use a 5-nanosecond Class I (eye-safe) laser pulse per image and collect the reflected laser light as 3D range point clouds and co-registered intensity data.

[0131] In at least one embodiment, the vehicle may further include one or more IMU sensors 766. In at least one embodiment, the IMU sensor(s) 766 may be located at the center of the rear axle of the vehicle 700. In at least one embodiment, the IMU sensor(s) 766 may include, for example, and without limitation, one or more accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other types of sensors. In at least one embodiment, such as in six-axis applications, the IMU sensor(s) 766 may include, without limitation, accelerometers and gyroscopes. In at least one embodiment, such as in nine-axis applications, the IMU sensor(s) 766 may include, without limitation, accelerometers, gyroscopes, and magnetometers.

[0132] In at least one embodiment, the IMU sensor(s) 766 may be implemented as a miniaturized, high-performance GPS-based inertial navigation system ("GPS / INS") that combines microelectromechanical systems ("MEMS") inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor(s) 766 may enable the vehicle 700 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS with the IMU sensor(s) 766. In at least one embodiment, the IMU sensor(s) 766 and GNSS sensor(s) 758 may be combined into a single integrated unit.

[0133] In at least one embodiment, the vehicle 700 may include one or more microphones 796 disposed in and / or around the vehicle 700. In at least one embodiment, the microphone(s) 796 may be used, among other things, for detecting and identifying emergency vehicles.

[0134] In at least one embodiment, vehicle 700 may further include any number of camera types, including stereo camera(s) 768, wide-angle camera(s) 770, infrared camera(s) 772, surround camera(s) 774, long-range camera(s) 798, mid-range camera(s) 776, and / or other camera types. In at least one embodiment, cameras may be used to capture image data around the entire perimeter of vehicle 700. In at least one embodiment, the types of cameras used depend on the vehicle 700. In at least one embodiment, any combination of camera types may be used to provide the required coverage around vehicle 700. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, vehicle 700 may include six, seven, ten, twelve, or another number of cameras.In at least one embodiment, the cameras may support, for example and without limitation, Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet. In at least one embodiment, each of the cameras is previously described herein with reference to . Fig. 7A and Fig. 7B is described in more detail.

[0135] In at least one embodiment, the vehicle 700 may further include one or more vibration sensors 742. In at least one embodiment, the vibration sensor(s) 742 may measure vibrations of components of the vehicle 700, such as the axle(s). For example, in at least one embodiment, changes in the vibrations may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 742 are used, differences between the vibrations may be used to determine friction or slippage of the road surface (e.g., when the difference in vibrations is between a driven axle and a free-spinning axle).

[0136] In at least one embodiment, vehicle 700 may include an ADAS system 738. The ADAS system 738 may, in some examples, include, without limitation, an SoC. In at least one embodiment, the ADAS system 738 may include, without limitation, any number and combination of an autonomous / adaptive / automatic cruise control (“ACC”) system, a cooperative adaptive cruise control (“CACC”) system, a forward crash warning (“FCW”) system, an automatic emergency braking (“AEB”) system, a lane departure warning (“LDW”) system, a lane keep assist (“LKA”) system, a blind spot warning (“BSW”) system, a rear cross traffic warning (“RCTW”) system, a forward collision warning (“CW”) system, a lane centering (“LC”) system, and / or other systems, features, and / or functions.

[0137] In at least one embodiment, the ACC system may utilize RADAR sensor(s) 760, LIDAR sensor(s) 764, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to the vehicle immediately in front of the vehicle 700 and automatically adjusts the speed of the vehicle 700 to maintain a safe distance from vehicles ahead. In at least one embodiment, the lateral ACC system provides follow-through and advises the vehicle 700 to change lanes if necessary. In at least one embodiment, the lateral ACC system is connected to other ADAS applications such as LC and CW.

[0138] In at least one embodiment, the CACC system utilizes information from other vehicles that may be received via the network interface 724 and / or the radio antenna(s) 726 from other vehicles over a wireless connection or indirectly via a network connection (e.g., over the Internet). In at least one embodiment, direct connections may be provided through a vehicle-to-vehicle ("V2V") communication link, while indirect connections may be provided through an infrastructure-to-vehicle ("I2V") communication link. In general, the V2V communication concept provides information about immediately ahead vehicles (e.g., vehicles immediately in front of and in the same lane as vehicle 700), while the I2V communication concept provides information about traffic further ahead.In at least one embodiment, the CACC system may include either one or both I2V and V2V information sources. In at least one embodiment, the CACC system may be more reliable given information about vehicles ahead of vehicle 700, and it has the potential to improve traffic flow and reduce congestion on the road.

[0139] In at least one embodiment, the FCW system is designed to warn the driver of a hazard so they can take corrective action. In at least one embodiment, the FCW system uses a forward-facing camera and / or RADAR sensor(s) 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, e.g., a display, speaker, and / or vibrating component. In at least one embodiment, the FCW system can provide a warning, e.g., in the form of a sound, a visual warning, a vibration, and / or a rapid braking pulse.

[0140] In at least one embodiment, the AEB system detects an impending forward collision with another vehicle or other object and may automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may utilize forward-facing camera(s) and / or radar sensor(s) 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, upon detecting a hazard, the AEB system typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision.In at least one embodiment, the AEB system may include techniques such as dynamic brake assist and / or crash imminent braking.

[0141] In at least one embodiment, the LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 700 crosses lane markings. In at least one embodiment, the LDW system is not activated if the driver indicates an intentional lane departure by activating a turn signal. In at least one embodiment, the LDW system may utilize forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to feedback to the driver, such as a display, speaker, and / or vibrating component. In at least one embodiment, the LKA system is a variation of the LDW system. The LKA system provides steering intervention or braking to correct the vehicle 700 when the vehicle 700 begins to depart from the lane.

[0142] In at least one embodiment, the BSW system detects and warns the driver of vehicles located in the vehicle's blind spot. In at least one embodiment, the BSW system may provide a visual, audible, and / or tactile warning to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system may provide an additional warning when the driver activates a turn signal. In at least one embodiment, the BSW system may utilize rear-facing camera(s) and / or RADAR sensor(s) 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0143] In at least one embodiment, the RCTW system may provide visual, audible, and / or tactile notification when an object is detected outside the range of the rearview camera when the vehicle 700 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a crash. In at least one embodiment, the RCTW system may utilize one or more rear-facing RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0144] In at least one embodiment, conventional ADAS systems may be prone to false positive results, which can be annoying and distracting for the driver, but are typically not catastrophic because conventional ADAS systems warn the driver and give them the opportunity to decide whether a safety condition truly exists and act accordingly. In at least one embodiment, in the event of conflicting results, the vehicle 700 itself decides whether to consider the result of a primary computer or a secondary computer (e.g., the first controller 736 or the second controller 736). In at least one embodiment, the ADAS system 738 may, for example, be a backup and / or secondary computer that provides perception information to a rationality module of the backup computer.In at least one embodiment, a backup computer rationality monitor may execute redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. In at least one embodiment, the outputs of the ADAS system 738 may be forwarded to a higher-level MCU. In at least one embodiment, if conflicts occur between the outputs of the primary computer and the secondary computer, the monitoring MCU determines how to resolve the conflict to ensure safe operation.

[0145] In at least one embodiment, the primary computer may be configured to provide the supervising MCU with a confidence score indicating the primary computer's confidence in the selected outcome. In at least one embodiment, the supervising MCU may follow the primary computer's instruction if the confidence score exceeds a threshold, regardless of whether the secondary computer provides a conflicting or inconsistent outcome. In at least one embodiment, where the confidence score does not meet the threshold and the primary and secondary computers indicate different outcomes (e.g., a conflict), the supervising MCU may arbitrate between the computers to determine the appropriate outcome.

[0146] In at least one embodiment, the monitoring MCU may be configured to execute a neural network(s) that is / are trained and configured to determine, based at least in part on the outputs of the primary computer and the secondary computer, the conditions under which the secondary computer triggers false alarms. In at least one embodiment, the neural network(s) in the monitoring MCU may learn when the output of the secondary computer can and cannot be trusted. For example, in at least one embodiment, if the secondary computer is a RADAR-based FCW system, a neural network in the monitoring MCU may learn when the FCW system identifies metallic objects that are not actually hazards, such as a drain grate or manhole cover, that trigger an alarm.In at least one embodiment, when the secondary computer is a camera-based LDW system, a neural network in the monitoring MCU can learn to override the LDW system when cyclists or pedestrians are present and lane departure is actually the safest maneuver. In at least one embodiment, the monitoring MCU can include a DLA or GPU suitable for executing neural networks with associated memory. In at least one embodiment, the monitoring MCU can comprise and / or be included as a component of the SoC(s) 704.

[0147] In at least one embodiment, the ADAS system 738 may include a secondary computer that executes the ADAS functionality using conventional computer vision rules. In at least one embodiment, the secondary computer may use classic computer vision (if-then) rules, and the presence of a neural network(s) in the parent MCU may improve reliability, safety, and performance. In at least one embodiment, the different implementations and intentional non-identity make the overall system more fault-tolerant, particularly against errors caused by software features (or software-hardware interfaces).For example, in at least one embodiment, if there is a software bug or error in the software running on the primary computer, and if non-identical software code running on the secondary computer produces the same overall result, then the monitoring MCU may have greater confidence that the overall result is correct and the bug in the software or hardware on the primary computer does not cause a significant failure.

[0148] In at least one embodiment, the output of the ADAS system 738 may be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 738 displays a forward collision warning due to an object immediately ahead, the perception block may use this information in identifying objects. In at least one embodiment, the secondary computer may have its own neural network trained to reduce the risk of false alarms, as described herein.

[0149] In at least one embodiment, the vehicle 700 may further include an infotainment SoC 730 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, in at least one embodiment, the infotainment system 730 may not be an SoC and may include, without limitation, two or more discrete components. In at least one embodiment, the infotainment SoC 730 may include, without limitation, a combination of hardware and software that may be used to provide audio (e.g., music, a personal digital assistant, navigation directions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g.,Navigation systems, rear parking sensors, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / close, air filter information, etc.) for the vehicle 700. The infotainment SoC 730 may include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, WiFi, steering wheel audio controls, a hands-free system, a heads-up display ("HUD"), an HMI display 734, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 730 may further be used to provide information (e.g., visual and / or audible) to the user(s) of the vehicle, such as:Information from the ADAS system 738, autonomous driving information such as planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0150] In at least one embodiment, the infotainment SoC 730 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 730 may communicate with other devices, systems, and / or components of the vehicle 700 via the bus 702 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, the infotainment SoC 730 may be coupled to a supervisory MCU so that the infotainment system's GPU may perform some self-driving functions in the event that the primary controller(s) 736 (e.g., primary and / or backup computer of the vehicle 700) fails. In at least one embodiment, the infotainment SoC 730 may place the vehicle 700 into a chauffeur-to-safe-stop mode, as described herein.

[0151] In at least one embodiment, the vehicle 700 may further include an instrument cluster 732 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, the instrument cluster 732 may include, without limitation, a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). In at least one embodiment, the instrument cluster 732 may include, without limitation, any number and combination of instruments, such as, but not limited to, a speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine trouble light(s), supplemental restraint system information (e.g., airbags), lighting controls, safety system controls, navigation information, etc.In some examples, the information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. In at least one embodiment, the instrument cluster 732 may include a portion of the infotainment SoC 730, or vice versa.

[0152] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0153] Fig. 7D is a diagram of a system 776 for communication between the cloud-based server(s) and the autonomous vehicle 700 of Fig. 7A, according to at least one embodiment. In at least one embodiment, the system 776 may include, without limitation, the server(s) 778, the network(s) 790, and any number and type of vehicles, including the vehicle 700. The server(s) 778 may include, without limitation, a plurality of GPUs 784(A)-784(H) (collectively referred to herein as GPUs 784), PCIe switches 782(A)-782(H) (collectively referred to herein as PCIe switches 782), and / or CPUs 780(A)-780(B) (collectively referred to herein as CPUs 780). GPUs 784, CPUs 780, and PCIe switches 782 may be interconnected via high-speed links, such as 10 / 100 / 1000 Mbit / s. For example, and without limitation, via NVIDIA-developed NVLink interfaces 788 and / or PCIe connections 786. In at least one embodiment, the GPUs 784 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 784 and PCIe switches 782 are connected via PCIe connections.In at least one embodiment, while eight GPUs 784, two CPUs 780, and four PCIe switches 782 are illustrated, this is not intended to be limiting. In at least one embodiment, each of the servers 778 may include, without limitation, any number of GPUs 784, CPUs 780, and / or PCIe switches 782 in any combination. For example, in at least one embodiment, the server(s) 778 may each include eight, sixteen, thirty-two, and / or more GPUs 784.

[0154] In at least one embodiment, the server(s) 778 may receive, via the network(s) 790 and from vehicles, image data representative of images depicting unexpected or changed road conditions, such as recently commenced roadwork. In at least one embodiment, the server(s) 778 may transmit, via the network(s) 790 and to vehicles, neural networks 792, updated neural networks 792, and / or map information 794 including, without limitation, information about traffic and road conditions. In at least one embodiment, the updates to the map information 794 may include, without limitation, updates to the HD map 722, such as information about construction, potholes, detours, flooding, and / or other obstacles.In at least one embodiment, neural networks 792, updated neural networks 792, and / or map information 794 may result from new training and / or experience represented in data received from any number of vehicles in the environment and / or may be based at least in part on training performed in a data center (e.g., using server(s) 778 and / or other servers).

[0155] In at least one embodiment, the server(s) 778 may be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, the training data may be generated from vehicles and / or in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., if the associated neural network benefits from supervised learning) and / or subjected to other preprocessing. In at least one embodiment, any amount of training data is unlabeled and / or preprocessed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once machine learning models are trained, vehicle machine learning models may be used (e.g.,Transmission to vehicles via network(s) 790, and / or machine learning models may be used by server(s) 778 to remotely monitor vehicles.

[0156] In at least one embodiment, the server(s) 778 may receive data from vehicles and apply data to real-time, state-of-the-art neural networks for intelligent inferencing in real time. In at least one embodiment, the server(s) 778 may include deep learning supercomputers and / or dedicated AI computers powered by GPU(s) 784, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, the server(s) 778 may include a deep learning infrastructure using CPU-powered data centers.

[0157] In at least one embodiment, the deep learning infrastructure of server(s) 778 may be capable of rapid, real-time inferencing and may utilize this capability to assess and verify the health of processors, software, and / or associated hardware in the vehicle 700. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from the vehicle 700, such as an image sequence and / or objects that the vehicle 700 has located in that image sequence (e.g., via computer vision and / or other machine object classification techniques).In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to the objects identified by the vehicle 700, and if the results do not match and the deep learning infrastructure concludes that the AI ​​in the vehicle 700 is malfunctioning, the server(s) 778 may send a signal to the vehicle 700 instructing a fail-safe computer of the vehicle 700 to take over control, notify the passengers, and perform a safe parking maneuver.

[0158] In at least one embodiment, the server(s) 778 may include GPU(s) 784 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). In at least one embodiment, the combination of GPU-driven servers and inference acceleration may enable real-time responsiveness. In at least one embodiment, e.g., where performance is less critical, servers with CPUs, FPGAs, and other processors may also be used for inferencing. COMPUTER SYSTEMS

[0159] Fig. 8 is a block diagram illustrating an example computer system, which may be a system with interconnected devices and components, a system-on-a-chip (SOC), or a combination thereof 800, including a processor including execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, computer system 800 may include, without limitation, a component, such as a processor 802, for employing execution units including logic for performing algorithms for processing data in accordance with the present disclosure, such as in the embodiment described herein. In at least one embodiment, computer system 800 may include processors, such asthe PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 800 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0160] Embodiments may also be used in other implementations such as handheld devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0161] In at least one embodiment, computer system 800 may include, without limitation, a processor 802, which may include, without limitation, one or more execution units 808 to perform training of a machine learning and / or inferencing model according to the techniques described herein. In at least one embodiment, system 800 is a single-processor desktop or server system, but in another embodiment, system 800 may be a multiprocessor system. In at least one embodiment, processor 802 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other device, such as a digital signal processor.In at least one embodiment, the processor 802 may be connected to a processor bus 810 that may transmit data signals between the processor 802 and other components in the computer system 800.

[0162] In at least one embodiment, processor 802 may include, without limitation, an internal Level 1 ("L1") cache ("cache") 804. In at least one embodiment, processor 802 may include a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache may be external to processor 802. Other embodiments may include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, register file 806 may store various data types in various registers, including, without limitation, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0163] In at least one embodiment, execution unit 808, which includes, without limitation, logic for performing integer and floating-point operations, is also located in processor 802. In at least one embodiment, processor 802 may also include microcode read-only memory ("ROM") ("ucode") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 808 may include logic for handling a packed instruction set 809. In at least one embodiment, by providing a packed instruction set 809 in an instruction set of a general-purpose processor 802, along with associated instruction execution circuitry, the operations used by many multimedia applications can be performed using packed data in a general-purpose processor 802.In one or more embodiments, many multimedia applications may be accelerated and run more efficiently by utilizing the full width of a processor's data bus to perform operations on packed data, eliminating the need to transfer smaller units of data across the processor's data bus to perform one or more operations on one data item at a time.

[0164] In at least one embodiment, execution unit 808 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 800 may include, without limitation, a memory 820. In at least one embodiment, memory 820 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 820 may store instruction(s) 819 and / or data 821 represented by data signals that may be executed by processor 802.

[0165] In at least one embodiment, the system logic chip may be connected to the processor bus 810 and the memory 820. In at least one embodiment, the system logic chip may include, without limitation, a memory control hub ("MCH") 816, and the processor 802 may communicate with the MCH 816 via the processor bus 810. In at least one embodiment, the MCH 816 may provide a high-bandwidth memory path 818 to the memory 820 for instruction and data storage, as well as for the storage of graphics instructions, data, and textures. In at least one embodiment, the MCH 816 may route data signals between the processor 802, the memory 820, and other components in the computer system 800, and may bridge data signals between the processor bus 810, the memory 820, and a system I / O 822. In at least one embodiment, the system logic chip may provide a graphics port for connection to a graphics controller.In at least one embodiment, the MCH 816 may be coupled to the memory 820 via a high-bandwidth memory path 818, and the graphics video card 812 may be coupled to the MCH 816 via an AGP connection 814.

[0166] In at least one embodiment, computer system 800 may use a system I / O bus 822, which is a proprietary hub interface bus, to connect MCH 816 to I / O controller hub ("ICH") 830. In at least one embodiment, ICH 830 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 820, chipset, and processor 802. Examples may include, but are not limited to, an audio controller 829, a firmware hub (“Flash BIOS”) 828, a wireless transceiver 826, a data store 824, a legacy I / O controller 823 with user input and keyboard interfaces, a serial expansion port 827, such as Universal Serial Bus (“USB”), and a network controller 834.In at least one embodiment, data storage 824 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0167] In at least one embodiment, Fig. 8 a system comprising interconnected hardware devices or ‘chips’, while in other embodiments Fig. 8 may show an exemplary system on a chip (“SoC”). In at least one embodiment, the Fig. 8 may be interconnected using proprietary connections, standardized connections (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of system 800 are interconnected via Compute Express Link (CXL) connections.

[0168] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0169] Fig. 9 is a block diagram illustrating an electronic device 900 for using a processor 910 according to at least one embodiment. In at least one embodiment, the electronic device 900 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0170] In at least one embodiment, system 900 may include, without limitation, a processor 910 communicatively coupled to any number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 910 is coupled via a bus or interface, such as a 1°C bus, a system management bus ("SMBus"), a low pin count bus (LPC), a serial peripheral interface ("SPI"), a high-definition audio bus ("HDA"), a serial advance technology attachment bus ("SATA"), a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter bus ("UART"). In at least one embodiment, Fig. 9 a system comprising interconnected hardware devices or ‘chips’, while in other embodiments Fig. 9 may show an exemplary system on a chip (“SoC”). In at least one embodiment, the Fig. 9 may be interconnected with proprietary connections, standardized connections (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 9 are connected via Compute Express Link (CXL) connections.

[0171] In at least one embodiment, Fig. 9 a display 924, a touchscreen 925, a touchpad 930, a Near Field Communications unit (“NFC”) 945, a sensor hub 940, a thermal sensor 939, an Express Chipset (“EC”) 935, a Trusted Platform Module (“TPM”) 938, BIOS / Firmware / Flash Memory (“BIOS, FW Flash”) 922, a DSP 960, a drive (“SSD or HDD”) 920 such as a Solid State Disk (“SSD”) or a Hard Drive (“HDD”), a wireless local area network unit (“WLAN”) 950, a Bluetooth unit 952, a wireless wide area network unit (“WWAN”) 956, a Global Positioning System (GPS) 955, a camera (“USB 3. 0 Camera”) 954, such as a Such components may include a USB 3.0 camera, or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3"), e.g., implemented in the LPDDR3 standard. These components may be implemented in any suitable manner.

[0172] In at least one embodiment, other components may be communicatively coupled to processor 910 via the components described above. In at least one embodiment, an accelerometer 941, an ambient light sensor ("ALS") 942, a compass 943, and a gyroscope 944 may be communicatively coupled to sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 936, and a touchpad 930 may be communicatively coupled to EC 935. In at least one embodiment, speaker 963, headphones 964, and a microphone ("mic") 965 may be communicatively coupled to an audio unit ("audio codec and class d amp") 964, which in turn may be communicatively coupled to DSP 960. In at least one embodiment, the audio unit 964 may include, for example and without limitation, an audio encoder / decoder (“codec”) and a Class D amplifier.In at least one embodiment, the SIM card ("SIM") 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and the Bluetooth unit 952, as well as the WWAN unit 956, may be implemented in a Next Generation Form Factor ("NGFF").

[0173] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0174] Fig. 10 illustrates a computer system 1000 according to at least one embodiment. In at least one embodiment, the computer system 1000 is configured to implement various processes and methods described in this disclosure.

[0175] In at least one embodiment, computer system 1000 includes, without limitation, at least one central processing unit ("CPU") 1002 coupled to a communications bus 1010 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or another bus or point-to-point communications protocol. In at least one embodiment, computer system 1000 includes, without limitation, main memory 1004 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 1004, which may take the form of random access memory ("RAM").In at least one embodiment, a network interface subsystem ("network interface") 1022 provides an interface to other computing devices and networks to receive data from the computer system 1000 and to transmit data to other systems.

[0176] In at least one embodiment, computer system 1000 includes, without limitation, input devices 1008, a parallel processing system 1012, and display devices 1006, which may be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from input devices 1008 such as a keyboard, mouse, touchpad, microphone, and others. In at least one embodiment, each of the foregoing modules may be arranged on a single semiconductor platform to form a processing system.

[0177] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0178] Fig. 11 illustrates a computer system 1100 according to at least one embodiment. In at least one embodiment, the computer system 1100 includes, without limitation, a computer 1110 and a USB flash drive 1120. In at least one embodiment, the computer 1110 may include, without limitation, any number and type of processor(s) (not shown) and memory (not shown). In at least one embodiment, the computer 1110 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.

[0179] In at least one embodiment, USB flash drive 1120 includes, without limitation, a processing unit 1130, a USB interface 1140, and USB interface logic 1150. In at least one embodiment, processing unit 1130 may be any instruction execution system, device, or apparatus capable of executing instructions. In at least one embodiment, processing unit 1130 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, processing core 1130 comprises an application-specific integrated circuit ("ASIC") optimized to perform any number and type of machine learning-related operations.For example, in at least one embodiment, processing core 1130 is a tensor processing unit ("TPC") optimized for performing machine learning inference operations. In at least one embodiment, processing core 1130 is a vision processing unit ("VPU") optimized for performing vision and machine learning operations.

[0180] In at least one embodiment, USB interface 1140 may be any type of USB plug or receptacle. For example, in at least one embodiment, USB interface 1140 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1140 is a USB 3.0 Type-A plug. In at least one embodiment, USB interface logic 1150 may include any amount and type of logic that enables processing unit 1130 to connect to a device (e.g., computer 1110) via USB port 1140.

[0181] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0182] Fig. 12A illustrates an exemplary architecture in which a plurality of GPUs 1210-1213 are communicatively coupled to a plurality of multi-core processors 1205-1206 via high-speed interconnects 1240-1243 (e.g., buses, point-to-point links, etc.). In one embodiment, the high-speed interconnects 1240-1243 support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or more. Various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0.

[0183] Additionally, and in one embodiment, two or more GPUs 1210-1213 are interconnected via high-speed interconnects 1229-1230, which may be implemented with the same or different protocols / connections than those used for high-speed interconnects 1240-1243. Similarly, two or more multi-core processors 1205-1206 may be interconnected via high-speed interconnects 1228, which may be symmetric multiprocessor (SMP) buses operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various Fig. 12A shown system components via the same protocols / lines (e.g. via a common connection structure).

[0184] In one embodiment, each multi-core processor 1205-1206 is communicatively connected to a processor memory 1201-1202 via memory interconnects 1226-1227, and each graphics processor 1210-1213 is communicatively connected to graphics processor memory 1220-1223 via graphics processor memory interconnects 1250-1253. Memory interconnects 1226-1227 and 1250-1253 may use the same or different memory access technologies. For example, processor memories 1201-1202 and GPU memories 1220-1223 may include volatile memories such as dynamic random access memories (DRAMs) (including stacked DRAMs), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-RAM. In one embodiment, a portion of processor memories 1201-1202 may be volatile memory and another portion may be non-volatile memory (e.g.,using a two-level memory hierarchy (2LM).

[0185] As described herein, while different processors 1205-1206 and GPUs 1210-1213 may be physically connected to a particular memory 1201-1202 and 1220-1223, respectively, a unified memory architecture may be implemented in which the same virtual system address space (also referred to as "effective address space") is distributed across different physical memories. For example, processor memories 1201-1202 may each comprise 64 GB of system address space, and GPU memories 1220-1223 may each comprise 32 GB of system address space (resulting in a total addressable memory of 256 GB in this example).

[0186] Fig. 12B shows additional details for a connection between a multi-core processor 1207 and a graphics acceleration module 1246 according to an exemplary embodiment. The graphics acceleration module 1246 may include one or more GPU chips integrated on a line card connected to the processor 1207 via a high-speed interconnect 1240. Alternatively, the graphics acceleration module 1246 may be integrated on the same package or chip as the processor 1207.

[0187] In at least one embodiment, the illustrated processor 1207 includes a plurality of cores 1260A-1260D, each with a translation lookaside buffer 1261A-1261D and one or more caches 1262A-1262D. In at least one embodiment, the cores 1260A-1260D may include various other components for executing instructions and processing data that are not illustrated. The caches 1262A-1262D may include Level 1 (L1) and Level 2 (L2) caches. Additionally, one or more shared caches 1256 may be present within the caches 1262A-1262D, which are shared by groups of cores 1260A-1260D. For example, one embodiment of processor 1207 has 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared between two adjacent cores.The processor 1207 and the graphics acceleration module 1246 are connected to the system memory 1214, which includes the processor memories 1201-1202 of . Fig. 12A.

[0188] The coherence of data and instructions stored in various caches 1262A-1262D, 1256, and system memory 1214 is maintained by inter-core communication over a coherency bus 1264. For example, each cache may have cache coherency logic / circuitry connected to it to communicate over the coherency bus 1264 in response to detected reads or writes to specific cache lines. In one implementation, a cache snooping protocol is implemented over the coherency bus 1264 to snoop cache accesses.

[0189] In one embodiment, a proxy circuit 1225 communicatively couples the graphics acceleration module 1246 to the coherence bus 1264 so that the graphics acceleration module 1246 can participate in a cache coherence protocol as a peer of the cores 1260A-1260D. An interface 1235 provides connectivity to the proxy circuit 1225 via the high-speed interconnect 1240 (e.g., a PCIe bus, NVLink, etc.), and an interface 1237 connects the graphics acceleration module 1246 to the interconnect 1240.

[0190] In one implementation, an accelerator integration circuit 1236 provides cache management, memory access, context management, and interrupt management services on behalf of a plurality of graphics processing engines 1231, 1232, N of the graphics acceleration module 1246. The graphics processing engines 1231, 1232, N may each comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1231, 1232, N may comprise various types of graphics processing engines within a graphics processor, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit modules.In at least one embodiment, the graphics acceleration module 1246 may be a GPU having a plurality of graphics processing units 1231-1232, N, or the graphics processing units 1231-1232, N may be individual GPUs integrated into a common package, line card, or chip.

[0191] In one embodiment, accelerator integration circuit 1236 includes a memory management unit (MMU) 1239 to perform various memory management functions such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing system memory 1214. MMU 1239 may also include a translation lookaside buffer (TLB) (not shown) to cache virtual / effective to physical / real address translations. In one embodiment, a cache 1238 stores instructions and data for efficient access by graphics processors 1231-1232, N. In one embodiment, the data stored in cache 1238 and graphics memories 1233-1234, M is kept coherent with core caches 1262A-1262D, 1256 and system memory 1214.As previously mentioned, this may be done via a proxy circuit 1225 on behalf of the cache 1238 and memories 1233-1234, M (e.g., sending updates to the cache 1238 related to changes / accesses to cache lines in the processor caches 1262A-1262D, 1256 and receiving updates from the cache 1238).

[0192] A set of registers 1245 stores context data for threads executed by graphics processing engines 1231-1232, N, and a context management circuit 1248 manages thread contexts. For example, context management circuit 1248 may perform save and restore operations to save and restore contexts of different threads during context switches (e.g., when a first thread is saved and a second thread is saved so that a second thread can be executed by a graphics processing engine). During a context switch, for example, context management circuit 1248 may store current register values ​​in a specific area in memory (e.g., identified by a context pointer). The register values ​​may then be restored upon return to a context.In one embodiment, an interrupt management circuit 1247 receives and processes interrupts received from system devices.

[0193] In one implementation, virtual / effective addresses from a graphics processing engine 1231 are translated by the MMU 1239 into real / physical addresses in system memory 1214. One embodiment of the accelerator integration circuit 1236 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 1246 and / or other accelerator devices. The graphics accelerator module 1246 may be dedicated to a single application executing on the processor 1207, or it may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which the resources of the graphics processors 1231-1232, N are shared among multiple applications or virtual machines (VMs).In at least one embodiment, the resources may be divided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.

[0194] In at least one embodiment, an accelerator integration circuit 1236 acts as a bridge to a system for the graphics acceleration module 1246 and provides address translation and system memory caching services. Furthermore, the accelerator integration circuit 1236 may provide virtualization functions to a host processor to manage the virtualization of the graphics processing modules 1231-1232, interrupts, and memory management.

[0195] Because the hardware resources of graphics processors 1231-1232, N are explicitly mapped to a real address space seen by host processor 1207, each host processor can directly address these resources with an effective address value. One function of accelerator integration circuit 1236, in one embodiment, is to physically separate graphics processing engines 1231-1232, N so that they appear to a system as independent entities.

[0196] In at least one embodiment, one or more graphics memories 1233-1234, M are coupled to each of the graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M store instructions and data processed by each of the graphics processing engines 1231-1232, N. The graphics memories 1233-1234, M may include volatile memories such as DRAMs (including stacked DRAMs), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories such as 3D XPoint or Nano-Ram.

[0197] In one embodiment, to reduce data traffic over interconnect 1240, biasing techniques are used to ensure that the data stored in graphics memories 1233-1234, M is data most frequently used by graphics processing engines 1231-1232, N and preferably not used (at least not frequently) by cores 1260A-1260D. Similarly, a biasing mechanism attempts to keep data needed by cores (and preferably not by graphics processing engines 1231-1232, N) in core caches 1262A-1262D, 1256, and in system memory 1214.

[0198] Fig. 12C shows another exemplary embodiment in which the accelerator integration circuit 1236 is integrated into the processor 1207. In this embodiment, the graphics processors 1231-1232, N communicate directly over the high-speed interconnect 1240 with the accelerator integration circuit 1236 via the interface 1237 and the interface 1235 (which, again, may use any form of bus or interface protocol). The accelerator integration circuit 1236 may perform the same operations as in Fig. 12B, but possibly with higher throughput due to its close proximity to the coherence bus 1264 and caches 1262A-1262D, 1256. One embodiment supports various programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and shared programming models (with virtualization), which may include programming models controlled by the accelerator integration circuit 1236 and programming models controlled by the graphics acceleration module 1246.

[0199] In at least one embodiment, the graphics processing engines 1231-1232, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can forward other application requests to the graphics processing engines 1231-1232, N, thereby enabling virtualization within a VM / partition.

[0200] In at least one embodiment, the graphics processing engines 1231-1232, N may be shared between multiple VM / application partitions. In at least one embodiment, shared models may use a system hypervisor to virtualize the graphics processing engines 1231-1232, N and allow access by any operating system. For single-partition systems without a hypervisor, the graphics processors 1231-1232, N belong to an operating system. In at least one embodiment, an operating system may virtualize the graphics processing engines 1231-1232, N to allow access to any process or application.

[0201] In at least one embodiment, the graphics acceleration module 1246 or an individual graphics processing engine 1231-1232,N selects a process element using a process handle. In one embodiment, process elements are stored in system memory 1214 and are addressable using an effective address to real address translation technique, which is described herein. In at least one embodiment, a process handle may be an implementation-specific value provided to a host process when it registers its context with the graphics processing engine 1231-1232,N (i.e., when it calls system software to add a process element to a linked process element list). In at least one embodiment, the lower 16 bits of a process handle may be an offset of the process element within a linked process element list.

[0202] Fig. 12D shows an example accelerator integration slice 1290. As used herein, a "slice" comprises a particular portion of the processing resources of accelerator integration circuit 1236. The effective application address space 1282 in system memory 1214 stores process elements 1283. In one embodiment, the process elements 1283 are stored in response to GPU calls 1281 from applications 1280 executing on the processor 1207. A process element 1283 contains the process state for the corresponding application 1280. A work descriptor (WD) 1284 contained in the process element 1283 may be a single job requested by an application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 1284 is a pointer to a job request queue in the address space 1282 of an application.

[0203] The graphics acceleration module 1246 and / or the individual graphics processing engines 1231-1232, N may be shared by all or a subset of the processes in a system. In at least one embodiment, an infrastructure for establishing process status and sending a WD 1284 to a graphics acceleration module 1246 to start a job may exist in a virtualized environment.

[0204] In at least one embodiment, a dedicated process programming model is implementation-specific. In this model, a single process owns the graphics acceleration module 1246 or a single graphics processing engine 1231. Because the graphics acceleration module 1246 is owned by a single process, a hypervisor initializes the accelerator integration circuit 1236 for an owning partition, and an operating system initializes the accelerator integration circuit 1236 for an owning process when the graphics acceleration module 1246 is allocated.

[0205] In operation, a WD fetch unit 1291 in accelerator integration slice 1290 fetches the next WD 1284, which includes an indication of the work to be performed by one or more graphics processing engines of graphics acceleration module 1246. The data from WD 1284 may be stored in registers 1245 and used by MMU 1239, interrupt management circuitry 1247, and / or context management circuitry 1248, as illustrated. For example, one embodiment of MMU 1239 includes segment / page walkup circuitry for accessing segment / page tables 1286 in operating system virtual address space 1285. Interrupt management circuitry 1247 may process interrupt events 1292 received from graphics acceleration module 1246.When performing graphics operations, an effective address 1293 generated by a graphics processing engine 1231-1232, N is translated into a real address by the MMU 1239.

[0206] In one embodiment, the same set of registers 1245 is duplicated for each graphics processing engine 1231-1232, N, and / or each graphics acceleration module 1246 and may be initialized by a hypervisor or operating system. Each of these duplicated registers may be present in an accelerator integration slice 1290. Example registers that may be initialized by a hypervisor are listed in Table 1. Table 1 - Registers initialized by the hypervisor 1 Slice-Steuerungsregister 2 Reale Adresse (RA) Bereichszeiger geplanter Prozesse 3 Autoritätsmasken-Überschreibungsregister 4 Unterbrechungsvektor-Tabelleneintrags-Offset 5 Unterbrechungsvektor-Tabelleneintragsgrenze 6 Statusregister 7 Logische Partitions-ID 8 Reale Adresse (RA) Hypervisor-Beschleuniger-Nutzungsdatensatzzeiger 9 Speicherbeschreibungsregister

[0207] Example registers that can be initialized by an operating system are listed in Table 2. Table 2 - Initialized operating system registers 1 Prozess- und Thread-Identifikation 2 Effective Address (EA) Context Store / Restore Pointer 3 Virtual Address (VA) Accelerator Usage Record Pointer 4 Virtual Address (VA) Pointer to the memory segment table 5 Authority mask 6 Work descriptor

[0208] In one embodiment, each WD 1284 is specific to a particular graphics acceleration module 1246 and / or the graphics processing engines 1231-1232, N. It contains all the information required by a graphics processing engine 1231-1232, N to perform work, or it may be a pointer to a memory location where an application has established a command queue of work to be performed.

[0209] Fig. 12E illustrates additional details for an exemplary embodiment of a joint model. This embodiment includes a real hypervisor address space 1298 in which a process element list 1299 is stored. The real hypervisor address space 1298 is accessible via a hypervisor 1296 that virtualizes graphics acceleration engine machines for the operating system 1295.

[0210] In at least one embodiment, shared programming models allow all or a subset of processes from all or a subset of partitions in a system to use a graphics acceleration module 1246. There are two programming models in which the graphics acceleration module 1246 is shared among multiple processes and partitions: time-sharing and graphics-directed sharing.

[0211] In this model, the system hypervisor 1296 owns the graphics acceleration module 1246 and makes its functionality available to all operating systems 1295. For a graphics acceleration module 1246 to support virtualization by the system hypervisor 1296, the graphics acceleration module 1246 can meet the following conditions: 1) An application's job request must be autonomous (i.e., state does not need to be maintained between jobs), or the graphics acceleration module 1246 must provide a context save and restore mechanism. 2) The graphics acceleration module 1246 guarantees that an application's job request will be completed within a specified amount of time, including any compilation errors, or the graphics acceleration module 1246 provides the ability to interrupt the processing of a job.3) The graphics acceleration module 1246 must be guaranteed fairness between processes when operating in a directed joint programming model.

[0212] In at least one embodiment, the application 1280 must execute an operating system 1295 system call with a graphics acceleration module 1246 type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module 1246 type describes a targeted acceleration function for a system call. In at least one embodiment, the graphics acceleration module 1246 type may be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 1246 and may be in the form of a graphics acceleration module 1246 instruction, an effective address pointer to a user-defined structure, an effective address pointer to an instruction queue, or another data structure that describes the work to be performed by the graphics acceleration module 1246.In one embodiment, an AMR value is an AMR state to be used for a current process. In at least one embodiment, a value passed to an operating system is similar to an application setting an AMR. If the implementations of accelerator integration circuit 1236 and graphics acceleration module 1246 do not support a User Authority Mask Override Register (UAMOR), an operating system may apply a current UAMOR value to an AMR value before passing an AMR in a hypervisor call. The hypervisor 1296 may optionally apply a current Authority Mask Override Register (AMOR) value before placing an AMR in a process element 1283.In at least one embodiment, CSRP is one of the registers 1245 that contains an effective address of a region in an application's address space 1282 for the graphics acceleration module 1246 to store and restore context state. This pointer is optional if no state needs to be saved between jobs or if a job terminates prematurely. In at least one embodiment, the context save / restore region may be located in system memory.

[0213] Upon receiving a system call, the operating system 1295 can verify whether the application 1280 is registered and has been granted permission to use the graphics acceleration module 1246. The operating system 1295 then calls the hypervisor 1296 with the information shown in Table 3. Table 3 - Hypervisor call parameters from the operating system 1 A work descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked) 3 An effective address (EA) of a context backup / restore region pointer (CSRP) 4 A process ID (PID) and optionally a thread ID (TID) 5 A virtual address (VA) accelerator usage record pointer (AURP) 6 Virtual address of a memory segment table pointer (SSTP) 7 A logical interrupt service number (LISN)

[0214] Upon receiving a hypervisor call, hypervisor 1296 checks whether operating system 1295 is registered and has been granted permission to use graphics acceleration module 1246. Hypervisor 1296 then places process element 1283 in a linked process element list for a corresponding graphics acceleration module type 1246. A process element may include the information shown in Table 4. Table 4 - Process element information 1 A work descriptor (WD) 2 An Authority Mask Register (AMR) value (possibly masked) 3 An effective address (EA) of a context backup / restore region pointer (CSRP) 4 A process ID (PID) and optionally a thread ID (TID) 5 A virtual address (VA) accelerator usage record pointer (AURP) 6 Virtual address of a memory segment table pointer (SSTP) 7 A logical interrupt service number (LISN) 8 Interrupt vector table derived from hypervisor call parameters 9 A status register (SR) value 10 A logical partition ID (LPID) 11 Real Address (RA) Hypervisor Accelerator Usage Record Pointer 12 Memory Description Register (SDR)

[0215] In at least one embodiment, the hypervisor initializes a plurality of registers 1245 for accelerator integration slices 1290.

[0216] As it is in Fig.12F, at least one embodiment uses a unified memory addressable via a common virtual memory address space used to access physical processor memories 1201-1202 and GPU memories 1220-1223. In this implementation, operations performed on GPUs 1210-1213 use the same virtual / effective memory address space to access processor memories 1201-1202 and vice versa, simplifying programmability. In one embodiment, a first portion of a virtual / effective address space is assigned to processor memory 1201, a second portion is assigned to second processor memory 1202, a third portion is assigned to GPU memory 1220, and so on.In at least one embodiment, this distributes an entire virtual / effective memory space (sometimes referred to as effective address space) across each of the processor memories 1201-1202 and GPU memories 1220-1223, allowing any processor or GPU to access any physical memory with a virtual address associated with that memory.

[0217] In one embodiment, bias / coherence management circuitry 1294A-1294E within one or more MMUs 1239A-1239E ensures cache coherence between the caches of one or more host processors (e.g., 1205) and GPUs 1210-1213 and implements biasing methods that indicate in which physical memories certain data types should be stored. While multiple instances of bias / coherence management circuitry 1294A-1294E in Fig.12F, the bias / coherence circuitry may be implemented within an MMU of one or more host processors 1205 and / or within the accelerator integration circuit 1236.

[0218] One embodiment enables GPU-attached memory 1220-1223 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, but without incurring performance penalties associated with full system cache coherence. In at least one embodiment, the ability to access GPU-attached memory 1220-1223 as system memory without burdensome cache coherence overhead provides a favorable operating environment for GPU offload. This arrangement enables host processor 1205 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. Such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses.In at least one embodiment, the ability to access GPU-attached memory 1220-1223 without cache coherence overheads may be critical to the execution time of an offloaded computation. For example, in cases with significant streaming write memory traffic, the cache coherence overhead may significantly reduce the effective write bandwidth of a GPU 1210-1213. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation may play a role in determining the effectiveness of a GPU offload.

[0219] In at least one embodiment, the selection of a GPU bias and a host processor bias is controlled by a bias tracker data structure. For example, a bias table may be used, which may be a page-granular structure (i.e., controlled at the granularity of a memory page) having 1 or 2 bits per GPU-attached memory page. In at least one embodiment, a bias table may be implemented in a stolen memory region of one or more GPU-attached memories 1220-1223, with or without a bias cache in GPU 1210-1213 (e.g., to cache frequently / recently used bias table entries). Alternatively, an entire bias table may be maintained in a GPU.

[0220] In at least one embodiment, prior to actually accessing a GPU memory, a bias table entry associated with each access to GPU-attached memory 1220-1223 is accessed, causing the following operations. First, local requests from GPU 1210-1213 that find their page in the GPU bias are forwarded directly to a corresponding GPU memory 1220-1223. Local requests from a GPU that find their page in the host bias are forwarded to processor 1205 (e.g., over a high-speed connection, as described above). In one embodiment, requests from processor 1205 that find a requested page in the host processor bias are completed like a normal memory read access. Alternatively, requests directed to a GPU-biased page may be forwarded to GPU 1210-1213.In at least one embodiment, a GPU may then transition a page to host processor bias when it is not currently using the page. In at least one embodiment, the bias state of a page may be changed either through a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited number of cases, a purely hardware-based mechanism.

[0221] A mechanism for changing the bias state uses an API call (e.g., OpenCL), which in turn calls a GPU's device driver, which in turn sends a message to a GPU (or queues a command descriptor) to instruct it to change a bias state and perform a cache flush operation in a host for some transitions. In at least one embodiment, the cache flush operation is used for a transition from the host processor 1205 bias to the GPU bias, but not for an opposite transition.

[0222] In one embodiment, cache coherence is maintained by temporarily rendering GPU-bound pages that cannot be cached by the host processor 1205. To access these pages, the processor 1205 may request access from the GPU 1210, which may not grant access immediately. Therefore, to reduce communication between the processor 1205 and the GPU 1210, it is advantageous to ensure that GPU-bound pages are those required by a GPU but not by the host processor 1205, and vice versa.

[0223] Fig.Figure 13 illustrates exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to the illustrated circuitry, additional logic and circuitry may be present in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0224] Fig.13 is a block diagram illustrating an exemplary system-on-a-chip integrated circuit 1300 that may be manufactured using one or more IP cores, according to at least one embodiment. In at least one embodiment, the integrated circuit 1300 includes one or more application processors 1305 (e.g., CPUs), at least one graphics processor 1310, and may additionally include an image processor 1315 and / or a video processor 1320, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 1300 includes peripheral or bus logic, including a USB controller 1325, a UART controller 1330, an SPI / SDIO controller 1335, and an I.sup.2S / I.sup.2C controller 1340.In at least one embodiment, integrated circuit 1300 may include a display device 1345 connected to one or more HDMI (High-Definition Multimedia Interface) controllers 1350 and a MIPI (Mobile Industry Processor Interface) display interface 1355. In at least one embodiment, memory may be provided by a flash memory subsystem 1360 including flash memory and a flash memory controller. In at least one embodiment, the memory interface may be provided via a memory controller 1365 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 1370.

[0225] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0226] Fig. 14A-14B illustrate example integrated circuits and associated graphics processors that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to the illustrated circuitry, other logic and circuitry may be present in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0227] Fig.14A-14B are block diagrams illustrating example graphics processors for use in an SoC according to the embodiments described herein. Fig. 14A shows an exemplary graphics processor 1410 of a system-on-chip integrated circuit that may be manufactured using one or more IP cores, in accordance with at least one embodiment. Fig. 14B shows another exemplary graphics processor 1440 of a system-on-a-chip integrated circuit that may be fabricated using one or more IP cores, in accordance with at least one embodiment. In at least one embodiment, the graphics processor 1410 is Fig. 14A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1440 is Fig.14B, a higher performance graphics processor core. In at least one embodiment, each of the graphics processors 1410, 1440 may be a variant of the graphics processor 1310 of Fig. be 13.

[0228] In at least one embodiment, graphics processor 1410 includes a vertex processor 1405 and one or more fragment processors 1415A-1415N (e.g., 1415A, 1415B, 1415C, 1415D through 1415N-1, and 1415N). In at least one embodiment, graphics processor 1410 may execute different shader programs via separate logic, such that vertex processor 1405 is optimized to perform operations for vertex shader programs, while one or more fragment processors 1415A-1415N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1405 performs a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data.In at least one embodiment, fragment processors 1415A-1415N use the primitive and vertex data generated by vertex processor 1405 to generate a frame buffer displayed on a display device. In at least one embodiment, fragment processor(s) 1415A-1415N are optimized for executing fragment shader programs, such as those provided in an OpenGL API, which can be used to perform similar operations as a pixel shader program, such as those provided in a Direct 3D API.

[0229] In at least one embodiment, graphics processor 1410 additionally includes one or more memory management units (MMUs) 1420A-1420B, one or more caches 1425A-1425B, and one or more circuit interconnects 1430A-1430B. In at least one embodiment, one or more MMUs 1420A-1420B provide virtual-to-physical address mapping for graphics processor 1410, including vertex processor 1405 and / or fragment processor(s) 1415A-1415N, which may reference vertex or image / texture data stored in memory in addition to the vertex or image / texture data stored in one or more caches 1425A-1425B.In at least one embodiment, one or more MMU(s) 1420A-1420B may be synchronized with other MMUs within the system, including one or more MMUs associated with one or more application processors 1305, image processors 1315, and / or video processors 1320. Fig. 13, so that each processor 1305-1320 can participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1430A-1430B enable the graphics processor 1410 to connect to other IP cores within the SoC, either via an internal bus of the SoC or via a direct connection.

[0230] In at least one embodiment, the graphics processor 1440 includes one or more MMU(s) 1420A-1420B, caches 1425A-1425B, and circuit interconnects 1430A-1430B of the graphics processor 1410 of Fig.14A. In at least one embodiment, the graphics processor 1440 includes one or more shader cores 1455A-1455N (e.g., 1455A, 1455B, 1455C, 1455D, 1455E, 1455F through 1455N-1 and 1455N), enabling a unified shader core architecture where a single core or type of core can execute all types of programmable shader code, including shader code implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores may vary.In at least one embodiment, the graphics processor 1440 includes an inter-core task manager 1445 acting as a thread dispatcher to distribute execution threads to one or more shader cores 1455A-1455N and a tiling unit 1458 to accelerate tiling operations for tile-based rendering, in which rendering operations for a scene are divided in image space, for example, to exploit local spatial coherence within a scene or to optimize the use of internal caches.

[0231] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0232] Fig. 15A and Fig.15B illustrate additional example graphics processor logic according to the embodiments described herein. Fig. 15A shows a graphics core 1500 that, in at least one embodiment, is included in the graphics processor 1310 of Fig. 13 and in at least one embodiment, a unified shader core 1455A-1455N as shown in Fig. 14B can be. Fig. 15B illustrates a highly parallel, general-purpose graphics processing unit 1530 suitable for use on a multi-chip module in at least one embodiment.

[0233] In at least one embodiment, graphics core 1500 includes a shared instruction cache 1502, a texture unit 1518, and a cache / shared memory 1520 common to the execution resources within graphics core 1500. In at least one embodiment, graphics core 1500 may include multiple slices 1501A-1501N or partitions for each core, and a graphics processor may include multiple instances of graphics core 1500. Slices 1501A-1501N may include support logic including a local instruction cache 1504A-1504N, a thread scheduler 1506A-1506N, a thread dispatcher 1508A-1508N, and a set of registers 1510A-1510N.In at least one embodiment, slices 1501A-1501N may include a set of additional functional units (AFUs 1512A-1512N), floating point units (FPU 1514A-1514N), integer arithmetic logic units (ALUs 1516-1516N), address calculation units (ACU 1513A-1513N), double precision floating point units (DPFPU 1515A-1515N), and matrix processing units (MPU 1517A-1517N).

[0234] In at least one embodiment, FPUs 1514A-1514N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while DPFPUs 1515A-1515N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, ALUs 1516A-1516N can perform variable-precision integer operations at 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, MPUs 1517A-1517N can also be configured for mixed-precision matrix operations, including floating-point and 8-bit half-precision integer operations. In at least one embodiment, MPUs 1517-1517N may perform a variety of matrix operations to accelerate machine learning application frameworks, including support for accelerated generalized matrix-matrix multiplication (GEMM).In at least one embodiment, AFUs 1512A-1512N may perform additional logical operations not supported by floating point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0235] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0236] Fig.15B illustrates a general-purpose processing unit (GPGPU) 1530, which, in at least one embodiment, may be configured to perform highly parallel computational operations through an array of graphics processing units. In at least one embodiment, the GPGPU 1530 may be directly connected to other instances of the GPGPU 1530 to form a multi-GPU cluster and improve training speed for deep neural networks. In at least one embodiment, the GPGPU 1530 includes a host interface 1532 to enable connection to a host processor. In at least one embodiment, the host interface 1532 is a PCI Express interface. In at least one embodiment, the host interface 1532 may be a vendor-specific communication interface or communication fabric.In at least one embodiment, GPGPU 1530 receives instructions from a host processor and uses a global scheduler 1534 to distribute the execution threads associated with those instructions across a number of compute clusters 1536A-1536H. In at least one embodiment, compute clusters 1536A-1536H share a cache 1538. In at least one embodiment, cache 1538 may serve as a parent cache for caches within compute clusters 1536A-1536H.

[0237] In at least one embodiment, GPGPU 1530 includes memory 1544A-1544B coupled to compute clusters 1536A-1536H via a series of memory controllers 1542A-1542B. In at least one embodiment, memory 1544A-1544B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate memory (GDDR).

[0238] In at least one embodiment, the compute clusters 1536A-1536H each comprise a set of graphics cores, such as the graphics core 1500 of Fig.15A, which may include multiple types of integer and floating-point logic units capable of performing computational operations at a range of precisions, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 1536A-1536H may be configured to perform 16-bit or 32-bit floating-point operations, while another subset of floating-point units may be configured to perform 64-bit floating-point operations.

[0239] In at least one embodiment, multiple instances of GPGPU 1530 may be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 1536A-1536H for synchronization and data exchange varies between embodiments. In at least one embodiment, multiple instances of GPGPU 1530 communicate via host interface 1532. In at least one embodiment, GPGPU 1530 includes an I / O hub 1539 that couples GPGPU 1530 to a GPU link 1540 that enables direct connection to other instances of GPGPU 1530. In at least one embodiment, GPU link 1540 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1530.In at least one embodiment, GPU link 1540 is coupled to a high-speed interconnect for sending and receiving data to other GPGPUs or parallel processors. In at least one embodiment, multiple instances of GPGPU 1530 are located in separate computing systems and communicate via a network device accessible via host interface 1532. In at least one embodiment, GPU interconnect 1540 may be configured to enable connection to a host processor in addition to, or alternatively to, host interface 1532.

[0240] In at least one embodiment, the GPGPU 1530 may be configured to train neural networks. In at least one embodiment, the GPGPU 1530 may be used within an inferencing platform. In at least one embodiment where the GPGPU 1530 is used for inferencing, the GPGPU may have fewer compute clusters 1536A-1536H than when the GPGPU is used to train a neural network. In at least one embodiment, the memory technology associated with the memory 1544A-1544B may differ between inferencing and training configurations, with higher-bandwidth memory technologies being allocated to the training configurations. In at least one embodiment, the inferencing configuration of the GPGPU 1530 may support inferencing-specific instructions.For example, in at least one embodiment, an inferencing configuration may provide support for one or more 8-bit integer dot product instructions that may be used during inferencing operations for deployed neural networks.

[0241] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0242] Fig.16 is a block diagram illustrating a computing system 1600 according to at least one embodiment. In at least one embodiment, computing system 1600 includes a processing subsystem 1601 having one or more processors 1602 and a system memory 1604 communicating via an interconnect path that may include a memory hub 1605. In at least one embodiment, memory hub 1605 may be a separate component within a chipset component or integrated with one or more processors 1602. In at least one embodiment, memory hub 1605 is connected to an I / O subsystem 1611 via a communications link 1606. In at least one embodiment, I / O subsystem 1611 includes an I / O hub 1607 that enables computing system 1600 to receive input from one or more input devices 1608.In at least one embodiment, the I / O hub 1607 may enable a display controller, which may be included in one or more processors 1602, to provide outputs to one or more display devices 1610A. In at least one embodiment, one or more display devices 1610A coupled to the I / O hub 1607 may comprise a local, internal, or embedded display device.

[0243] In at least one embodiment, processing subsystem 1601 includes one or more parallel processors 1612 connected to storage hub 1605 via a bus or other communication link 1613. In at least one embodiment, communication link 1613 may be any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or a vendor-specific communication interface or communication structure. In at least one embodiment, one or more parallel processors 1612 form a computationally focused parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a Many Integrated Core (MIC) processor.In at least one embodiment, one or more parallel processors 1612 form a graphics processing subsystem that can output pixels to one or more display devices 1610A coupled via the I / O hub 1607. In at least one embodiment, one or more parallel processors 1612 may also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1610B.

[0244] In at least one embodiment, a system storage unit 1614 may be coupled to the I / O hub 1607 to provide a storage mechanism for the computer system 1600. In at least one embodiment, an I / O switch 1616 may be used to provide an interface mechanism to enable connections between the I / O hub 1607 and other components, such as a network adapter 1618 and / or a wireless network adapter 1619 that may be integrated into the platform, and various other devices that may be added via one or more add-in devices 1620. In at least one embodiment, the network adapter 1618 may be an Ethernet adapter or other wired network adapter.In at least one embodiment, the wireless network adapter 1619 may include one or more Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices that include one or more wireless radios.

[0245] In at least one embodiment, the computing system 1600 may include other components not explicitly shown, including USB or other ports, optical storage drives, video capture devices, and the like, that may also be connected to the I / O hub 1607. In at least one embodiment, communication paths connecting various components in Fig.16 interconnect, be implemented using any suitable protocols, such as PCI (Peripheral Component Interconnect)-based protocols (e.g. PCI Express) or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed interconnect or interconnect protocols.

[0246] In at least one embodiment, one or more parallel processors 1612 include circuitry optimized for graphics and video processing, for example, including video output circuitry and constituting a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1612 include circuitry optimized for general processing. In at least one embodiment, components of computing system 1600 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1612, a memory hub 1605, a processor(s) 1602, and an I / O hub 1607 may be integrated into an integrated circuit comprising a system on a chip (SoC).In at least one embodiment, the components of computing system 1600 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 1600 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules to form a modular computing system.

[0247] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing. PROCESSORS

[0248] Fig.17A illustrates a parallel processor 1700 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 1700 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 1700 is a variant of one or more parallel processors 1612 described in Fig. 16 according to an exemplary embodiment.

[0249] In at least one embodiment, parallel processor 1700 includes a parallel processing unit 1702. In at least one embodiment, parallel processing unit 1702 includes an I / O unit 1704 that enables communication with other devices, including other instances of parallel processing unit 1702. In at least one embodiment, I / O unit 1704 may be directly connected to other devices. In at least one embodiment, I / O unit 1704 is connected to other devices via a hub or switch interface, such as storage hub 1705. In at least one embodiment, the connections between storage hub 1705 and I / O unit 1704 form a communication link.In at least one embodiment, the I / O unit 1704 is coupled to a host interface 1706 and a memory switch 1716, where the host interface 1706 receives commands to perform processing operations and the memory switch 1716 receives commands to perform memory operations.

[0250] In at least one embodiment, when host interface 1706 receives a command buffer via I / O unit 1704, host interface 1706 may direct work operations to a front end 1708 to execute those commands. In at least one embodiment, front end 1708 is coupled to a scheduler 1710 configured to dispatch commands or other work items to a processing cluster arrangement 1712. In at least one embodiment, scheduler 1710 ensures that processing cluster arrangement 1712 is properly configured and in a valid state before dispatching tasks to processing cluster arrangement 1712. In at least one embodiment, scheduler 1710 is implemented via firmware logic executing on a microcontroller.In at least one embodiment, the microcontroller-implemented scheduler 1710 is configured to perform complex scheduling and work distribution operations at coarse and fine granularity, enabling rapid interruption and context switching of threads executing on the processing array 1712. In at least one embodiment, host software may allocate workloads for scheduling on the processing array 1712 via one of several graphics processing doorbells. In at least one embodiment, the workloads may then be automatically distributed on the processing array 1712 by the logic of the scheduler 1710 within a microcontroller including the scheduler 1710.

[0251] In at least one embodiment, the processing cluster arrangement 1712 may include up to "N" processing clusters (e.g., cluster 1714A, cluster 1714B, through cluster 1714N). In at least one embodiment, each cluster 1714A-1714N of the processing cluster arrangement 1712 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1710 may allocate work to the clusters 1714A-1714N of the processing cluster arrangement 1712 using various scheduling and / or work distribution algorithms, which may vary depending on the workload incurred for each type of program or computation. In at least one embodiment, the scheduling may be performed dynamically by the scheduler 1710 or may be assisted in part by compiler logic during compilation of the program logic configured for execution by the processing cluster arrangement 1712.In at least one embodiment, different clusters 1714A-1714N of the processing cluster array 1712 may be assigned for processing different types of programs or for performing different types of computations.

[0252] In at least one embodiment, processing cluster assembly 1712 may be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster assembly 1712 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, processing cluster assembly 1712 may include logic to perform processing tasks, including filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.

[0253] In at least one embodiment, processing cluster assembly 1712 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster assembly 1712 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic to perform texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, processing cluster assembly 1712 may be configured to execute graphics processing-related shader programs, such as vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing unit 1702 may transfer data from system memory via I / O unit 1704 for processing.In at least one embodiment, the transferred data may be stored in on-chip memory (e.g., parallel processor memory 1722) during processing and then written back to system memory.

[0254] In at least one embodiment, when parallel processing unit 1702 is used to perform graphics processing, scheduler 1710 may be configured to divide a processing load into approximately equal-sized tasks to enable better distribution of graphics processing operations across multiple clusters 1714A-1714N of processing cluster array 1712. In at least one embodiment, portions of processing cluster array 1712 may be configured to perform different types of processing.For example, in at least one embodiment, a first section may be configured to perform vertex shading and topology generation, a second section may be configured to perform tessellation and geometry shading, and a third section may be configured to perform pixel shading or other screen-space operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more clusters 1714A-1714N may be stored in buffers to allow intermediate data to be transferred between clusters 1714A-1714N for further processing.

[0255] In at least one embodiment, the processing cluster arrangement 1712 may receive processing tasks to be executed via the scheduler 1710, which receives commands defining processing tasks from the front end 1708. In at least one embodiment, the processing tasks may include indices of the data to be processed, e.g., surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands that define how the data is to be processed (e.g., which program is to be executed). In at least one embodiment, the scheduler 1710 may be configured to retrieve indices corresponding to the tasks or to receive indices from the front end 1708. In at least one embodiment, the front end 1708 may be configured to ensure that the processing cluster arrangement 1712 is configured in a valid state before executing a command defined by incoming command buffers (e.g.,Batch buffer, push buffer, etc.) specified workload is initiated.

[0256] In at least one embodiment, each of one or more instances of parallel processing unit 1702 may be coupled to parallel processor memory 1722. In at least one embodiment, parallel processor memory 1722 may be accessed via memory switch 1716, which may receive memory requests from processing cluster arrangement 1712 as well as I / O unit 1704. In at least one embodiment, memory switch 1716 may access parallel processor memory 1722 via a memory interface 1718. In at least one embodiment, memory interface 1718 may include a plurality of partition units (e.g., partition unit 1720A, partition unit 1720B through partition unit 1720N), each of which may be coupled to a portion (e.g., a memory unit) of parallel processor memory 1722.In at least one embodiment, a number of partition units 1720A-1720N is configured to be equal to a number of storage units, such that a first partition unit 1720A has a corresponding first storage unit 1724A, a second partition unit 1720B has a corresponding storage unit 1724B, and an Nth partition unit 1720N has a corresponding Nth storage unit 1724N. In at least one embodiment, a number of partition units 1720A-1720N may not be equal to a number of storage devices.

[0257] In at least one embodiment, memory units 1724A-1724N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate memory (GDDR). In at least one embodiment, memory units 1724A-1724N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets, such as frame buffers or texture maps, may be stored across memory units 1724A-1724N so that partition units 1720A-1720N can write portions of each rendering target in parallel to efficiently utilize the available bandwidth of parallel processor memory 1722.In at least one embodiment, a local instance of parallel processor memory 1722 may be eliminated in favor of a unified memory design that utilizes system memory in conjunction with the local cache memory.

[0258] In at least one embodiment, each of the clusters 1714A-1714N of the processing cluster arrangement 1712 can process data written to each of the memory units 1724A-1724N in the parallel processor memory 1722. In at least one embodiment, the memory switch 1716 can be configured to transfer an output of each cluster 1714A-1714N to any partition unit 1720A-1720N or to another cluster 1714A-1714N that can perform additional processing operations on an output. In at least one embodiment, each cluster 1714A-1714N can communicate with the memory interface 1718 via the memory switch 1716 to read from or write to various external devices.In at least one embodiment, memory switch 1716 has a connection to memory interface 1718 to communicate with I / O unit 1704, as well as a connection to a local instance of parallel processor memory 1722 so that the processing units in the various processing clusters 1714A-1714N can communicate with system memory or other memory not local to parallel processing unit 1702. In at least one embodiment, memory switch 1716 can use virtual channels to separate traffic flows between clusters 1714A-1714N and partition units 1720A-1720N.

[0259] In at least one embodiment, multiple instances of the parallel processing unit 1702 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 1702 may be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 1702 may include higher-precision floating-point units compared to other implementations.In at least one embodiment, systems including one or more instances of the parallel processing unit 1702 or the parallel processor 1700 may be implemented in a variety of embodiments and form factors, including, but not limited to, desktop, laptop, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0260] Fig. 17B is a block diagram of a partition unit 1720 according to at least one embodiment. In at least one embodiment, the partition unit 1720 is an instance of one of the partition units 1720A-1720N of Fig.17A. In at least one embodiment, partition unit 1720 includes an L2 cache 1721, a frame buffer interface 1725, and a ROP (raster operation unit) 1726. The L2 cache 1721 is a read / write cache configured to perform load and store operations received from the memory switch 1716 and the ROP 1726. In at least one embodiment, read misses and urgent write-back requests are issued from the L2 cache 1721 to the frame buffer interface 1725 for processing. In at least one embodiment, updates may also be sent to a frame buffer via the frame buffer interface 1725 for processing. In at least one embodiment, the frame buffer interface 1725 is coupled to one of the memory units in the parallel processor memory, such as the memory units 1724A-1724N of Fig. 17 (e.g. within the parallel processor memory 1722).

[0261] In at least one embodiment, ROP 1726 is a processing unit that performs raster operations such as stenciling, Z-testing, blending, and the like. In at least one embodiment, ROP 1726 then outputs processed graphics data, which is stored in graphics memory. In at least one embodiment, ROP 1726 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic may be lossless compression logic using one or more of several compression algorithms. In at least one embodiment, the type of compression performed by ROP 1726 may vary based on statistical characteristics of the data to be compressed.For example, in at least one embodiment, delta color compression is performed on depth and color data on a per-tile basis.

[0262] In at least one embodiment, the ROP 1726 is in each processing cluster (e.g., clusters 1714A-1714N of Fig. 17) and not present in the partition unit 1720. In at least one embodiment, read and write requests for pixel data are transmitted over the memory switch 1716 instead of pixel fragment data. In at least one embodiment, processed graphics data may be displayed on a display device, such as one of one or more display devices 1610 of Fig. 16, for further processing by processor(s) 1602 or for further processing by one of the processing units within the parallel processor 1700 of Fig. 17A.

[0263] Fig.17C is a block diagram of a processing cluster 1714 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is an instance of one of the processing clusters 1714A-1714N of Fig.17. In at least one embodiment, the processing cluster 1714 may be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, single-instruction, multiple-data (SIMD) instruction issuing techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single-instruction multiple-thread (SIMT) techniques are used to support the parallel execution of a large number of generally synchronized threads, where a common instruction unit is configured to issue instructions to a set of processing engines within each of the processing clusters.

[0264] In at least one embodiment, the operation of the processing cluster 1714 may be controlled by a pipeline manager 1732 that distributes the processing tasks to parallel SIMT processors. In at least one embodiment, the pipeline manager 1732 receives instructions from the scheduler 1710 of the Fig.17 and manages the execution of these instructions via a graphics multiprocessor 1734 and / or a texture unit 1736. In at least one embodiment, the graphics multiprocessor 1734 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be present in the processing cluster 1714. In at least one embodiment, one or more instances of the graphics multiprocessor 1734 may be present in a processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 may process data, and a data switch 1740 may be used to distribute the processed data to one of several possible destinations, including other shader units.In at least one embodiment, the pipeline manager 1732 may facilitate the distribution of the processed data by specifying destinations for the processed data to be distributed across the data switch 1740.

[0265] In at least one embodiment, each graphics multiprocessor 1734 within the processing cluster 1714 may include an identical set of functional execution logic (e.g., arithmetic logic units, load storage units, etc.). In at least one embodiment, the functional execution logic may be pipelined so that new instructions may be issued before previous instructions complete. In at least one embodiment, the functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifting, and the computation of various algebraic functions. In at least one embodiment, the same hardware with functional units may be used to perform different operations, and any combination of functional units may be present.

[0266] In at least one embodiment, the instructions transferred to processing cluster 1714 form a thread. In at least one embodiment, a set of threads executing across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program with different input data. In at least one embodiment, each thread within a thread group may be assigned to a different processing engine within a graphics multiprocessor 1734. In at least one embodiment, a thread group may have fewer threads than the number of processing units in the graphics multiprocessor 1734. In at least one embodiment, when a thread group has fewer threads than a number of processing engines, one or more of the processing engines may be idle during the cycles in which that thread group is processing.In at least one embodiment, a thread group may also include more threads than the number of processing engines in the graphics multiprocessor 1734. In at least one embodiment, when a thread group includes more threads than the number of processing engines in the graphics multiprocessor 1734, processing may occur over consecutive clock cycles. In at least one embodiment, multiple thread groups may execute concurrently on a graphics multiprocessor 1734.

[0267] In at least one embodiment, the graphics multiprocessor 1734 includes an internal cache to perform load and store operations. In at least one embodiment, the graphics multiprocessor 1734 may forgo an internal cache and use a cache (e.g., L1 cache 1748) within the processing cluster 1714. In at least one embodiment, each graphics multiprocessor 1734 also has access to L2 caches within partition units (e.g., partition units 1720A-1720N of Fig.17) that are shared by all processing clusters 1714 and can be used to transfer data between threads. In at least one embodiment, the graphics multiprocessor 1734 can also access off-chip global memory, which can include one or more local parallel processor memories and / or system memories. In at least one embodiment, any memory external to the parallel processing unit 1702 can be used as global memory. In at least one embodiment, the processing cluster 1714 includes multiple instances of the graphics multiprocessor 1734, which can share common instructions and data, which can be stored in the L1 cache 1748.

[0268] In at least one embodiment, each processing cluster 1714 may include a memory management unit (MMU) 1745 configured to translate virtual addresses into physical addresses. In at least one embodiment, one or more instances of the MMU 1745 may reside within the memory interface 1718 of Fig.17. In at least one embodiment, the MMU 1745 includes a set of page table entries (PTEs) used to map a virtual address to a physical address of a tile, and optionally a cache line index. In at least one embodiment, the MMU 1745 may include address translation lookaside buffers (TLBs) or caches, which may be located in the graphics multiprocessor 1734 or the L1 cache or the processing cluster 1714. In at least one embodiment, the physical address is processed to distribute access locality to the surface data to enable efficient request interleaving between the partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0269] In at least one embodiment, a processing cluster 1714 may be configured such that each graphics multiprocessor 1734 is coupled to a texture unit 1736 to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1734 and retrieved from an L2 cache, local parallel processor memory, or system memory as needed.In at least one embodiment, each graphics multiprocessor 1734 outputs processed tasks to the data switch 1740 to make the processed task available to another processing cluster 1714 for further processing or to store the processed task in an L2 cache, local parallel processor memory, or system memory via the memory switch 1716. In at least one embodiment, a preROP 1742 (Pre-Raster Operations Unit) is configured to receive data from the graphics multiprocessor 1734 and forward data to ROP units that may be located in the partition units described herein (e.g., partition units 1720A-1720N of FIG. Fig. 17). In at least one embodiment, the PreROP unit 1742 may perform color mixing optimizations, organize pixel color data, and perform address translations.

[0270] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0271] Fig.17D shows a graphics multiprocessor 1734 according to at least one embodiment. In at least one embodiment, the graphics multiprocessor 1734 is coupled to the pipeline manager 1732 of the processing cluster 1714. In at least one embodiment, the graphics multiprocessor 1734 includes an execution pipeline including, among other things, an instruction cache 1752, an instruction unit 1754, an address mapping unit 1756, a register file 1758, one or more GPGPU cores 1762, and one or more load / store units 1766. The GPGPU cores 1762 and the load / store units 1766 are connected to the cache memory 1772 and the shared memory 1770 via a memory and cache interconnect 1768.

[0272] In at least one embodiment, instruction cache 1752 receives a stream of instructions to be executed from pipeline manager 1732. In at least one embodiment, the instructions are cached in instruction cache 1752 and forwarded for execution by instruction unit 1754. In at least one embodiment, instruction unit 1754 may dispatch the instructions as thread groups (e.g., warps), with each thread of the thread group assigned to a different execution unit within GPGPU core 1762. In at least one embodiment, an instruction may access a local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, address mapping unit 1756 may be used to translate addresses in a unified address space into a unique memory address accessible by load / store units 1766.

[0273] In at least one embodiment, register file 1758 provides a set of registers for functional units of graphics multiprocessor 1734. In at least one embodiment, register file 1758 provides temporary storage for operands associated with data paths of functional units (e.g., GPGPU cores 1762, load / store units 1766) of graphics multiprocessor 1734. In at least one embodiment, register file 1758 is partitioned between each functional unit such that each functional unit is assigned its own section of register file 1758. In at least one embodiment, register file 1758 is partitioned among different warps executed by graphics multiprocessor 1734.

[0274] In at least one embodiment, the GPGPU cores 1762 may each include floating-point units (FPUs) and / or integer arithmetic logic units (ALUs) used to execute instructions of the graphics multiprocessor 1734. The GPGPU cores 1762 may be similar or different in architecture. In at least one embodiment, a first portion of the GPGPU cores 1762 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU cores includes a double-precision FPU. In at least one embodiment, the FPUs may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic.In at least one embodiment, graphics multiprocessor 1734 may additionally include one or more fixed-function or special-function units to perform specific functions such as rectangle copying or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed or special-purpose functional logic.

[0275] In at least one embodiment, GPGPU cores 1762 include SIMD logic capable of executing a single instruction on multiple data sets. In at least one embodiment, GPGPU cores 1762 can physically execute SIMD4, SIMD8, and SIMD16 instructions, and logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, SIMD instructions for GPGPU cores can be generated at compile time by a shader compiler or automatically during the execution of programs written and compiled for SPMD or SIMT (Single Program Multiple Data) architectures. In at least one embodiment, multiple threads of a program designed for a SIMT execution model can execute via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can execute in parallel via a single SIMD8 logic unit.

[0276] In at least one embodiment, the memory and cache interconnect 1768 is an interconnect network that connects each functional unit of the graphics multiprocessor 1734 to the register file 1758 and the shared memory 1770. In at least one embodiment, the memory and cache interconnect 1768 is a cross-connect that enables the load / store unit 1766 to perform load and store operations between the shared memory 1770 and the register file 1758. In at least one embodiment, the register file 1758 may operate at the same frequency as the GPGPU cores 1762, such that data transfer between the GPGPU cores 1762 and the register file 1758 has very low latency. In at least one embodiment, the shared memory 1770 may be used to enable communication between threads executing on functional units within the graphics multiprocessor 1734.For example, in at least one embodiment, cache 1772 may be used as a data cache to cache texture data transferred between functional units and texture unit 1736. In at least one embodiment, shared memory 1770 may also be used as a programmatic cache. In at least one embodiment, threads executing on GPGPU cores 1762 may programmatically store data in shared memory in addition to the automatically cached data stored in cache 1772.

[0277] In at least one embodiment, a parallel processor or a GPGPU, as described herein, is communicatively coupled with host / processor cores to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general GPU (GPGPU) functions. In at least one embodiment, the GPU can be communicatively connected to the host processor (the processor cores) via a bus or another connection (e.g., a high-speed connection such as PCIe or NVLink). In at least one embodiment, the GPU can be integrated in the same housing or chip as the cores and communicate with the cores via an internal processor bus or an internal connection (i.e., within the housing or chip).In at least one embodiment, regardless of the GPU's connection type, the processor cores may allocate work to the GPU in the form of command sequences / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0278] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0279] Fig.18 shows a multi-GPU computing system 1800 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1800 may include a processor 1802 connected to a plurality of general-purpose graphics processing units (GPGPUs) 1806A-D via a host interface switch 1804. In at least one embodiment, the host interface switch 1804 is a PCI Express switch device that connects the processor 1802 to a PCI Express bus over which the processor 1802 can communicate with the GPGPUs 1806A-D. The GPGPUs 1806A-D may be interconnected via a series of high-speed point-to-point GPU-to-GPU interconnects 1816. In at least one embodiment, the GPU-to-GPU connections 1816 are connected to each of the GPGPUs 1806A-D via a separate GPU connection.In at least one embodiment, the P2P GPU connections 1816 enable direct communication between the individual GPGPUs 1806A-D without requiring communication over the host interface bus 1804 to which the processor 1802 is connected. In at least one embodiment where GPU-to-GPU traffic is directed on P2P GPU connections 1816, the host interface bus 1804 remains available for system memory access or for communication with other instances of the multi-GPU computer system 1800, for example, over one or more network devices. While in at least one embodiment the GPGPUs 1806A-D are connected to the processor 1802 via the host interface switch 1804, in at least one embodiment the processor 1802 has direct support for P2P GPU connections 1816 and may be directly connected to the GPGPUs 1806A-D.

[0280] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0281] Fig.19 is a block diagram of a graphics processor 1900 according to at least one embodiment. In at least one embodiment, the graphics processor 1900 includes a ring interconnect 1902, a pipeline front end 1904, a media engine 1937, and graphics cores 1980A-1980N. In at least one embodiment, the ring interconnect 1902 connects the graphics processor 1900 to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, the graphics processor 1900 is one of many processors integrated into a multi-core processing system.

[0282] In at least one embodiment, graphics processor 1900 receives batches of commands via ring interconnect 1902. In at least one embodiment, the incoming commands are interpreted by a command streamer 1903 in pipeline front end 1904. In at least one embodiment, graphics processor 1900 includes scalable execution logic to perform 3D geometry processing and media processing via graphics core(s) 1980A-1980N. In at least one embodiment, command streamer 1903 provides commands to geometry pipeline 1936 for 3D geometry processing commands. In at least one embodiment, command streamer 1903 provides commands to a video front end 1934 coupled to a media engine 1937 for at least some media processing commands.In at least one embodiment, the media engine 1937 includes a video quality engine (VQE) 1930 for video and image post-processing and a multi-format encoder / decoder engine (MFX) 1933 to enable hardware-accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 1936 and the media engine 1937 each generate execution threads for thread execution resources provided by at least one graphics core 1980A.

[0283] In at least one embodiment, graphics processor 1900 includes scalable threaded execution resources comprising modular cores 1980A-1980N (sometimes referred to as core slices), each having a plurality of sub-cores 1950A-1950N, 1960A-1960N (sometimes referred to as core sub-slices). In at least one embodiment, graphics processor 1900 may have any number of graphics cores 1980A-1980N. In at least one embodiment, graphics processor 1900 includes a graphics core 1980A having at least a first sub-core 1950A and a second sub-core 1960A. In at least one embodiment, graphics processor 1900 is a low-performance processor having a single sub-core (e.g., 1950A). In at least one embodiment, the graphics processor 1900 includes a plurality of graphics cores 1980A-1980N, each of which includes a set of first sub-cores 1950A-1950N and a set of second sub-cores 1960A-1960N.In at least one embodiment, each sub-core in the first sub-cores 1950A-1950N includes at least a first set of execution units 1952A-1952N and media / texture samplers 1954A-1954N. In at least one embodiment, each sub-core in the second sub-cores 1960A-1960N includes at least a second group of execution units 1962A-1962N and samplers 1964A-1964N. In at least one embodiment, each sub-core 1950A-1950N, 1960A-1960N shares a set of shared resources 1970A-1970N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.

[0284] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig.1-5 to perform ray tracing.

[0285] Fig.20 is a block diagram illustrating the microarchitecture of a processor 2000, which may include logic circuitry for executing instructions according to at least one embodiment. In at least one embodiment, the processor 2000 may execute instructions including x86 instructions, ARM instructions, special instructions for application-specific integrated circuits (ASICs), etc. In at least one embodiment, the processor 2010 may include registers for storing packed data, such as 64-bit wide MMX™ registers in microprocessors employing MMX technology from Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers, which are available as both integer and floating-point registers, may operate on packed data elements associated with SIMD (Single Instruction, Multiple Data) and SSE (Streaming SIMD Extensions) instructions.In at least one embodiment, 128-bit XMM registers related to SSE2, SSE3, SSE4, AVX, or beyond technologies (commonly referred to as "SSEx") may contain such packed data operands. In at least one embodiment, processors 2010 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inferencing.

[0286] In at least one embodiment, processor 2000 includes an in-order front-end ("front-end") 2001 to fetch instructions to be executed and prepare instructions to be used later in the processor pipeline. In at least one embodiment, front-end 2001 may include multiple units. In at least one embodiment, an instruction prefetcher 2026 fetches instructions from memory and passes them to an instruction decoder 2028, which in turn decodes or interprets instructions. For example, in at least one embodiment, instruction decoder 2028 decodes a received instruction into one or more operations, referred to as "micro-instructions" or "micro-operations" (also called "micro-ops" or "uops"), that may be executed by the machine.In at least one embodiment, instruction decoder 2028 decomposes the instruction into opcode and corresponding data and control fields that may be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, a trace cache 2030 may assemble decoded uops into program-ordered sequences or traces in a uop queue 2034 for execution. In at least one embodiment, when trace cache 2030 encounters a complex instruction, a microcode ROM 2032 provides the uops required to complete the operation.

[0287] In at least one embodiment, some instructions may be converted into a single micro-op, while others may require multiple micro-ops to fully complete operation. In at least one embodiment, if an instruction requires more than four micro-ops to execute, the instruction decoder 2028 may access the microcode ROM 2032 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops for processing in the instruction decoder 2028. In at least one embodiment, an instruction may be stored in the microcode ROM 2032 if a number of micro-ops are required to perform the operation.In at least one embodiment, the trace cache 2030 refers to a programmable logic array ("PLA") as an entry point to determine a correct microinstruction pointer for reading microcode sequences to complete one or more instructions from the microcode ROM 2032. In at least one embodiment, after the microcode ROM 2032 finishes sequencing microinstructions for an instruction, the machine front end 2001 may resume fetching microinstructions from the trace cache 2030.

[0288] In at least one embodiment, the out-of-order execution engine ("out-of-order engine") 2003 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic includes a series of buffers to smooth and reorder the flow of instructions to optimize performance as they traverse the pipeline and are scheduled for execution. The out-of-order execution engine 2003 includes, without limitation, an allocator / register renamer 2040, a memory uop queue 2042, an integer / floating point uop queue 2044, a memory scheduler 2046, a fast scheduler 2002, a slow / general FP scheduler 2004, and a simple FP scheduler 2006.In at least one embodiment, the fast scheduler 2002, the slow / general floating-point scheduler 2004, and the simple floating-point scheduler 2006 are also collectively referred to herein as "uop schedulers 2002, 2004, 2006." In at least one embodiment, the allocator / register renamer 2040 allocates machine buffers and resources required by each uop for its execution. In at least one embodiment, the allocator / register renamer 2040 renames logical registers to entries in a register file. In at least one embodiment, the allocator / register renamer 2040 also assigns each uop an entry in one of two uop queues, the memory uop queue 2042 for memory operations and the integer / floating point uop queue 2044 for non-memory operations, prior to the memory scheduler 2046 and the uop schedulers 2002, 2004, 2006.In at least one embodiment, the uop schedulers 2002, 2004, 2006 determine when a uop is ready to execute based on the readiness of their dependent input register operand sources and the availability of the execution resources the uops require to complete their operation. In at least one embodiment, the fast scheduler 2002 may schedule every half of the main clock cycle, while the slow / general floating-point scheduler 2004 and the simple floating-point scheduler 2006 may schedule once per main processor clock cycle. In at least one embodiment, the uop schedulers 2002, 2004, 2006 arbitrate for dispatch ports to schedule uops for execution.

[0289] In at least one embodiment, execution block b11 includes, without limitation, an integer register file / bypass network 2008, a floating-point register file / bypass network ("an FP register file / bypass network") 2010, address generation units ("AGUs") 2012 and 2014, fast arithmetic logic units (ALUs) ("fast ALUs") 2016 and 2018, a slow arithmetic logic unit ("slow ALU") 2020, a floating-point ALU ("FP") 2022, and a floating-point move unit ("FP move") 2024. In at least one embodiment, an integer register file / bypass network 2008 and a floating-point register file / bypass network 2010 are also referred to herein as "register files 2008, 2010."In at least one embodiment, the AGUSs 2012 and 2014, the fast ALUs 2016 and 2018, the slow ALU 2020, the floating-point ALU 2022, and the floating-point movement unit 2024 are also referred to herein as "execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024." In at least one embodiment, the execution block b11 may include, without limitation, any number (including zero) and type of register files, bypass networks, address generation units, and execution units in any combination.

[0290] In at least one embodiment, register files 2008, 2010 may be located between uop schedulers 2002, 2004, 2006 and execution units 2012, 2014, 2016, 2018, 2020, 2022, and 2024. In at least one embodiment, integer register file / bypass network 2008 performs integer operations. In at least one embodiment, floating-point register file / bypass network 2010 performs floating-point operations. In at least one embodiment, each of register files 2008, 2010 may include, without limitation, a bypass network that may redirect or forward just-completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register files 2008, 2010 may exchange data with each other.In at least one embodiment, the integer register file / bypass network 2008 may include, without limitation, two separate register files, one register file for 32 bits of low-order data and a second register file for 32 bits of high-order data. In at least one embodiment, the floating-point register file / bypass network 2010 may include, without limitation, 128-bit wide entries, as floating-point instructions typically have operands 64 to 128 bits wide.

[0291] In at least one embodiment, execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024 may execute instructions. In at least one embodiment, register files 2008, 2010 store integer and floating-point data operand values ​​required to execute microinstructions. In at least one embodiment, processor 2000 may include, without limitation, any number and combination of execution units 2012, 2014, 2016, 2018, 2020, 2022, 2024. In at least one embodiment, floating-point ALU 2022 and floating-point move unit 2024 may perform floating-point, MMX, SIMD, AVX, and SSE or other operations, including special machine learning instructions. In at least one embodiment, the floating-point ALU 2022 may include, without limitation, a 64-bit by 64-bit floating-point divider to perform division, square root, and remainder micro-operations.In at least one embodiment, instructions involving a floating-point value may be processed with floating-point hardware. In at least one embodiment, ALU operations may be forwarded to fast ALUs 2016, 2018. In at least one embodiment, the fast ALUs 2016, 2018 may perform fast operations with an effective latency of half a clock cycle. In at least one embodiment, most complex integer operations go to the slow ALU 2020, as the slow ALU 2020 may include, without limitation, integer execution hardware for long-latency operations, such as a multiplier, shift units, flag logic, and branch processing. In at least one embodiment, memory load / store operations may be performed by ALUs 2012, 2014.In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may perform integer operations on 64-bit data operands. In at least one embodiment, the fast ALU 2016, the fast ALU 2018, and the slow ALU 2020 may be implemented to support a variety of data bit sizes, including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating-point ALU 2022 and the floating-point movement unit 2024 may be implemented to support a range of operands with different bit widths. In at least one embodiment, the floating-point ALU 2022 and the floating-point movement unit 2024 may operate with 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0292] In at least one embodiment, the uop schedulers 2002, 2004, 2006 initiate dependent operations before the parent load completes execution. In at least one embodiment, because uops may be speculatively scheduled and executed in the processor 2000, the processor 2000 may also include logic to handle memory errors. In at least one embodiment, if a data load into the data cache fails, there may be dependent operations in the pipeline that have exited the scheduler with temporarily incorrect data. In at least one embodiment, a retry mechanism tracks the instructions that use incorrect data and reexecutes them. In at least one embodiment, dependent operations may be required to reexecute while independent operations are allowed to complete.In at least one embodiment, schedulers and a retry mechanism of at least one embodiment of a processor may also be configured to intercept instruction sequences for text string comparison operations.

[0293] In at least one embodiment, the term "registers" may refer to internal processor memory locations that can be used as part of instructions to identify operands. In at least one embodiment, the registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, the registers may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform functions described herein. In at least one embodiment, the registers described herein may be implemented by circuitry within a processor using any number of different techniques, such as:Dedicated physical registers, dynamically allocated physical registers using register renaming, combinations of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. In at least one embodiment, a register file also includes eight multimedia SIMD registers for packed data.

[0294] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0295] Fig.21 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 2100 includes one or more processors 2102 and one or more graphics processors 2108, and may be a single-processor desktop system, a multiprocessor workstation system, or a server system with a large number of processors 2102 or processor cores 2107. In at least one embodiment, system 2100 is a processing platform integrated into a system-on-a-chip (SoC) integrated circuit for use in mobile, wearable, or embedded devices.

[0296] In at least one embodiment, system 2100 may include or be integrated with a server-based gaming platform, a gaming console, including a gaming and media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, system 2100 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, processing system 2100 may also include, be coupled to, or integrated with a wearable device, such as a wearable device for a smart watch, smart glasses, an augmented reality device, or a virtual reality device.In at least one embodiment, the processing system 2100 is a television or set-top box device having one or more processors 2102 and a graphical interface generated by one or more graphics processors 2108.

[0297] In at least one embodiment, one or more processors 2102 each include one or more processor cores 2107 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of one or more processor cores 2107 is configured to process a particular instruction set 2109. In at least one embodiment, the instruction set 2109 may enable Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or Very Long Instruction Word (VLIW) computing. In at least one embodiment, the processor cores 2107 may each process a different instruction set 2109, which may include instructions to facilitate emulation of other instruction sets.In at least one embodiment, the processor core 2107 may also include other processing devices, such as a digital signal processor (DSP).

[0298] In at least one embodiment, processor 2102 includes a cache memory 2104. In at least one embodiment, processor 2102 may include a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared by various components of processor 2102. In at least one embodiment, processor 2102 also uses an external cache (e.g., a Level 3 (L3) cache or Last Level Cache (LLC)) (not shown), which may be shared by processor cores 2107 using known cache coherence techniques. In at least one embodiment, a register file 2106 is additionally present in processor 2102, which may include various types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and an instruction pointer register).In at least one embodiment, register file 2106 may include general-purpose registers or other registers.

[0299] In at least one embodiment, one or more processors 2102 are coupled to one or more interface buses 2110 to communicate communication signals, such as address, data, or control signals, between processor 2102 and other components in system 2100. In at least one embodiment, interface bus 2110 may be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, interface 2110 is not limited to a DMI bus and may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, processor(s) 2102 include an integrated memory controller 2116 and a platform control hub 2130.In at least one embodiment, the memory controller 2116 facilitates communication between a memory device and other components of the system 2100, while the platform controller hub (PCH) 2130 provides connections to I / O devices via a local I / O bus.

[0300] In at least one embodiment, memory device 2120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or other memory device with suitable performance to serve as process memory. In at least one embodiment, memory device 2120 may operate as system memory for system 2100 to store data 2122 and instructions 2121 for use when one or more processors 2102 are executing an application or process. In at least one embodiment, memory controller 2116 is also coupled to an optional external graphics processor 2112 that can communicate with one or more graphics processors 2108 in processors 2102 to perform graphics and media operations.In at least one embodiment, a display device 2111 may be connected to the processor(s) 2102. In at least one embodiment, the display device 2111 may comprise one or more internal display devices, such as in a mobile electronic device or a laptop, or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 2111 may comprise a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.

[0301] In at least one embodiment, the platform control hub 2130 enables the connection of peripherals to the storage device 2120 and the processor 2102 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, among others, an audio controller 2146, a network controller 2134, a firmware interface 2128, a wireless transceiver 2126, touch sensors 2125, and a data storage device 2124 (e.g., hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2124 may be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a Peripheral Component Interconnect Bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensors 2125 may include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, the wireless transceiver 2126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a cellular transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2128 enables communication with the system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2134 may enable network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2110. In at least one embodiment, the audio controller 2146 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2100 includes an optional legacy I / O controller 2140 for coupling legacy devices (e.g.,Personal System 2 (PS / 2)) interfaces with the system. In at least one embodiment, the platform control hub 2130 may also be connected to one or more Universal Serial Bus (USB) controllers 2142 that connect input devices such as keyboard and mouse combinations 2143, a camera 2144, or other USB input devices.

[0302] In at least one embodiment, an instance of the memory controller 2116 and the platform control hub 2130 may be integrated into a discrete external graphics processor, such as the external graphics processor 2112. In at least one embodiment, the platform control hub 2130 and / or the memory controller 2116 may be external to one or more processors 2102. For example, in at least one embodiment, the system 2100 may include an external memory controller 2116 and a platform control hub 2130, which may be configured as a memory control hub and peripheral control hub within a system chipset in communication with the processor(s) 2102.

[0303] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0304] Fig.22 is a block diagram of a processor 2200 having one or more processor cores 2202A-2202N, an integrated memory controller 2214, and an integrated graphics processor 2208, according to at least one embodiment. In at least one embodiment, the processor 2200 may include additional cores, up to and including the additional core 2202N, represented by dashed boxes. In at least one embodiment, each of the processor cores 2202A-2202N includes one or more internal cache units 2204A-2204N. In at least one embodiment, each processor core also has access to one or more shared cache units 2206.

[0305] In at least one embodiment, the internal cache units 2204A-2204N and the shared cache units 2206 represent a cache hierarchy within the processor 2200. In at least one embodiment, the cache units 2204A-2204N may include at least one level of instruction and data cache within each processor core and one or more levels of shared mid-level cache, such as a Level 2 (L2), Level 3 (L3), Level 4 (L4), or other cache levels, with a highest cache level prior to external memory classified as LLC. In at least one embodiment, the cache coherence logic maintains coherence between different cache units 2206 and 2204A-2204N.

[0306] In at least one embodiment, the processor 2200 may also include a set of one or more bus control units 2216 and a system agent core 2210. In at least one embodiment, one or more bus control units 2216 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, the system agent core 2210 provides management functions for various processor components. In at least one embodiment, the system agent core 2210 includes one or more integrated memory controllers 2214 to manage access to various external memory devices (not shown).

[0307] In at least one embodiment, one or more of the processor cores 2202A-2202N include support for concurrent multithreading. In at least one embodiment, the system agent core 2210 includes components for coordinating and operating the cores 2202A-2202N during multithreaded processing. In at least one embodiment, the system agent core 2210 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of the processor cores 2202A-2202N and the graphics processor 2208.

[0308] In at least one embodiment, processor 2200 additionally includes a graphics processor 2208 for performing graphics processing operations. In at least one embodiment, graphics processor 2208 is coupled to shared cache units 2206 and system agent core 2210, which includes one or more integrated memory controllers 2214. In at least one embodiment, system agent core 2210 also includes a display controller 2211 for controlling the output of the graphics processor to one or more coupled displays. In at least one embodiment, display controller 2211 may also be a separate module connected to graphics processor 2208 via at least one interconnect, or it may be integrated into graphics processor 2208.

[0309] In at least one embodiment, a ring-based interconnect 2212 is used to connect internal components of processor 2200. In at least one embodiment, an alternative interconnect may be used, such as a point-to-point connection, a switched connection, or other techniques. In at least one embodiment, graphics processor 2208 is connected to ring interconnect 2212 via an I / O connection 2213.

[0310] In at least one embodiment, I / O connection 2213 represents at least one of several types of I / O connections, including an on-package I / O connection that enables communication between various processor components and an embedded high-performance memory module 2218, such as an eDRAM module. In at least one embodiment, each of the processor cores 2202A-2202N and the graphics processor 2208 use embedded memory modules 2218 as a shared last-level cache.

[0311] In at least one embodiment, processor cores 2202A-2202N are homogeneous cores executing a common instruction set architecture. In at least one embodiment, processor cores 2202A-2202N are instruction set architecture (ISA) heterogeneous, where one or more processor cores 2202A-2202N execute a common instruction set, while one or more other cores of processor cores 2202A-2202N execute a subset of a common instruction set or a different instruction set. In at least one embodiment, processor cores 2202A-2202N are microarchitecturally heterogeneous, where one or more relatively higher power cores are coupled with one or more lower power cores. In at least one embodiment, processor 2200 may be implemented on one or more chips or as an integrated circuit (SoC).

[0312] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0313] Fig.23 is a block diagram of a graphics processor 2300, which may be a discrete graphics processing unit or a graphics processor integrated with a plurality of processor cores. In at least one embodiment, the graphics processor 2300 communicates with registers on the graphics processor 2300 and with instructions stored in memory via an I / O interface associated with memory. In at least one embodiment, the graphics processor 2300 includes a memory interface 2314 for accessing memory. In at least one embodiment, the memory interface 2314 is an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.

[0314] In at least one embodiment, graphics processor 2300 also includes a display controller 2302 for controlling display output data to a display device 2320. In at least one embodiment, display controller 2302 includes hardware for one or more overlay layers for display device 2320 and composing multiple layers of video or user interface elements. In at least one embodiment, display device 2320 may be an internal or external display device. In at least one embodiment, display device 2320 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device.In at least one embodiment, graphics processor 2300 includes a video codec engine 2306 to encode, decode, or transcode media to, from, or between one or more media coding formats, including, but not limited to, Moving Picture Experts Group (MPEG) formats such as MPEG-2, Advanced Video Coding (AVC) formats such as H.264 / MPEG-4 AVC, as well as Society of Motion Picture & Television Engineers (SMPTE) 421MNC-1 and Joint Photographic Experts Group (JPEG) formats such as JPEG and Motion JPEG (MJPEG) formats.

[0315] In at least one embodiment, graphics processor 2300 includes a BLIT (Block Image Transfer) engine 2304 for performing two-dimensional (2D) rasterization operations, including, for example, bit-boundary block transfers. However, in at least one embodiment, 2D graphics operations are performed with one or more components of graphics processing engine (GPE) 2310. In at least one embodiment, GPE 2310 is a compute engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.

[0316] In at least one embodiment, GPE 2310 includes a 3D pipeline 2312 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions that act on 3D primitives (e.g., rectangle, triangle, etc.). 3D pipeline 2312 includes programmable and fixed function elements that perform various tasks and / or spawn execution threads to a 3D / media subsystem 2315. While 3D pipeline 2312 can be used to perform media operations, in at least one embodiment, GPE 2310 also includes a media pipeline 2316 used to perform media operations, such as video post-processing and image enhancement.

[0317] In at least one embodiment, media pipeline 2316 includes fixed functional or programmable logic units to perform one or more specialized media operations, such as video decoding acceleration, video descrambling, and video encoding acceleration, instead of or on behalf of video codec engine 2306. In at least one embodiment, media pipeline 2316 additionally includes a thread spawning unit to spawn threads for execution in 3D / media subsystem 2315. In at least one embodiment, the spawned threads perform computations for media operations on one or more graphics execution units present in 3D / media subsystem 2315.

[0318] In at least one embodiment, the 3D / Media subsystem 2315 includes logic for executing threads spawned by the 3D pipeline 2312 and the media pipeline 2316. In at least one embodiment, the 3D pipeline 2312 and the media pipeline 2316 send thread execution requests to the 3D / Media subsystem 2315, which includes thread dispatch logic to arbitrate and distribute various requests to available thread execution resources. In at least one embodiment, the execution resources include an array of graphics execution units for processing 3D and media threads. In at least one embodiment, the 3D / Media subsystem 2315 includes one or more internal caches for thread instructions and data.In at least one embodiment, subsystem 2315 also includes shared memory, including registers and addressable memory, to share data between threads and store output data.

[0319] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0320] Fig. 24 is a block diagram of a graphics processing engine 2410 of a graphics processor according to at least one embodiment. In at least one embodiment, the graphics processing engine (GPE) 2410 is a version of the Fig.23. In at least one embodiment, the media pipeline 2416 is optional and may not be explicitly present in the GPE 2410. In at least one embodiment, a separate media and / or image processor is coupled to the GPE 2410.

[0321] In at least one embodiment, the GPE 2410 is coupled to or includes an instruction streamer 2403 that provides an instruction stream to the 3D pipeline 2412 and / or the media pipelines 2416. In at least one embodiment, the instruction streamer 2403 is coupled to memory, which may be system memory or one or more internal caches and shared caches. In at least one embodiment, the instruction streamer 2403 receives instructions from memory and sends instructions to the 3D pipeline 2412 and / or the media pipeline 2416. In at least one embodiment, the instructions are instructions, primitives, or micro-operations fetched from a circular buffer that stores instructions for the 3D pipeline 2412 and the media pipeline 2416. In at least one embodiment, a ring buffer may additionally include batch instruction buffers that store batches of multiple instructions.In at least one embodiment, the instructions for the 3D pipeline 2412 may also include references to data stored in memory, such as vertex and geometry data for the 3D pipeline 2412 and / or image data and memory objects for the media pipeline 2416. In at least one embodiment, the 3D pipeline 2412 and the media pipeline 2416 process instructions and data by performing operations or forwarding one or more threads of execution to a graphics core array 2414. In at least one embodiment, the graphics core array 2414 includes one or more blocks of graphics cores (e.g., graphics core(s) 2415A, graphics core(s) 2415B), each block including one or more graphics cores.In at least one embodiment, each graphics core comprises a set of graphics execution resources, including general-purpose and graphics-specific execution logic for performing graphics and compute operations, as well as fixed-function texture processing logic and / or acceleration logic for machine learning and artificial intelligence.

[0322] In at least one embodiment, the 3D pipeline 2412 includes fixed function and programmable logic to process one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 2414. In at least one embodiment, the graphics core array 2414 provides a unified block of execution resources for processing shader programs. In at least one embodiment, the general-purpose execution logic (e.g., execution units) in the graphics cores 2415A-2415B of the graphics core array 2414 includes support for various 3D API shader languages ​​and can execute multiple concurrent execution threads associated with multiple shaders.

[0323] In at least one embodiment, the graphics core assembly 2414 also includes execution logic for performing media functions such as video and / or image processing. In at least one embodiment, the execution units additionally include general-purpose logic programmable to perform parallel general-purpose computation operations in addition to the graphics processing operations.

[0324] In at least one embodiment, output data generated by threads executing on the graphics core array 2414 may be output to memory in a unified return buffer (URB) 2418. The URB 2418 may store data for multiple threads. In at least one embodiment, the URB 2418 may be used to send data between different threads executing on the graphics core array 2414. In at least one embodiment, the URB 2418 may additionally be used for synchronization between threads on the graphics core array 2414 and the fixed functional logic within the shared functional logic 2420.

[0325] In at least one embodiment, the graphics core array 2414 is scalable such that the graphics core array 2414 includes a variable number of graphics cores, each of which has a variable number of execution units based on a target power and performance level of the GPE 2410. In at least one embodiment, the execution resources are dynamically scalable such that the execution resources can be enabled or disabled as needed.

[0326] In at least one embodiment, the graphics core assembly 2414 is coupled to shared functional logic 2420, which includes a plurality of resources shared by the graphics cores within the graphics core assembly 2414. In at least one embodiment, the shared functions performed by the shared functional logic 2420 are embodied in hardware logic units that provide specific additional functionality to the graphics core assembly 2414. In at least one embodiment, the shared functional logic 2420 includes, among other things, a sampler 2421, math 2422, and inter-thread communication (ITC) 2423 logic. In at least one embodiment, one or more caches 2425 are included in or coupled to the shared functional logic 2420.

[0327] In at least one embodiment, a shared function is used when the demand for a specialized function is insufficient to include it in the graphics core assembly 2414. In at least one embodiment, a single instantiation of a specialized function is used in the shared function logic 2420 and shared by other execution resources within the graphics core assembly 2414. In at least one embodiment, certain shared functions within the shared function logic 2420 that are heavily used by the graphics core assembly 2414 may be present in the shared function logic 2416 within the graphics core assembly 2414. In at least one embodiment, the shared function logic 2416 within the graphics core assembly 2414 may include some or all of the logic of the shared function logic 2420.In at least one embodiment, all logic elements within shared functional logic 2420 may be duplicated within shared functional logic 2416 of graphics core assembly 2414. In at least one embodiment, shared functional logic 2420 is excluded in favor of shared functional logic 2416 within graphics core assembly 2414.

[0328] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0329] Fig.25 is a block diagram of the hardware logic of a graphics processor core 2500, as described herein in at least one embodiment. In at least one embodiment, the graphics processor core 2500 is present in a graphics core array. In at least one embodiment, the graphics processor core 2500, sometimes referred to as a core slice, may be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 2500 is an example of a graphics core slice, and a graphics processor as described herein may have multiple graphics core slices based on targeted power and performance envelopes. In at least one embodiment, each graphics core 2500 may have a fixed functional block 2530 coupled to a plurality of sub-cores 2501A-2501F, also referred to as sub-cores orSub-slices are modular blocks with general-purpose and fixed-function logic.

[0330] In at least one embodiment, fixed function block 2530 includes a geometry / fixed function pipeline 2536 that may be shared by all subcores in graphics processor 2500, e.g., in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry / fixed function pipeline 2536 includes a 3D fixed function pipeline, a video front-end unit, a thread spawner and thread dispatcher, and a unified return buffer manager that manages unified return buffers.

[0331] In at least one embodiment, fixed functional block 2530 also includes a graphics SoC interface 2537, a graphics microcontroller 2538, and a media pipeline 2539. Graphics SoC interface 2537 provides an interface between graphics core 2500 and other processor cores within a system-on-chip integrated circuit. In at least one embodiment, graphics microcontroller 2538 is a programmable subprocessor that can be configured to manage various functions of graphics processor 2500, including thread dispatch, scheduling, and preemption. In at least one embodiment, media pipeline 2539 includes logic to facilitate decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data.In at least one embodiment, the media pipeline 2539 implements media operations via requests to compute or sensing logic within the subcores 2501-2501F.

[0332] In at least one embodiment, the SoC interface 2537 enables the graphics core 2500 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within an SoC, including memory hierarchy elements such as a shared last-level cache, system RAM, and / or embedded on-chip or on-package DRAM. In at least one embodiment, the SoC interface 2537 may also enable communication with fixed-function devices within an SoC, such as camera imaging pipelines, and enables the use and / or implementation of global memory atoms that may be shared between the graphics core 2500 and CPUs within an SoC.In at least one embodiment, the SoC interface 2537 may also implement power management controls for the graphics core 2500 and enable an interface between a clock domain of the graphics core 2500 and other clock domains within an SoC. In at least one embodiment, the SoC interface 2537 enables receipt of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within a graphics processor. In at least one embodiment, commands and instructions may be sent to the media pipeline 2539 when media operations are to be performed or to a geometry and fixed function pipeline (e.g., geometry and fixed function pipeline 2536, geometry and fixed function pipeline 2514) when graphics processing operations are to be performed.

[0333] In at least one embodiment, graphics microcontroller 2538 may be configured to perform various scheduling and management tasks for graphics core 2500. In at least one embodiment, graphics microcontroller 2538 may perform scheduling of graphics and / or compute tasks on various parallel graphics engines within arrays 2502A-2502F, 2504A-2504F of execution units (EUs) within subcores 2501A-2501F. In at least one embodiment, host software executing on a CPU core of an SoC including graphics core 2500 may submit workloads to one of several graphics processor doorbells, which invokes a scheduling operation on an appropriate graphics engine.In at least one embodiment, the scheduling operations include determining the next workload to execute, submitting a workload to an instruction streamer, preempting existing workloads running on a machine, monitoring the progress of a workload, and notifying host software upon completion of a workload. In at least one embodiment, the graphics microcontroller 2538 may also facilitate low-power or idle states for the graphics core 2500 by providing the graphics core 2500 with the ability to save and restore registers within the graphics core 2500 via low-power state transitions independent of an operating system and / or graphics driver software on a system.

[0334] In at least one embodiment, the graphics core 2500 may include more or fewer sub-cores than the illustrated sub-cores 2501A-2501F, up to N modular sub-cores. In at least one embodiment, the graphics core 2500 may also include, for each set of N sub-cores, shared functional logic 2510, shared and / or cache memory 2512, a geometry / fixed function pipeline 2514, and additional fixed function logic 2516 to accelerate various graphics and compute processing operations. In at least one embodiment, the shared functional logic 2510 may include logical units (e.g., samplers, math, and / or inter-thread communication logic) that may be shared by each of the N sub-cores within the graphics core 2500.Shared and / or cache memory 2512 may be a last-level cache for N sub-cores 2501A-2501F within graphics core 2500 and may also serve as shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 2514 may be present within fixed function block 2530 instead of geometry / fixed function pipeline 2536 and may include the same or similar logic units.

[0335] In at least one embodiment, the graphics core 2500 includes additional fixed function logic 2516, which may include various fixed function acceleration logic for use by the graphics core 2500. In at least one embodiment, the additional fixed function logic 2516 includes an additional geometry pipeline for use in positional shading. In positional shading, there are at least two geometry pipelines: a full geometry pipeline within the geometry / fixed function pipeline 2516, 2536, and a cull pipeline, which is an additional geometry pipeline and in which additional fixed function logic 2516 may be included. In at least one embodiment, the cull pipeline is a stripped-down version of a full geometry pipeline.In at least one embodiment, a full pipeline and a cull pipeline may execute different instances of an application, with each instance having its own context. In at least one embodiment, positional shading may hide long cull runs from discarded triangles, allowing shading to complete sooner in some embodiments. For example, in at least one embodiment, the cull pipeline logic within the additional fixed-function logic 2516 may execute position shaders in parallel with a main application and generally generates critical results faster than a full pipeline because the cull pipeline retrieves and shades vertices' position attributes without performing rasterization and rendering pixels into a frame buffer.In at least one embodiment, the cull pipeline may use the generated critical results to compute visibility information for all triangles, regardless of whether those triangles are culled. In at least one embodiment, the full pipeline (which in this case may be referred to as a retry pipeline) may use visibility information to skip culled triangles and shade only visible triangles, which are ultimately passed to a rasterization phase.

[0336] In at least one embodiment, the additional fixed function logic 2516 may also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementations that include optimizations for machine learning training or inferencing.

[0337] In at least one embodiment, each graphics sub-core 2501A-2501F includes a set of execution resources that can be used to perform graphics, media, and compute operations in response to requests from graphics pipeline, media pipeline, or shader programs. In at least one embodiment, the graphics sub-cores 2501A-2501F include a plurality of EU arrays 2502A-2502F, 2504A-2504F, thread dispatch and inter-thread communication logic (TD / IC) 2503A-2503F, a 3D sampler (e.g., texture) 2505A-2505F, a media sampler 2506A-2506F, a shader processor 2507A-2507F, and shared local memory (SLM) 2508A-2508F.The EU devices 2502A-2502F, 2504A-2504F each include a plurality of execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logic operations on a graphics, media, or compute operation, including graphics, media, or compute shader programs. In at least one embodiment, the TDIIC logic 2503A-2503F performs local thread dispatch and thread control operations for execution units within a subcore and facilitates communication between threads executing on execution units of a subcore. In at least one embodiment, the 3D sampler 2505A-2505F can read texture or other 3D graphics data into memory.In at least one embodiment, the 3D sampler may read texture data differently based on a configured sampling state and a texture format associated with a particular texture. In at least one embodiment, the media sampler 2506A-2506F may perform similar read operations based on a type and format associated with the media data. In at least one embodiment, each graphics subcore 2501A-2501F may alternately include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each of the subcores 2501A-2501F may utilize the shared local memory 2508A-2508F within each subcore to enable threads executing within a thread group to execute using a common pool of on-chip memory.

[0338] In at least one embodiment, at least one component shown or described with respect to the figure described above is used to implement methods and / or functions associated with one or more of the Fig. 1-5 to perform ray tracing.

[0339] Fig. 26A and Fig. 26B illustrate thread execution logic 2600 comprising an arrangement of processing elements of a graphics processor core according to at least one embodiment. Fig. 26A illustrates at least one embodiment in which thread execution logic 2600 is used. Fig. 26B illustrates exemplary internal details of an execution unit according to at least one embodiment.

[0340] As it is in Fig.26A, in at least one embodiment, thread execution logic 2600 includes a shader processor 2602, a thread dispatcher 2604, an instruction cache 2606, a scalable execution unit array having a plurality of execution units 2608A-2608N, a sampler 2610, a data cache 2612, and a data port 2614. In at least one embodiment, a scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., one of execution units 2608A, 2608B, 2608C, 2608D through 2608N-1, and 2608N) based on the computational requirements of a workload. In at least one embodiment, the scalable execution units are interconnected via an interconnect fabric that connects to each execution unit.In at least one embodiment, thread execution logic 2600 includes one or more connections to memory, such as system memory or cache memory, via one or more of the following: instruction cache 2606, data port 2614, sampler 2610, and execution units 2608A-2608N. In at least one embodiment, each execution unit (e.g., 2608A) is a standalone, general-purpose programmable compute unit capable of executing multiple concurrent hardware threads, processing multiple data elements in parallel for each thread. In at least one embodiment, the arrangement of execution units 2608A-2608N is scalable to include any number of individual execution units.

[0341] In at least one embodiment, execution units 2608A-2608N are primarily used to execute shader programs. In at least one embodiment, shader processor 2602 may process various shader programs and dispatch the execution threads associated with the shader programs via a thread dispatcher 2604. In at least one embodiment, thread dispatcher 2604 includes logic to mediate thread initiation requests from graphics and media pipelines and instantiate requested threads on one or more execution units within execution units 2608A-2608N. In at least one embodiment, a geometry pipeline may, for example, forward vertex, tessellation, or geometry shaders to the thread execution logic for processing.In at least one embodiment, the thread dispatcher 2604 may also process runtime thread creation requests from executing shader programs.

[0342] In at least one embodiment, execution units 2608A-2608N support an instruction set that includes native support for many standard 3D graphics shader instructions, allowing shader programs from graphics libraries (e.g., Direct 3D and OpenGL) to execute with minimal translation. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general processing (e.g., compute and media shaders). In at least one embodiment, each of execution units 2608A-2608N, including one or more arithmetic logic units (ALUs), is capable of single instruction multiple data (SIMD) execution, and multi-threaded operation enables an efficient execution environment despite higher latency memory accesses.In at least one embodiment, each hardware thread within each execution unit has its own high-bandwidth register file and associated independent thread state. In at least one embodiment, execution occurs with multiple threads per clock on pipelines capable of performing integer, floating-point, and double-precision operations, SIMD branch capability, logical operations, transcendental operations, and other miscellaneous operations. In at least one embodiment, dependency logic in execution units 2608A-2608N causes a waiting thread to sleep until the requested data is returned while waiting for data from memory or one of the shared functions. In at least one embodiment, while a waiting thread is sleeping, hardware resources may be used to process other threads.For example, in at least one embodiment, during a delay associated with a vertex shader operation, an execution unit may perform operations on a pixel shader, fragment shader, or other type of shader program that includes a different vertex shader.

[0343] In at least one embodiment, each execution unit in execution units 2608A-2608N operates on arrays of data elements. In at least one embodiment, a number of data elements is the "execution size" or the number of channels for an instruction. In at least one embodiment, an execution channel is a logical execution unit for accessing data elements, masking, and flow control within instructions. In at least one embodiment, the number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) for a particular graphics processor. In at least one embodiment, execution units 2608A-2608N support integer and floating-point data types.

[0344] In at least one embodiment, the instruction set of an execution unit comprises SIMD instructions. In at least one embodiment, different data elements may be stored as a packed data type in a register, and the execution unit processes different elements based on the data size of the elements. For example, in at least one embodiment, when operating on a 256-bit wide vector, 256 bits of a vector are stored in a register, and an execution unit processes a vector as four separate packed 64-bit data elements (quad-word (QW) data elements), eight separate packed 32-bit data elements (double-word (DW) data elements), sixteen separate packed 16-bit data elements (word (W) data elements), or thirty-two separate 8-bit data elements (byte (B) data elements).However, in at least one embodiment, other vector widths and register sizes are also possible.

[0345] In at least one embodiment, one or more execution units may be combined into a fused execution unit 2609A-2609N with thread control logic (2607A-2607N) common to the fused EUs. In at least one embodiment, multiple EUs may be fused into an EU group. In at least one embodiment, each EU in a fused EU group may be configured to execute a separate SIMD hardware thread. The number of EUs in a fused EU group may vary depending on the embodiment. In at least one embodiment, various SIMD widths may be executed per EU, including, but not limited to, SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 2609A-2609N has at least two execution units.For example, in at least one embodiment, the fused execution unit 2609A includes a first EU 2608A, a second EU 2608B, and thread control logic 2607A common to the first EU 2608A and the second EU 2608B. In at least one embodiment, the thread control logic 2607A controls threads executing on the fused graphics execution unit 2609A such that each EU within the fused execution units 2609A-2609N can execute using a common instruction pointer register.

[0346] In at least one embodiment, thread execution logic 2600 includes one or more internal instruction caches (e.g., 2606) to cache thread instructions for execution units. In at least one embodiment, one or more data caches (e.g., 2612) are present to cache thread data during thread execution. In at least one embodiment, a sampler 2610 is present to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, sampler 2610 includes specialized texture or media sampling functionality to process texture or media data during the sampling process before passing the sampled data to an execution unit.

[0347] In at least one embodiment, during execution, graphics and media pipelines send thread initiation requests to thread execution logic 2600 via thread creation and dispatch logic. In at least one embodiment, once a group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 2602 is invoked to further compute output information and cause the results to be written to output surfaces (e.g., color buffers, depth buffers, stencil buffers, etc.). In at least one embodiment, a pixel shader or fragment shader computes the values ​​of various vertex attributes to be interpolated over a rasterized object.In at least one embodiment, pixel processor logic within shader processor 2602 then executes a pixel or fragment shader program provided via an application programming interface (API). In at least one embodiment, shader processor 2602 dispatches threads to an execution unit (e.g., 2608A) via thread dispatcher 2604 to execute a shader program. In at least one embodiment, shader processor 2602 uses texture sampling logic in sampler 2610 to access texture data in texture maps stored in memory. In at least one embodiment, arithmetic operations on texture data and input geometry data calculate pixel color data for each geometric fragment or exclude one or more pixels from further processing.

[0348] In at least one embodiment, data port 2614 provides a memory access mechanism for thread execution logic 2600 to output processed data to memory for further processing on a graphics processor output pipeline. In at least one embodiment, data port 2614 includes or is coupled to one or more caches (e.g., data cache 2612) to temporarily store data for memory access via a data port.

[0349] As in Fig.26B, in at least one embodiment, a graphics execution unit 2608 may include an instruction fetch unit 2637, a general register file array (GRF) 2624, an architectural register file array (ARF) 2626, a thread dispatcher 2622, a dispatch unit 2630, a branch unit 2632, a set of SIMD floating-point units (FPUs) 2634, and, in at least one embodiment, a set of dedicated integer SIMD ALUs 2635. In at least one embodiment, the GRF 2624 and the ARF 2626 include a set of general register files and architectural register files associated with each concurrent hardware thread that may be active in the graphics execution unit 2608. In at least one embodiment, per-thread architectural state is managed in the ARF 2626, while data used during thread execution is stored in the GRF 2624.In at least one embodiment, the execution state of each thread, including instruction pointers for each thread, may be maintained in thread-specific registers in the ARF 2626.

[0350] In at least one embodiment, graphics execution unit 2608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on a target number of concurrent threads and the number of registers per execution unit, with the execution unit resources being divided among the logic used to execute multiple concurrent threads.

[0351] In at least one embodiment, the graphics execution unit 2608 may issue multiple instructions together, each of which may be different instructions. In at least one embodiment, the thread dispatcher 2622 of the graphics execution unit 2608 thread may forward instructions to one of the dispatch units 2630, branch units 2642, or SIMD FPU(s) 2634 for execution. In at least one embodiment, each thread may access 128 general-purpose registers within the GRF 2624, where each register may store 32 bytes accessible as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread has access to 4 KB within the GRF 2624, although embodiments are not so limited, and other implementations may provide more or fewer register resources.In at least one embodiment, up to seven threads can execute concurrently, although the number of threads per execution unit can also vary depending on the embodiment. In at least one embodiment where seven threads can access 4 KB, the GRF 2624 can store a total of 28 KB. In at least one embodiment, flexible addressing modes can allow registers to be addressed together to effectively form wider registers or to represent strided rectangular block data structures.

[0352] In at least one embodiment, memory operations, scan operations, and other longer-latency system communications are handled via "send" instructions executed by a message passing send unit 2630. In at least one embodiment, branch instructions are forwarded to a dedicated branch unit 2632 to enable divergence and eventual convergence with respect to SIMD.

[0353] In at least one embodiment, the graphics execution unit 2608 includes one or more SIMD floating-point units (FPU(s)) 2634 to perform floating-point operations. In at least one embodiment, the FPU(s) 2634 also support integer calculations. In at least one embodiment, the FPU(s) 2634 can perform up to M 32-bit floating-point (or integer) operations or up to 2M 16-bit integer or 16-bit floating-point operations with respect to SIMD. In at least one embodiment, at least one of the FPU(s) provides extended mathematical capabilities to support high-throughput transcendental mathematical ...

Claims

[1] Processor comprising: one or more circuits for causing an amount of memory to be used by one or more ray tracing software programs to be identified for a user. [2] The processor of claim 1, wherein the amount of memory is provided to enable one or more rays resulting from execution of the one or more ray tracing software programs to be selected. [3] The processor of claim 1 or 2, wherein the one or more circuits are configured to generate a ray tracing workload prior to ray tracing based at least in part on less than all data associated with rays to be processed. [4] A processor according to any preceding claim, wherein the memory size is calculated based at least in part on information relating to ray intersections in a virtual environment, the ray intersections comprising one or more of specular reflection, diffraction, refraction and diffuse reflection. [5] A processor according to any preceding claim, wherein the amount of memory is calculated using less than all of the data associated with rays to be processed, the less than all of the data corresponding to a minimum amount of data required to calculate one or more numbers of one or more types of secondary rays of a primary beam. [6] Processor according to one of the preceding claims, wherein the one or more circuits are configured: to divide rays to be processed into two or more groups of rays based at least in part on a ray tracing workload and available memory of one or more processing units; and to perform ray tracing of the two or more groups of rays using the one or more processing units. [7] A processor according to any preceding claim, wherein the one or more circuits are configured to perform lightweight ray tracing by processing only one or more numbers and one or more types of secondary rays. [8] System comprising: one or more processors to cause an amount of memory to be used by one or more ray tracing software programs to be identified for a user. [9] The system of claim 8, wherein the amount of memory is provided to enable one or more rays resulting from execution of the one or more ray tracing software programs to be selected. [10] The system of claim 8 or 9, wherein the one or more processors are configured to generate a ray tracing workload prior to ray tracing based at least in part on less than all data associated with rays to be processed. [11] The system of any one of claims 8 to 10, wherein the memory size is calculated at least in part based on information relating to ray intersections in a virtual environment, the ray intersections comprising one or more of specular reflection, diffraction, refraction, and diffuse reflection. [12] A system according to any one of claims 8 to 11, wherein the memory size is calculated using less than all data associated with rays to be processed, the less than all data corresponding to a minimum amount of data required to calculate one or more numbers of one or more types of secondary rays of a primary beam. [13] The system of any one of claims 8 to 12, wherein the one or more processors are configured to: to divide rays to be processed into two or more groups of rays based at least in part on a ray tracing workload and available memory of one or more processing units; and to perform ray tracing of the two or more groups of rays using the one or more processing units. [14] The system of any of claims 8 to 13, wherein the one or more ray tracing software programs are executed in parallel by two or more graphics processing units, GPUs. [15] Procedure comprising: Using one or more circuits to cause an amount of memory to be used by one or more ray tracing software programs to be identified for a user. [16] The method of claim 15, wherein the amount of memory is provided to enable one or more rays resulting from execution of the one or more ray tracing software programs to be selected. [17] A method according to claim 15 or 16, further comprising: Before ray tracing, generate a ray tracing workload based at least in part on less than all of the data associated with rays to be processed. [18] A method according to any one of claims 15 to 17, wherein the storage amount is calculated at least in part based on information relating to ray intersections in a virtual environment, the ray intersections comprising one or more of specular reflection, diffraction, refraction and diffuse reflection. [19] A method according to any one of claims 15 to 18, wherein the memory size is calculated using less than all data associated with rays to be processed, the less than all data corresponding to a minimum amount of data required to calculate one or more numbers of one or more types of secondary rays of a primary beam. [20] A method according to any one of claims 15 to 19, further comprising: Splitting rays to be processed into two or more groups of rays based at least in part on a ray tracing workload and available memory of one or more processing units; and Performing ray tracing of the two or more groups of rays using the one or more processing units.