Multi-format GPU docking board

By designing a multi-format GPU docking board, using the first switch and the second switch to realize efficient communication between the GPU board and the CPU, the problem of GPU board format limitation and low communication efficiency in the prior art is solved, and the system compatibility and performance of AI/ML applications are improved.

CN115836281BActive Publication Date: 2025-05-06NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080102959.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-05-06
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

Existing GPU docking boards usually support a form factor or format connector, which limits the compatibility and application range of the GPU board. Especially in complex communication solutions, the communication efficiency between the CPU and the GPU is inefficient, making it difficult to meet high-performance computing and AI/ML requirements.

Method used

A multi-format GPU docking board is designed, including a first switch and a second switch, which is used to realize communication between GPU boards of different formats and between CPU and GPU boards, and supports GPU board connections of multiple form factors, including peripheral component interconnection fast connectors and NVLINKTM format connectors.

Benefits of technology

It realizes efficient communication between GPU boards of different formats and between CPU and GPU boards, enhances system compatibility and application range, and improves the performance of high-performance computing and AI/ML applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115836281B_ABST
    Figure CN115836281B_ABST
Patent Text Reader

Abstract

A multi-format graphics processing unit (GPU) docking board is disclosed. The multi-format GPU docking board includes a first switch for enabling communication between a first-format GPU board and a second-format GPU board, and a second switch for enabling communication between a central processing unit (CPU) and the first-format GPU board and between the CPU and the second-format GPU board.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment relates to a multi-format graphics processing unit (GPU) docking board. In at least one embodiment, the multi-format GPU docking board includes a first switch for enabling communication between a first format (GPU) board and a second format GPU board, and a second switch for enabling communication between a central processing unit (CPU) and the first format GPU board and between the CPU and the second format GPU board. Background Art

[0002] The GPU docking board may support a form factor or format of connector. The GPU board format may be or have a peripheral component interconnect express The form factor of the connector, or the format can be or have NVLINK TM Format connectors (such as SXM TM or SXM2 TM connector). SXM2 TM The connector is a type of processor module form factor defined by Nvidia that supports Nvidia's NVLINK TM Interconnect standard (also known as Nvidia High Speed ​​Signaling Interconnect or NVHS). NVLINK TM An interconnection standard is a Standard higher performance interconnect standards. For example, compared to Standard, NVLINK1 TM and NVLINK2 TM (NVLINK for short) TM ) interconnection speeds to achieve ultra-high-speed data transmission. For example, NVLINK1 TM The interconnection speed can reach 20GB / s, NVLINK2 TM The interconnection speed can reach 25GB / s. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:

[0004] Figure 1 is a system diagram of an example GPU docking board and environment that would benefit from the improvements described in at least one embodiment;

[0005] Figure 2 is a system having a multi-format GPU docking board according to at least one embodiment;

[0006] Figure 3A is a top view of features of a multi-format GPU docking board according to at least one embodiment;

[0007] Figure 3B is a side view of features of a multi-format GPU docking board according to at least one embodiment;

[0008] Figure 3C is a block diagram illustrating a communication mode of a system having a multi-format GPU docking board according to at least one embodiment;

[0009] Figure 4A is a block diagram of coupling features between components of a multi-format GPU docking board according to at least one embodiment;

[0010] Figure 4B is a further block diagram detailing clock and power requirements for a multi-format GPU docking board according to at least one embodiment;

[0011] Figure 5 According to at least one embodiment, it is possible to use or make Figure 2-4B The process flow of the steps of the method for multi-format GPU docking board described;

[0012] Fig. 6A shows an example data center where Figure 2-5 at least one embodiment of;

[0013] Figure 6B , Figure 6C Inference and / or training logic for using, enabling, and / or supporting a multi-format GPU docking board, such as in FIG. Fig. 6A and reasoning and / or training logic used in at least one embodiment of the present disclosure;

[0014] Fig. 7A is a block diagram illustrating a computer system, which may be a system having interconnected devices and components, a system on a chip (SOC), or some combination thereof formed together with a processor, wherein the processor may include an execution unit for executing instructions to use, support and / or implement the multi-format GPU docking board described herein, according to at least one embodiment;

[0015] Figure 7B is a block diagram illustrating an electronic device for utilizing a processor to use, support and / or implement the multi-format GPU docking board described herein according to at least one embodiment;

[0016] Figure 7C is a block diagram illustrating an electronic device for utilizing a processor to use, support and / or implement the multi-format GPU docking board described herein according to at least one embodiment;

[0017] Figure 8Another example computer system for implementing various processes and methods for a multi-format GPU docking board described throughout this disclosure is shown in accordance with at least one embodiment;

[0018] Fig. 9A An architecture according to at least one embodiment of the present disclosure is shown, wherein a GPU is communicatively coupled to a multi-core processor via a high-speed link to implement and / or support a multi-format GPU docking board;

[0019] Fig. 9B shows additional details of the interconnection between a multi-core processor and a graphics acceleration module according to at least one embodiment;

[0020] Fig. 9C Another embodiment according to at least one embodiment disclosed herein is shown, wherein an accelerator integrated circuit is integrated within a processor for implementing and / or supporting a multi-format GPU docking board;

[0021] Fig.9D An accelerator integration slice 990 for using, implementing and / or supporting a multi-format GPU docking board according to at least one embodiment disclosed herein is shown;

[0022] Fig.9E shows additional details of one embodiment of a shared model for using, supporting and / or implementing a multi-format GPU docking board in accordance with at least one embodiment disclosed herein;

[0023] Fig.9F Additional details of at least one embodiment of a unified memory addressable via a common virtual memory address space for accessing physical processor memory and GPU memory to use, support and / or implement a multi-format GPU docking board are shown;

[0024] Fig. 10A An integrated circuit and associated graphics processor for a multi-format GPU docking board according to embodiments described herein are shown;

[0025] Figures 10B-10C An integrated circuit and associated graphics processor for using, supporting and / or implementing a multi-format GPU docking board according to at least one embodiment are shown;

[0026] Figures 10D-10E Additional graphics processor logic for using, supporting and / or implementing a multi-format GPU docking board is shown in accordance with at least one embodiment;

[0027] Fig.11A is a block diagram illustrating a computing system for using, supporting and / or implementing a multi-format GPU docking board according to at least one embodiment;

[0028] Fig. 11B A parallel processor for using, supporting and / or implementing a multi-format GPU docking board according to at least one embodiment is shown;

[0029] Fig. 11C is a block diagram of a partitioning unit according to at least one embodiment;

[0030] Fig.11D A graphics multiprocessor for a multi-format GPU docking board is shown according to at least one embodiment;

[0031] Fig.11E A graphics multiprocessor is shown in accordance with at least one embodiment;

[0032] Fig. 12A A multi-GPU computing system is shown in accordance with at least one embodiment;

[0033] Fig. 12B is a block diagram of a graphics processor according to at least one embodiment;

[0034] Fig.13 is a block diagram illustrating a microarchitecture for a processor, which may include logic circuitry for executing instructions, according to at least one embodiment;

[0035] Fig.14 A deep learning application processor according to at least one embodiment is shown;

[0036] Fig.15 is a block diagram of a neuromorphic processor according to at least one embodiment;

[0037] Fig.16A is a block diagram of a processing system according to at least one embodiment;

[0038] Fig. 16B is a block diagram of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor according to at least one embodiment;

[0039] Fig. 16C is a block diagram of the hardware logic of a graphics processor core according to at least one embodiment;

[0040] Figures 16D-16E Thread execution logic including an array of processing elements of a graphics processor core is shown according to at least one embodiment;

[0041] Fig.17A illustrates a parallel processing unit according to at least one embodiment;

[0042] Fig. 17B illustrates a general processing cluster in accordance with at least one embodiment;

[0043] Fig. 17C A memory partitioning unit of a parallel processing unit according to at least one embodiment is shown; and

[0044] Fig.17D A streaming multiprocessor in accordance with at least one embodiment is shown. DETAILED DESCRIPTION

[0045] The increase in artificial intelligence (AI) demand and processor-intensive operations or calculations has led to a demand for GPU-based products, particularly for multi-GPU-based products. One advantage provided by GPU-based products is the benefit of graphics engines that are designed from the ground up to address such processor-intensive operations or calculations relative to CPUs. For example, a GPU may include more smaller-sized logic cores, such as arithmetic logic units or ALUs, control units, and associated memory caches. At least these features advance parallel processing or calculations while simplifying the input requirements for performing processor-intensive operations or calculations.

[0046] At least for similar reasons, servers or high-performance and high-performance computing rely more on GPU-based products, such as servers with many processing cores. These high-performance and high-performance computing resources are well suited to the needs of AI and machine learning (ML). However, current GPU docking boards may be specific to the format or form factor requirements of the GPU board. For example, a GPU docking board may adopt a universal slot, such as a peripheral component interconnect express (PCI) Inlay connector or with corresponding The GPU's functionality is then limited only by the Inlay Connector and GPU Board Connector bottom The functional limitations of the standard. It can be a standard for general computing (which can be computing that is not closely related to AI or ML needs). In addition, this general computing relies on One or two GPUs in a slot interface.

[0047] This architecture may be of limited benefit for AI and ML needs, and Standards can limit productivity. An inlay connector refers to one or both of the male or female connectors that couple together to enable coupling between a component and a GPU docking board or motherboard. For a component such as a CPU or GPU, it can be a socket interface, which is a male-side collection of pins, solder balls, or pads in a 1xN configuration (with a row of pins) or NxN dimensions (with a ball grid array) that couples to the female-side holes of a socket on a GPU docking board or motherboard. The male or female side of a socket is called an inlay connector, a slot interface, or a socket.

[0048] Such problems are further evident in the complex communication schemes that enable communication between the CPU and GPU to execute applications on one or more processors. In at least one aspect, applications can also limit the distribution of workloads to the available architectures. When the limits are reached within the available components, the performance of the application is lower than expected. In most applications, the performance threshold of the GPU defines the performance threshold of the system. Further, Nvidia's NVLINK TM Standards rely on SXM TM or SXM2 TM (interchangeably referred to as SXM TM , including backward and forward compatibility with the SXM3 and SXM4 standards) connector, or slot interface operates largely and is physically identical to The inlay connector or socket interface is different. GPU boards and SXM-based TM GPU boards may have different GPU docking boards available. Therefore, GPU boards may be selected based in part on the availability of GPU docking boards, or in part on the availability of GPU docking boards. Or choose GPU docking board based on the availability of SXM GPU board. Then the restriction is for the same GPU board or the same interface (need or SXM TM ) can be used in one chassis at the same time, but both cannot be enabled at the same time. The present disclosure includes at least one embodiment to solve at least one of the above disadvantages, but can also solve all the disadvantages.

[0049] In at least one embodiment, a multi-format docking board is disclosed. The multi-format docking board includes a first switch for enabling communication between a first-format graphics processing unit (GPU) board and a second-format GPU board. The multi-format docking board also includes a second switch for enabling communication between a central processing unit (CPU) and the first-format GPU board and between the CPU and the second-format GPU board. In at least one embodiment, a first-format GPU board for operating with a second-format graphics processing unit (GPU) board on the multi-format docking board is also disclosed. The first-format GPU board includes a bridge connector connected to the first switch for enabling communication between the first-format GPU board and the second-format GPU board. In addition, the first-format GPU board also includes an inlay connector connected to the second switch for enabling communication between the central processing unit (CPU) and the first-format GPU board and between the CPU and the second-format GPU board. Further, in at least one embodiment, a switch for enabling the multi-format docking board to operate with the second switch is disclosed. The switch includes a bridge connector for enabling communication with the first-format graphics processing unit (GPU) board, and includes one or more inlay connectors for enabling communication with the second-format GPU board and the central processing unit (CPU).

[0050] Figure 1 is a system diagram of an example GPU docking board and environment 100 that benefits from the improvements described in at least one embodiment. A server tray or server box 102 is provided with one or more boards 104; 106 on which components are connected to perform the computing needs of the data center hosting the server tray or box 102. One or more boards 104, 106 may be a motherboard or a GPU docking board. When a board 104; 106 is not a motherboard, it may include an inlay connector or pin interface for fitting within a companion inlay connector, socket or slot interface of a motherboard. Inlay connectors, such as pins, pads, and balls, refer to board-level physical connections provided between components and boards 104; 106. The motherboard or GPU docking board carries one or more GPU boards 120; 122 of some form factor or format, as well as other components such as a CPU 118, memory devices based on persistent memory or random access memory (RAM), and input / output ports 124.

[0051] One or more boards 104, 106 include a slot interface, socket or inlay connector 108-116, which are used interchangeably herein, for receiving the above-mentioned components, including one or more GPU boards 120; 122 of a certain form factor or format, CPU 118, and persistent or RAM-based storage devices. On the component side, one or more GPU boards 120; 122 of a certain form factor or format, CPU 118, and persistent or RAM-based storage devices also have inlay connectors 126, which mate with appropriate and respective inlay connectors 108-116. The inlay connectors 126 on the component side can be solder ball contacts, pads, bumps, pins or other physical input / output features for providing input-output and receiving signals from reference sources, such as address signals, enable signals, and clock signals.

[0052] Figure 2 A system 200 having a multi-format GPU docking board 202 is provided according to at least one embodiment. Figure 1 Unlike the boards 104, 106 of the multi-format GPU docking board 202, the multi-format GPU docking board 202 is capable of supporting GPU boards of multiple form factors. In at least one embodiment, the multi-format GPU docking board 202 includes slot interfaces, inlay connectors or sockets 204, 208, 212, 214, 218, 226. The slot interfaces, inlay connectors or sockets 204, 208, 212, 214, 218, 226 can be female connectors for accepting corresponding male slot interfaces, inlay connectors or socket pins or solder balls. These slot interfaces, inlay connectors or sockets 204, 208, 212, 214, 218, 226 are according to at least one embodiment, so that additional slot interfaces, inlay connectors or sockets can be provided. In at least one embodiment, there is a first format GPU board (such as GPU board 228) and for a second format GPU board (such as NVLINK TM (or SXM TM ) GPU board 224). The slots 222, 226 for GPU boards 224, 228 are shown as physically different to support multi-format or different form factor GPU boards 224, 228. In at least one embodiment, the physical layout of the slots 222, 226 may be different than shown. Some slot interfaces 204, 212 may be used for persistent or temporary memory, including random access memory (RAM) or non-volatile memory sticks. Socket 208 enables coupling of CPU 210, while socket 218 enables coupling of switches 216, 220 to the multi-format GPU docking board 202.

[0053] In at least one embodiment, the multi-format GPU docking board 202 includes a first switch 220 for enabling communication between a first format graphics processing unit (GPU) board 228 and a second format GPU board 224. In at least one embodiment, the first switch is an NVLINK TM An interconnect standard switch adapted to enable communication between a first format graphics processing unit (GPU) board 228 and a second format GPU board 224. Board input or output of an accessory may be provided via port 230.

[0054] In at least one embodiment, the multi-format GPU docking board 202 includes a second switch 216 for enabling communication between the central processing unit (CPU) 210 and the first format GPU board 228 and between the CPU 210 and the second format GPU board 224. In at least one embodiment, the second switch 216 is In at least one embodiment, the multi-format docking board 202 is a system that includes connectors associated with one or more components thereon. In at least one embodiment, the multi-format docking board 202 also includes a bridge connector 228A located on the first format GPU board 228. The bridge connector 228A enables coupling between the first format GPU board 228 and the first switch 220.

[0055] In at least one embodiment, the multi-format docking board 202 includes an inlay connector 228B that enables coupling between the first switch 220 and the second format GPU board 224. Coupling between the first switch 220 and the second format GPU board 224 is additionally enabled when the first format GPU board 228 is plugged into one of the slot interfaces 226 via the inlay connector 228B; when the first switch 220 is coupled to the board 220 at the receptacle 218; when the bridge connector 228A is coupled to the first switch 220 via, for example, a flexible ribbon cable; and when the second format GPU board 224 is coupled to one of the receptacles 222 available for coupling. Unless otherwise specified, coupling refers to enabling communication, power, and operation of the respective components discussed.

[0056] Figure 3A 3 is a top view of features 300 of a multi-format GPU docking board 302 according to at least one embodiment. In at least one embodiment, ribbon cable 306 implements bridge connector 316 (on the top portion of a first format GPU board, such as Figure 3B The bottom or other portion of the first format GPU board (such as Figure 3B338 in the reference section) has an inlay connector for coupling with the multi-format GPU docking board 302 at the inlay connector 306. In at least one embodiment, the board line 314 implements the inlay connector of the first switch 302 (in at least one embodiment, this can be Figure 2 solder ball 216A in the second format GPU board (in Figure 3A 302). In addition, second, third, or additional board lines 314 couple the second switch 304 to the inlay connectors 306, 308 of the first and second format GPU boards. Even though shown as solid lines 314, in at least one embodiment, the board lines 312, 314 represent printed circuit board (PCB) lines. The board lines 312, 314 may be within one or more PCB layers of the PCB and may not be visible.

[0057] In at least one embodiment, the multi-format docking board 302 also includes an inlay connector 306 for coupling between the first format GPU board and the second switch 304 via the board line 314. In at least one embodiment, a second or additional inlay connector 308 is provided for coupling between the second format GPU board and the second switch 304 via the board line 314. Further, in at least one embodiment, the first switch 302 is Switch, the second switch 304 is NVLINK TM In at least one embodiment, both sides of the ribbon cable 306 have NVLINK TM Therefore, the bridge connector 316 and the inlay connector 310 are NVLINK TM Connector.

[0058] Figure 3B is a side view of a feature 320 of a multi-format GPU docking board 322 according to at least one embodiment. In at least one embodiment, a third inlay connector 324 is provided, such as a Inlay connector 338 or NVLINK TM The male SXM2 connectors of the connectors 310 and 316 are used for coupling between the second switch 328 and the CPU. TM The connector implementation is based in part on the NVLINK of the Pascal microarchitecture TM Communication protocol. For a bidirectional bandwidth of 40 GB / s and a total aggregate bandwidth of 160 GB / s, NVLINK TM The connector can support up to 20GB / s. For 50GB / s bidirectional bandwidth, NVLINK TM Another version of the connector is capable of supporting signaling at up to 25Gbps (25GT / s) per line.

[0059] In at least one embodiment, The switch 330 is adapted to manage or control the bridge chip. The switch 330 is connected to the base station via a flexible ribbon cable 332. The SXM-based GPU board 326 communicates with the SXM-based GPU board 334 via the board line in the board 322. The SXM-based GPU board 326 communicates with the SXM-based GPU board 334 via the board line in the board 322. TM (or SXM2 TM , both referred to as SXM) connector 334 is inserted into board 322. NVLINK TM Switch 328 implements SXM-based GPU board 326 and SXM-based GPU board 326 via board lines in board 322. Communication between GPU boards 334.

[0060] In at least one embodiment, the SXM-based GPU board 326 uses the SXM TM Backward and forward compatible protocols. Similarly, SXM-based GPU board 326 can use SXM2.0 TM / SXM3.0 TM / SXM4.0 communication protocol. Similarly, based on NVLINK TM Components, including NVLINK TM Switch 328, NVLINK TM And based on NVLINK TM The ribbon cable 332 is forward and backward compatible and can be used with NVLINK 2.0 TM / NVLINK3.0 TM / NVLINK4.0 TM In at least one embodiment, the SXM-based GPU board is connected to the SXM TM The connector is directly connected to board 322. In at least one embodiment, a SXM-based GPU board can be used with a SXM TM A compatible adapter is connected indirectly to board 322 .

[0061] In at least one embodiment, based on GPU board 334 via NVLINK TM A flexible ribbon cable 332 is connected (at top portion 336) to board 322 and is connected via The connector is connected to the bottom portion 338 of the board. In at least one embodiment, based on The GPU board 334 is connected via Switch 330 is connected to board 322, the The switch 330 receives the NVLINK on its side. TM A flexible ribbon cable 332 is provided and is coupled to the board 322 via a receptacle, such as the receptacle 218 that may receive the receptacle interface 216A in at least one embodiment.

[0062] In at least one embodiment, a rigid PCB-supported NVLINK may also be used. TM A connector, such as the commercially available P4932 / P4933 connector, is used to replace the ribbon cable 332. In at least one embodiment, the soft-flex NVLINK TM The cable may be based in part on NVLINK in accordance with the disclosure herein. TM protocol is designed to Switch 330 may replicate based on The coupling between GPU board 334 and board 322.

[0063] Figure 3C is a block diagram illustrating communication modes 356-362 of a system 350 including a multi-format GPU docking board 364, according to at least one embodiment. The multi-format GPU docking board 350 has connectors 352; 354 ​​of different formats or form factors for GPU boards of different formats or form factors. The connectors 352; 354 ​​of different formats or form factors are different physical interfaces. The connectors 352; 354 ​​of different formats or form factors are suitable for implementing the multi-format GPU docking board 364 in one chassis of a rack in a data center to couple or connect from different types of GPU boards.

[0064] In at least one embodiment, the present disclosure Features are multi-generational or support several generations Standards, including 2.0 / 3.0 / 4.0 interfaces. Therefore, in at least one embodiment, The connector 352 is connected to the multi-format GPU docking board 364 having a connector 352 (such as Connector) in the provided area based on GPU board to implement The GPU board is coupled to the multi-format GPU docking board 364. In addition, as shown in reference Figure 2 and Figure 3A , Figure 3B The discussion is based on A separate connector is available on the GPU board to support NVLINK TM NVLINK for switches TM In at least one embodiment, the Support generation in a similar way, in order to take advantage of SXM TMStandard multi-generation support feature, which can provide several SXMs on docking board 364 TM The interface or connector 354 can be used to insert different generations of SXM2.0 / 3.0 GPU boards on the multi-format GPU docking board 364, so as to couple the multi-format GPU docking board with the SXM-based GPU board.

[0065] In at least one embodiment, due to the The GPU board has two connectors that support and support NVLINK TM , so communication 360 is considered to represent both types of communications as if coming from a single interface or connector 352. So communication 360 can be or NVLINK TM In at least one embodiment, the formatted communication is a reference to the protocol followed to enable communication using the corresponding connector. Communication 360 can be to one or more components external to multi-format GPU docking board 364. In at least one embodiment, communication 360 can be in (a A GPU board Interface 352 and (the second based on Another NVLINK for the GPU board TM In at least one embodiment, communication 360 is with NVLINK TM NVLINK for switches TM Format of communication or Switch In at least one embodiment, communication 360 may be via Switch for CPU For these reasons, in at least one embodiment, communication 360 can occur between GPU boards of the same or different form factors, or based on between the GPU board and other components (such as the CPU).

[0066] In at least one embodiment, communication 358 is via SXM TM Connector NVLINK TM / In at least one embodiment, NVLINK is implemented TM Communication or interconnection SXM TM The connector can be two arrays with 400 pins on each array. In at least one embodiment, one of the two arrays is used for NVLINK TM Communication, the other of the two arrays is used for power supply, transmission of control signals and Communication. Thanks to SXM TM The connector is an inlay connector, so no bridge connector is required to achieve communication 358. As in the case of communication 360, NVLINK TM Communications can be directed to NVLINK via communications 356, 362 TM Switch 368, and from communications 358, 360 Communications can be directed to endpoints 364, 366 via In at least one embodiment, the endpoints may be Receive port on switch. Although communication 362 shows direct flow from connectors 352, 354, in at least one embodiment, such communication 362 occurs via communication 356 to switch 368. These features enable off-board communication, where GPU boards of different formats can be accessed by off-board devices such as CPU or memory devices via one or more communications 358, 360; but can also be accessed via NVLINK TM The switch enables on-board communication 362 between GPU boards of different formats.

[0067] In at least one embodiment, communication 358 may be to one or more components external to multi-format GPU docking board 364. In at least one embodiment, communication 358 may be on an SXM (of an SXM-based GPU board) TM Interface 354 and another SXM (of a second SXM-based GPU board) TM In at least one embodiment, communication 358 is to NVLINK TM NVLINK for switches TM In addition, in at least one embodiment, communication 360 may be via Switch for CPU For these reasons, in at least one embodiment, communication 360 can occur between different GPU boards or based on In at least one embodiment, the communication 372 is at least powered by the power supply 370. Therefore, in at least one embodiment, the docking board 364 has an external power supply 370.

[0068] Figure 4A is a block diagram of coupling features 400 between components 402-408 of a multi-format GPU docking board according to at least one embodiment. Figure 2-3C As discussed, the connector and protocol of the present disclosure are communicated via NVLINK TM Switch 406 is based on NVLINK is implemented between the GPU board 402 and the SXM-based GPU board 404 TM In at least one embodiment, communication from the different formats of the GPU boards 402, 404 to on-board or off-board non-GPU components (such as the CPU 410) is performed via Switch 408 occurs. In at least one embodiment, The switch includes at least chip, which allows Access based on the GPU board 402 to perform any communication, including failover and redundancy. This feature is based on at least The GPU board can be accessed or shared by multiple devices, including at least CPU 410.

[0069] In at least one embodiment, Switch 408 via The connector is coupled to the CPU 410 . Switch 408 Chip management based on In at least one embodiment, the NVLINK between GPU boards of different formats TM Communication can be via NVLINK TM Switch 406 occurs, and the different formats of GPU boards 402, 404 to CPU 410 are Communication can be done via The switch 408 occurs, but the two communications may occur in different routines or clock cycles. Although the different formats of the GPU boards 402, 404 may operate simultaneously to perform their respective calculations, interface with each other and interface with the CPU 410, the different formats of the GPU boards 402, 404 may operate at different core frequencies and may be in different bandwidth and can operate within the same NVLINK TM In at least one embodiment, The switch 408 includes a memory having instructions for the management software and a processor configured to execute the instructions to enable the The switch 408 is capable of executing various functions for the processor. In at least one embodiment, at least one function performed by the PCI switch 408 is processing resource allocation. In at least one embodiment, resource allocation enables GPU boards 402, 404 of different formats to perform different workloads for the CPU 410. In at least one embodiment, the NVLINK TMThe switch 406 handles the handshake requirements of the GPU boards 402, 404 of different formats because NVLINK TM Switch 406 uses NVLINK that can be used for different formats of GPU boards 402, 404 TM Public frequency for communication.

[0070] Figure 4B 4 is a further block diagram detailing the clock and power requirements 450 of the multi-format GPU docking board 462 according to at least one embodiment. In at least one embodiment, each GPU board 402, 404 of different formats has a different device identifier. In at least one embodiment, an operating system associated with a CPU 410 can identify all GPU boards 402, 404 of different formats coupled to the docking board 462. In at least one embodiment, the CPU 410 can use a GPU driver associated with the operating system to make such identification. In at least one embodiment, the operating system can allocate the necessary memory for one or more GPU tasks on the corresponding GPU, or allocate associated memory accessible to the corresponding GPU of the GPU board 402, 404 of different formats. In at least one embodiment, the operating system has full memory management capabilities for the corresponding GPU. In at least one embodiment, one or more CPUs can be associated with 8 to 16 individual GPUs on one or more different format boards of the system, and can be adapted to manage memory requirements as well as tasks of the GPUs.

[0071] In at least one embodiment, The switch 408 includes a processor and a memory having instructions to be executed on the processor. The instructions may cause the processor to perform a function of initializing one or more GPU boards 402, 404 of different formats. In at least one embodiment, the function in the processor is to initialize and The core associated with switch 408. Then, in at least one embodiment, The switch 408 can poll the connected GPU boards, such as the different formats of the GPU boards 402, 404, to provide at least a security identifier and any relevant information for each board, such as available memory, operating clock frequency, and operating mode. In at least one embodiment, the operating mode can include a code or can be a portion of the operating clock frequency. In at least one embodiment, the SXM-based GPU board can be connected to the NVLINK TM Frequency and In at least one embodiment, The frequency is 100MHz, NVLINK TM The frequency is 157 MHz. In at least one embodiment, NVLINK TMThe GPU board is tuned to use the clock from the docking board to reference the 1000+MHz internal clock.

[0072] In at least one embodiment, The core associated with the switch 408 is adapted to communicate the identifier of each of the differently formatted GPU boards 402, 404 and any associated information to the CPU. In at least one embodiment, the CPU polls The memory of the switch 408 is used to protect the identifiers of each of the different formats of the GPU boards 402, 404. In at least one embodiment, the GPU driver in the operating system associated with the CPU can be configured to communicate with the GPU board 402, 404 via The switch 408 facilitates further communication between the CPU and each GPU board 402, 404 of different formats, and is able to do so by addressing each GPU board 402, 404 of different formats. To the CPU, each GPU board 402, 404 of different formats is simply considered a different GPU board without specifically paying attention to their different formats. In at least one embodiment, one benefit recognized from the present architecture is to enable multiple GPU boards of different formats to be connected via a high-speed NVLINK. TM The architectures communicate with each other for specialized functions such as machine learning (ML) and artificial intelligence (AI). In at least one embodiment, each of the different formats of the GPU board is adapted to perform gradient calculations for a neural network layer in ML and AI applications. The calculations from the different formats of the GPU board can be averaged and then used to weight or cause changes to the neural network as part of the training applied to the neural network. Initializing the neural network can be done via instructions from the CPU via Communications (such as through In at least one embodiment, the Initialization of the switch 408 to register (or identify) different formats of GPU boards occurs during the power-on self-test (POST) of the operating system associated with the CPU.

[0073] In at least one embodiment, the SXM-based GPU board 404 includes a power regulator, a memory, one or more GPUs, a clock generator, and a power input for receiving power via an SXM connector from a docking board 462 associated with the SXM-based GPU board 404. In at least one embodiment, the clock generator of the SXM-based GPU board 404 is an embedded clock that enables NVLINK TM In at least one embodiment, the docking board 462 has a second clock generator, which can be different from the clock generator on the SXM-based GPU board 404, but it receives a reference from the clock generator. In addition, in at least one embodiment, the SXM-based GPU board 404 can be used to communicate with the SXM-based GPU board 404. The GPU board also receives a clock signal from the clock generator of the docking board via a flexible ribbon cable. This allows, for example, GPU boards of different formats to be used on the same NVLINK TM The docking board 462 also has power pins in the inlay connector 458 for providing power supply to the SXM-based GPU board 404.

[0074] Therefore, in at least one embodiment, the multi-format docking board has one or more clock circuits for providing a first clock frequency for the first format GPU board 402 and a second clock frequency for the second format GPU board 404. Alternatively, in at least one embodiment, the multi-format docking board has a clock generator that forms one or more clock circuits. The clock generator provides a reference for the second clock generator of the first format GPU board 402 and the third clock generator of the second format GPU board 404. The second and third clock generators provide the first clock frequency and the second clock frequency for the first format GPU board 402 and the second format GPU board 404, respectively. In at least one embodiment, the third clock generator uses the second clock generator as a reference so that the first format GPU board 402 and the second format GPU board 404 operate in a synchronous mode, for example, to implement the same NVLINK TM In at least one embodiment, the multi-format docking board includes a bridge connector to couple, for example, a flexible ribbon cable, between a first format GPU board and a first switch to provide at least a clock signal referenced from a second format GPU board.

[0075] In at least one embodiment, the multi-format docking board has one or more communication channels, such as for and for NVLINK TM In at least one embodiment, a first communication channel of the one or more communication channels operates at a first bandwidth for use with a first format GPU board and a first switch (such as Switch or NVLINK TM In at least one embodiment, a second communication channel of the one or more channels operates at a second bandwidth for a second communication between a second format GPU board and the first switch. In at least one embodiment, the first or second (or different) communication channel is configured to act as a common communication channel by operating at a common clock frequency for the first communication and the second communication. In at least one embodiment, NVLINK TM The switch operates at a single frequency, so the first and second communications are with NVLINK TM The switches are at the same clock frequency to achieve NVLINK TM communication.

[0076] In at least one embodiment, the multi-format docking board includes at least one processor associated with a first switch. The first switch also includes a memory having instructions for execution on the at least one processor to cause the at least one processor to perform functions. Figure 4A and Figure 4B As indicated in the associated description, when the first switch is When the switch The switch is adapted to perform functions such as allocating at least communication resources to the first format GPU board and the second format GPU board to enable communication with the CPU. In at least one embodiment, the SXM TM The switch includes at least one processor and a memory having instructions for execution by the at least one processor to perform functions such as allocating at least communication resources for a first format GPU board and a second format GPU board to communicate via NVLINK TM Interconnection standards to communicate with each other.

[0077] In at least one embodiment, the multi-format docking board includes a first format GPU board and a peripheral component interconnect fast link between the multi-format docking board. In at least one embodiment, the multi-format docking board also includes an NVLINK between the first format GPU board and the first switch. TM In at least one embodiment, the multi-format docking board includes an NVLINK for supporting a second format GPU board and a multi-format docking board. TM SXM interconnect standard TM or SXM2 TM Connector.

[0078] Figure 5 A method for using or making a multi-format GPU docking board (such as Figure 2-4B In at least one embodiment, step 502 is used to provide a multi-format GPU docking board, which has a first switch, a second switch, and a device (provision) for receiving a first format GPU board and a second format GPU board. In at least one embodiment, the device is at least one Inlay connector and at least one SXM inlay connector. In at least one embodiment, the first switch is an NVLINK TM switch, the second switch is switch.

[0079] In at least one embodiment, step 502 is performed by Figure 2-4B The disclosure of the preparation of the docking board is completed, including providing the necessary clock, power pins, management software (at least In at least one embodiment, step 504 implements the first switch to communicate between the first format GPU board and the second format GPU board. In at least one embodiment, step 504 implements the use of the first switch (acting as an NVLINK TM In at least one embodiment, step 506 determines whether a request is received for one or both of the first format GPU board and the second format GPU board to communicate with the CPU. Step 508 implements the second switch to communicate with the CPU and one or both of the first format GPU board and the second format GPU board that made the request.

[0080] In at least one embodiment, the method 500 includes coupling a first format GPU board to a first switch using at least a bridge connector, and includes coupling the first switch to a second format GPU board using at least an inlay connector. In at least one embodiment, the method 500 includes the step of implementing a ribbon cable between the bridge connector of the first format GPU board and the second bridge connector of the first switch. In at least one embodiment, by The GPU board provides at least NVLINK TM Bridge connector and NVLINK TM NVLINK is provided on the switch TM Bridge connector to implement ribbon cable.

[0081] In at least one embodiment, the method 500 includes the step of implementing a board line between an inlay connector of the first switch and a second inlay connector of the second format GPU board (at least as part of step 504). The board line is located in one or more layers of the docking board and enables the first format GPU board to communicate with the second format GPU board. In at least one embodiment, the method 500 includes providing an inlay connector for coupling between the first format GPU board and the second switch. In at least one embodiment, such a step is part of step 508, based on Communication between both the GPU board and the SXM-based GPU board to the CPU occurs via an inlay connector, such as the one on the docking board. The SXM connector on the connector and the mating board has a TM The communication setting format is part of the pin and the Therefore, in at least one embodiment, step 508 includes providing a second inlay connector for coupling between the second format GPU board and the second switch. In addition, in at least one embodiment, step 508 includes providing a third inlay connector for coupling between the second switch and the CPU.

[0082] In at least one embodiment, one or more of steps 504 and 508 are further implemented by providing a first clock frequency for the first format GPU board and a second clock frequency for the second format GPU board using one or more clock circuits on the docking board. In at least one embodiment, the one or more clock circuits are separate clock generators on each of the first format and the second format GPU board, and the separate clock generator can reference the clock generator of the docking board.

[0083] In at least one embodiment, one or more of steps 504 and 508 are implemented by enabling a first bandwidth for a first communication between a first format GPU board and a first switch within a first communication channel. Further, steps 504 and 508 may also be supported by enabling a second bandwidth for a second communication between a second format GPU board and the first switch within a second communication channel. In at least one embodiment, method 500 includes, for at least step 504, using a protocol (such as a protocol for NVLINK) implemented for the first communication and the second communication. TM A common communication channel with a common clock frequency for communication.

[0084] In at least one embodiment, method 500 includes: allocating at least communication resources for the first format GPU board and the second format GPU board by at least a processor associated with the first switch. This implements at least steps 504 and 508. In addition, in at least one embodiment, method 500 includes: providing a clock signal referenced from at least the second format GPU board using a bridge connector between the first format GPU board and the first switch.

[0085] In at least one embodiment, the multi-format docking board and its multi-format GPU board of the present disclosure can be implemented to execute a deep learning application processor (such as Fig.14 The processor 1400 in FIG. 1 may be used, and the neuron 1502 and its components implemented using circuits or logic may be used, including Fig.15 In addition, the deep learning application processor can be executed on multiple GPUs of multiple formats of GPU boards so that on the one hand, it can be in a training state at any time, and on the other hand, it can be in an application state at any time. In at least one embodiment, aspects of the deep learning processing can use the collected information processed according to multiple features determined by the discriminant analysis of the application in question, as referenced Fig.14 , Fig.15Further discussed. In one example, the process uses multiple neuron levels of a machine learning model (which is loaded with one or more values) to perform training, which can represent the correlation of values ​​under certain weights and the extrapolation of values ​​for new inputs provided. In at least one embodiment, the neuron level can store values ​​associated with the training process and can represent the correlation or correlation between the values ​​at the output relative to the values ​​at the input. In at least one embodiment, the GPU can be a processor with multiple cores, such as Fig. 9A The cores described in the multi-core processors 905 and 906.

[0086] Data Center

[0087] Fig. 6A An example data center 600 is shown in which Figure 2-Figure 5 In at least one embodiment, the data center 600 includes a data center infrastructure layer 610, a framework layer 620, a software layer 630, and an application layer 640. In at least one embodiment, such as Figure 2 As described, features in components 204-214 may be performed within or in cooperation with the example data center 600. In at least one embodiment, the infrastructure layer 610, the framework layer 620, the software layer 630, and the application layer 640 may be provided in part or in whole via computing components on server trays located in racks 210 of the data center 200. In addition, various aspects of the data center, including the data center infrastructure layer 610, the framework layer 620, the software layer 630, and the application layer 640 may be provided as described with reference to at least the above Figure 2-Figure 5 The multi-format GPU docking board in question uses, implements and / or supports. Therefore, reference Figures 6A-17D The discussion may be understood to apply to the use, implementation and / or support of e.g. Figure 2-Figure 5 The hardware and software features required for the multi-format GPU docking board.

[0088] In at least one embodiment, Fig. 6AAs shown, the data center infrastructure layer 610 may include a resource coordinator 612, group computing resources 614, and node computing resources ("node CRs") 616 (1)-616 (N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 616 (1)-616 (N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node CRs 616 (1)-616 (N) may be a server having one or more of the above-mentioned computing resources.

[0089] In at least one embodiment, the grouped computing resources 614 may include a separate grouping (not shown) of node CRs housed in one or more racks, or many racks (also not shown) housed in data centers at various geographic locations. The separate grouping of node CRs within the grouped computing resources 614 may include computing, networks, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0090] In at least one embodiment, resource coordinator 612 may configure or otherwise control one or more nodes CR 616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 612 may include a software design infrastructure ("SDI") management entity for data center 600. In at least one embodiment, resource coordinator 612 may include hardware, software, or some combination thereof.

[0091] In at least one embodiment, Fig. 6AAs shown, the framework layer 620 includes a job scheduler 622, a configuration manager 624, a resource manager 626, and a distributed file system 628. In at least one embodiment, the framework layer 620 may include a framework that supports software 632 of the software layer 630 and / or one or more applications 642 of the application layer 640. In at least one embodiment, the software 632 or the application 642 may include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 620 may be, but is not limited to, a free and open source software network application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 628 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 622 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 600. In at least one embodiment, the configuration manager 624 may be able to configure different layers, such as the software layer 630 and the framework layer 620 including Spark and a distributed file system 628 for supporting large-scale data processing. In at least one embodiment, the resource manager 626 can manage clustered or grouped computing resources mapped to or allocated to support the distributed file system 628 and the job scheduler 622. In at least one embodiment, the clustered or grouped computing resources can include grouped computing resources 614 on the data center infrastructure layer 610. In at least one embodiment, the resource manager 626 can coordinate with the resource coordinator 612 to manage these mapped or allocated computing resources.

[0092] In at least one embodiment, software 632 included in software layer 630 may include software used by at least a portion of node CRs 616(1)-616(N), grouped computing resources 614, and / or distributed file system 628 of framework layer 620. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0093] In at least one embodiment, one or more applications 642 included in the application layer 640 may include one or more types of applications used by at least a portion of the node CRs 616(1)-616(N), the grouped computing resources 614, and / or the distributed file system 628 of the framework layer 620. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0094] In at least one embodiment, any of configuration manager 624, resource manager 626, and resource coordinator 612 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of data center 600 from making potentially bad configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.

[0095] In at least one embodiment, the data center 600 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments of the present invention. In at least one embodiment, the machine learning model can be trained by calculating weight parameters according to the neural network architecture using the software and computing resources described above for the data center 600. In at least one embodiment, by using the weight parameters calculated by one or more training techniques herein, the resources described above with respect to the data center 600 can be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks. Any appropriate learning network and the computing capabilities of the data center 600 can be used to advance deep learning. Therefore, in this way, the hardware in the data center can be used to simultaneously or concurrently support deep neural networks (DNNs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs). For example, once the network is trained and successfully evaluated to identify data in a subset or slice, the trained network can provide similar representative data for use with the collected data.

[0096] In at least one embodiment, the data center 600 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as pressure, flow rate, temperature, and location information or other artificial intelligence services.

[0097] Reasoning and training logic

[0098] Reasoning and / or training logic 615 may be used to perform reasoning and / or training operations associated with one or more embodiments. In at least one embodiment, reasoning and / or training logic 615 may be used in the system Fig. 6A Inference and / or training logic 615 may be used to reason or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. In at least one embodiment, reasoning and / or training logic 615 may include, but is not limited to, hardware logic in which computing resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, reasoning and / or training logic 615 may be used in conjunction with an application specific integrated circuit (ASIC), such as the one from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g. "LakeCrest") processor.

[0099] In at least one embodiment, the reasoning and / or training logic 615 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the reasoning and / or training logic 615 includes, but is not limited to, a code and / or data storage model that can be used to store code (e.g., graphics code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameters or hyperparameter information. In at least one embodiment, each code and / or data storage module is associated with a dedicated computing resource. In at least one embodiment, the dedicated computing resource includes computing hardware that also includes one or more ALUs that perform mathematical functions (e.g., linear algebra functions) only on information stored in the code and / or data storage module, and stores the results stored therefrom in an activated storage module of the reasoning and / or training logic 615.

[0100] Figure 6B , Figure 6CInference and / or training logic according to at least one embodiment is shown, such as in Fig. 6A The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with at least one embodiment of the present disclosure. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. Distinguished from the computational hardware 602, 606 by the use of an arithmetic logic unit (ALU) 610 Figure 6B and Figure 6C In at least one embodiment, each of computing hardware 602 and computing hardware 606 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on information stored in code and / or data memory 601 and information stored in code and / or data memory 605, respectively, with the results stored in activation memory 620. Thus, unless otherwise specified, Figure 6B and Figure 6C may be substituted and used interchangeably.

[0101] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 601 to store forward and / or output weights and / or input / output data and / or other parameters of neurons or layers of a neural network trained and / or used for inference in at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 601 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, the code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 601 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with at least one embodiment during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of at least one embodiment. In at least one embodiment, any portion of code and / or data storage 601 may be included in other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.

[0102] In at least one embodiment, any portion of code and / or data storage 601 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or code and / or data storage 601 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 601 is internal or external to a processor, for example, or includes DRAM, SRAM, flash memory, or some other storage type, may depend on the available storage space on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the data batch size used in inference and / or training of a neural network, or some combination of these factors.

[0103] In at least one embodiment, the inference and / or training logic 615 may include, but is not limited to, code and / or data storage 605 to store backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in at least one embodiment aspect of the neural network. In at least one embodiment, the code and / or data storage 605 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with at least one embodiment during backpropagation of input / output data and / or weight parameters during training and / or inference using at least one embodiment. In at least one embodiment, the training logic 615 may include or be coupled to code and / or data storage 605 for storing graph code or other software to control the timing and / or order in which weights and / or other parameter information is loaded to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).

[0104] In at least one embodiment, code (such as graph code) loads weights or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 605 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 605 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 605 is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other storage type, depends on whether the available storage is on-chip or off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in the inference and / or training of the neural network, or some combination of these factors.

[0105] In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be separate storage structures. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be the same storage structure. In at least one embodiment, code and / or data storage 601 and code and / or data storage 605 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 601 and code and / or data storage 605 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0106] In at least one embodiment, inference and / or training logic 615 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 610, ALUs 610 comprising integer and / or floating point units for performing logic and / or mathematical operations based at least in part on or as directed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from a layer or neuron within a neural network) stored in activation storage 620, which are functions of input / output and / or weight parameter data stored in code and / or data storage 601 and / or code and / or data storage 605. In at least one embodiment, activations are performed in response to executing instructions or other code, linear algebra and / or matrix-based mathematics performed by ALU 610 to generate activations stored in activation storage 620, where weight values ​​stored in code and / or data storage 605 and / or in code and / or data storage 601 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 605 and / or code and / or data storage 601 or other on-chip or off-chip storage.

[0107] In at least one embodiment, one or more ALUs 610 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 610 may be outside a processor or other hardware logic device or circuit that uses them (e.g., a coprocessor). In at least one embodiment, one or more ALUs 610 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by an execution unit of a processor, which may be within the same processor or distributed between different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, code and / or data storage 601, code and / or data storage 605, and activation storage 620 may be on the same processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 620 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to a processor or other hardware logic or circuitry and may be retrieved and / or processed using the processor's fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0108] In at least one embodiment, activation storage 620 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, activation storage 620 may be completely or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, activation storage 620 may be selected to be internal or external to a processor, for example, or to include DRAM, SRAM, flash memory, or other storage types, depending on the storage available on-chip or off-chip, the latency requirements for performing training and / or inference functions, the batch size of data used in inferencing and / or training neural networks, or some combination of these factors. In at least one embodiment, Figure 6B The inference and / or training logic 615 shown in FIG. 6 may be used in conjunction with an application specific integrated circuit (“ASIC”), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6B The illustrated inference and / or training logic 615 may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as a field programmable gate array (“FPGA”).

[0109] In at least one embodiment, Figure 6C Inference and / or training logic 615 is shown, which may include, but is not limited to, hardware logic, wherein computing resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network, in accordance with at least one various embodiment. In at least one embodiment, Figure 6C The inference and / or training logic 615 shown in FIG. 6 can be used in conjunction with an application specific integrated circuit (ASIC), such as the ASIC from Google. Processing unit from Graphcore TM Inference Processing Unit (IPU) from Intel Corp (e.g., "Lake Crest") processor. In at least one embodiment, Figure 6CThe inference and / or training logic 615 shown in can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 615 includes, but is not limited to, code and / or data storage 601 and code and / or data storage 605, which can be used to store code (e.g., chart code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 6C In at least one embodiment shown in , each of code and / or data store 601 and code and / or data store 605 are associated with dedicated computing resources (eg, computing hardware 602 and computing hardware 606 ), respectively.

[0110] In at least one embodiment, each of the code and / or data stores 601 and 605 and the corresponding computing hardware 602 and 606 corresponds to a different layer of the neural network, such that activations from one "storage / compute pair 601 / 602" of the code and / or data store 601 and computing hardware 602 are provided as inputs to the next "storage / compute pair 605 / 606" of the code and / or data store 605 and computing hardware 606, so as to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / compute pair 601 / 602 and 605 / 606 can correspond to more than one neural network layer. In at least one embodiment, additional storage / compute pairs (not shown) can be included in the inference and / or training logic 615 after or in parallel with the storage / compute pairs 601 / 602 and 605 / 606.

[0111] Computer Systems

[0112] Fig. 7A A block diagram of a computer system 700A is shown in accordance with at least one embodiment, the exemplary computer system may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof with a processor, the processor may include an execution unit for executing instructions to use, support, and / or implement a multi-format GPU docking board as described herein. In at least one embodiment, the computer system 700A may include, but is not limited to, components (such as processor 702) to execute algorithms for process data using execution units including logic in accordance with the present disclosure (such as the embodiments herein). In at least one embodiment, the computer system 700A may include a processor such as an Intel® processor available from Intel Corporation of Santa Clara, California. Processor family, XeonTM, XScaleTM and / or StrongARMTM, Core TM or Nervana TM microprocessor, but other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 700B may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0113] In at least one embodiment, computer system 700A may combine components 110-116 (from Figure 1 ) to use, support and / or implement a multi-format GPU docking board as described herein. At least for this reason, in one embodiment, Fig. 7A The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Fig. 7A A system on a chip ("SoC") may be shown. In at least one embodiment, Fig. 7A The devices shown in the figure can be connected to proprietary interconnects, standardized interconnects (e.g., ) or some combination thereof. In at least one embodiment, one or more components of computer system 700B are interconnected using a compute express link (CXL) interconnect. Inference and / or training logic 615 is used to perform inference and / or training operations associated with one or more embodiments, for example, as previously described with respect to Fig. 6A -C discussed below. Fig. 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Fig. 7A for inferring or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0114] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (Internet Protocol) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0115] In at least one embodiment, the computer system 700A may include, but is not limited to, a processor 702, which may include, but is not limited to, one or more execution units 708 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 700A is a single-processor desktop or server system, but in another embodiment, the computer system 700A may be a multi-processor system. In at least one embodiment, the processor 702 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 702 may be coupled to a processor bus 710, which may transmit data signals between the processor 702 and other components in the computer system 700A.

[0116] In at least one embodiment, processor 702 may include, but is not limited to, level 1 ("L1") internal cache memory ("cache") 704. In at least one embodiment, processor 702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may reside external to processor 702. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, register file 706 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.

[0117] In at least one embodiment, an execution unit 708, including but not limited to logic to perform integer and floating point operations, is also located in the processor 702. In at least one embodiment, the processor 702 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 708 may include logic for processing a packed instruction set 709. In at least one embodiment, by including the packed instruction set 709 in the instruction set of a general-purpose processor, and the associated circuitry to execute the instructions, packed data in the general-purpose processor 702 may be used to perform operations used by many multimedia applications. In one or more embodiments, many multimedia applications may be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on the packed data, which may not require the transfer of smaller units of data on the processor's data bus to perform one or more operations one data element at a time.

[0118] In at least one embodiment, execution unit 708 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 700A may include, but is not limited to, memory 720. In at least one embodiment, memory 720 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 720 may store instructions 719 and / or data 721 represented by data signals that may be executed by processor 702.

[0119] In at least one embodiment, the system logic chip may be coupled to the processor bus 710 and the memory 720. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub ("MCH") 716, and the processor 702 may communicate with the MCH 716 via the processor bus 710. In at least one embodiment, the MCH 716 may provide a high bandwidth memory path 718 to the memory 720 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 716 may initiate data signals between the processor 702, the memory 720, and other components in the computer system 700A, and bridge data signals between the processor bus 710, the memory 720, and the system I / O 722. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 716 may be coupled to the memory 720 via a high bandwidth memory path 718 , and the graphics / video card 712 may be coupled to the MCH 716 via an Accelerated Graphics Port (“AGP”) interconnect 714 .

[0120] In at least one embodiment, the computer system 700A can use the system I / O 722 as a proprietary hub interface bus to couple the MCH 716 to an I / O controller hub ("ICH") 730. In at least one embodiment, the ICH 730 can provide direct connections to certain I / O devices through a local I / O bus. In at least one embodiment, the local I / O bus can include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 720, the chipset, and the processor 702. Examples can include, but are not limited to, an audio controller 729, a firmware hub ("Flash BIOS") 728, a wireless transceiver 726, a data store 724, a traditional I / O controller 723 including a user input and keyboard interface, a serial expansion port 727 (e.g., a universal serial bus (USB) port), and a network controller 734. The data store 724 can include a hard drive, a floppy drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0121] Figure 7B 700B for utilizing a processor 710 to use, support, and / or implement a multi-format GPU docking board as described herein, according to at least one embodiment. In at least one embodiment, the electronic device 700B may be, for example, but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device. In at least one embodiment, the electronic device 700B may be combined with components (from Figure 2-4B ) to use, support and / or implement the multi-format GPU docking board described herein.

[0122] In at least one embodiment, system 700B may include, but is not limited to, a processor 710 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 710 is coupled using a bus or interface, such as an I2C bus, a system management bus ("SMBus"), a low pin count (LPC) bus, a serial peripheral interface ("SPI"), a high-definition audio ("HDA") bus, a serial advanced technology attachment ("SATA") bus, a universal serial bus ("USB") (versions 1, 2, 3), or a universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Figure 7B A system is shown that includes interconnected hardware devices or "chips", while in other embodiments, Figure 7B A system on a chip ("SoC") may be shown. In at least one embodiment, Figure 7B The devices shown in can be connected to proprietary interconnects, standardized interconnects (e.g., ) or some combination thereof. In at least one embodiment, Figure 7B One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.

[0123] In at least one embodiment, Figure 7B The display 724, the touch screen 725, the touch pad 730, the near field communication unit ("NFC") 745, the sensor hub 740, the thermal sensor 746, the fast chipset ("EC") 735, the trusted platform module ("TPM") 738, the BIOS / firmware / flash memory ("BIOS, FW Flash") 722, the DSP 760, the drive 720 (such as a solid state disk ("SSD") or a hard disk drive ("HDD")), the wireless local area network unit ("WLAN") 750, the Bluetooth unit 752, the wireless wide area network unit ("WWAN") 756, the global positioning system (GPS) unit 755, the camera ("USB 3.0 camera") 754 (such as a USB 3.0 camera) and / or the low power double data rate ("LPDDR") memory unit ("LPDDR3") 715 implemented in, for example, the LPDDR3 standard. Each of these components can be implemented in any suitable manner.

[0124] In at least one embodiment, other components may be communicatively coupled to the processor 710 via the following components. In at least one embodiment, the accelerometer 741, ambient light sensor (“ALS”) 742, compass 743, and gyroscope 744 may be communicatively coupled to the sensor hub 740. In at least one embodiment, the thermal sensor 739, fan 737, keyboard 746, and touchpad 730 may be communicatively coupled to the EC 735. In at least one embodiment, the speaker 763, earphone 764, and microphone (“mic”) 765 may be communicatively coupled to the audio unit (“audio codec and class D amplifier”) 762, which in turn may be communicatively coupled to the DSP 760. In at least one embodiment, the audio unit 764 may include, for example, but not limited to, an audio encoder / decoder (“codec”) and a class D amplifier. In at least one embodiment, the SIM card (“SIM”) 757 may be communicatively coupled to the WWAN unit 756. In at least one embodiment, components such as the WLAN unit 750 and the Bluetooth unit 752 and the WWAN unit 756 may be implemented as a next generation form factor (NGFF).

[0125] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CProvides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 7B for use in reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0126] Figure 7C A computer system 700C is shown according to at least one embodiment for using, supporting and / or implementing the multi-format GPU docking board described herein. In at least one embodiment, the computer system 700C includes, but is not limited to, a computer 771 and a USB disk 770. In at least one embodiment, the computer 771 may include, but is not limited to, any number and type of processors (not shown) and memories (not shown). In at least one embodiment, the computer 771 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0127] In at least one embodiment, the USB disk 770 includes, but is not limited to, a processing unit 772, a USB interface 774, and a USB interface logic 773. In at least one embodiment, the processing unit 772 can be any instruction execution system, device, or device capable of executing instructions. In at least one embodiment, the processing unit 772 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit or core 772 includes an application specific integrated circuit ("ASIC") that is optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 772 is a tensor processing unit ("TPC") that is optimized to perform machine learning reasoning operations. In at least one embodiment, the processing core 772 is a visual processing unit ("VPU") that is optimized to perform machine vision and machine learning reasoning operations.

[0128] In at least one embodiment, the USB interface 774 can be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 774 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 774 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 773 can include any number and type of logic that enables the processing unit 772 to connect to a device (e.g., computer 771) via the USB connector 774.

[0129] Reasoning and / or training logic 615 (e.g., regarding Figure 6B and Figure 6CThe invention is described in detail and is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 can be used to Figure 7C In a system of the present invention, operations can be inferred or predicted based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0130] Figure 8 A further exemplary computer system 800 is shown for using, supporting, and / or implementing the various processes and methods of the multi-format GPU docking board described throughout the present disclosure in accordance with at least one embodiment. In at least one embodiment, the computer system 800 includes, but is not limited to, at least one central processing unit ("CPU") 802 connected to a communication bus 810 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 800 includes, but is not limited to, a main memory 804 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data may be stored in the main memory 804 in the form of random access memory ("RAM"). In at least one embodiment, a network interface subsystem ("network interface") 822 provides an interface to other computing devices and networks for receiving data from the computer system 800 and transmitting data to other systems.

[0131] In at least one embodiment, computer system 800 includes, but is not limited to, input device 808, parallel processing system 812, and display device 806, which may be implemented using cathode ray tubes ("CRT"), liquid crystal displays ("LCD"), light emitting diodes ("LED"), plasma displays, or other suitable display technologies. In at least one embodiment, user input is received from input device 808 (such as a keyboard, mouse, touch pad, microphone, and more). In at least one embodiment, each of the above modules may be located on a single semiconductor platform to form a processing system.

[0132] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments, such as those previously described with respect to Fig. 6A -C discussed. The following combination Fig. 6A -C provides details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be implemented in the system Figure 8 In at least one embodiment, the inference and / or training logic 615 may be used in the system to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein. Figure 8 to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases herein.

[0133] Fig. 9A An architecture is shown according to at least one embodiment, wherein a plurality of GPUs 910-913 are communicatively coupled to a plurality of multi-core processors 940-943 via high-speed links 905-906 (e.g., buses / point-to-point interconnects, etc.). In one embodiment, the high-speed links 940-943 support 4 GB / s, 30 GB / s, 80 GB / s, or higher communication throughput. Various interconnect protocols may be used, including but not limited to 4.0 or 5.0 and NVLINK TM 2.0.

[0134] Furthermore, in one embodiment, two or more GPUs 910-913 are interconnected via high-speed links 929-930, which may be implemented using the same or different protocols / links as used for high-speed links 940-943. Similarly, two or more multi-core processors 905-906 may be connected via high-speed link 928, which may be a symmetric multiprocessor (SMP) bus running at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, the same protocol / links may be used (e.g., via a common interconnect fabric) to accomplish the above. Fig. 9A All communications between the various system components shown in .

[0135] In one embodiment, each multi-core processor 905-906 is communicatively coupled to processor memory 901-902 via memory interconnects 926-927, respectively, and each GPU 910-913 is communicatively coupled to GPU memory 920-923 via GPU memory interconnects 950-953, respectively. The memory interconnects 926-927 and 950-953 may utilize the same or different memory access technologies. By way of example and not limitation, the processor memory 901-902 and the GPU memory 920-923 may be volatile memory, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memory, such as 3D XPoint or Nano-Ram. In one embodiment, some portion of the processor memory 901-902 may be volatile memory, while another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0136] As follows, although the various processors 905-906 and GPUs 910-913 may be physically coupled to specific memories 901-902, 920-923, respectively, a unified memory architecture may be implemented in which a virtual system address space (also referred to as an "effective address" space) is distributed among the various physical memories. In at least one embodiment, the processor memories 901-902 may each include 64GB of system memory address space, and the GPU memories 920-923 may each include 32GB of system memory address space (resulting in a total addressable memory size of 256GB in this example).

[0137] As discussed elsewhere in this disclosure, flow rates and associated temperatures can be established for at least the first level of an intelligent learning system (such as a neural network system). Since the first level represents previous data, it also represents a smaller subset of data that can be used to improve the system by retraining the system. Testing and training can be performed in parallel using multiple processor units so that the intelligent learning system is stable. Fig. 9A When the intelligent learning system achieves convergence, the number of data points used to lead to convergence and the data in the data points are recorded. The data and data points can be used as described in the reference Figure 2-Figure 5 Multi-format GPU docking board generation.

[0138] Fig. 9BAdditional details for interconnection between multi-core processor 907 and graphics acceleration module 946 are shown according to at least one embodiment. Graphics acceleration module 946 may include one or more GPU chips integrated on a line card that is coupled to processor 907 via high-speed link 940. Optionally, graphics acceleration module 946 may be integrated with processor 907 on the same package or chip.

[0139] In at least one embodiment, the processor 907 shown includes multiple cores 960A-960D, each core having a translation lookaside buffer 961A-961D and one or more caches 962A-962D. In at least one embodiment, the cores 960A-960D may include various other components not shown for executing instructions and processing data. The caches 962A-962D may include level 1 (L1) and level 2 (L2) caches. In addition, one or more shared caches 956 may be included in the caches 962A-962D and shared by each group of cores 960A-960D. In at least one embodiment, one embodiment of the processor 907 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. The processor 907 and the graphics acceleration module 946 are connected to the system memory 914, which may include Fig. 9A Processor memory 901-902 in.

[0140] Coherence is maintained for data and instructions stored in the various caches 962A-962D, 956 and system memory 914 via inter-core communications over the coherence bus 964. In at least one embodiment, each cache may have cache coherence logic / circuitry associated therewith to communicate in response to detecting a read or write to a particular cache line over the coherence bus 964. In one implementation, a cache snooping protocol is implemented over the coherence bus 964 to snoop cache accesses.

[0141] In at least one embodiment, the proxy circuit 925 communicatively couples the graphics acceleration module 946 to the coherence bus 964, thereby allowing the graphics acceleration module 946 to participate in the cache coherence protocol as a peer of the cores 960A-960D. In particular, in at least one embodiment, the interface 935 communicates with the graphics acceleration module 946 via the high-speed link 940 (e.g., bus, NVLINK, etc.) provides a connection to the proxy circuit 925, and an interface 937 connects the graphics acceleration module 946 to the link 940.

[0142] In one implementation, the accelerator integrated circuit 936 provides cache management, memory access, context management, and interrupt management services on behalf of multiple graphics processing engines 931, 932, N of the graphics acceleration module. The graphics processing engines 931, 932, N may each include a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 931, 932, N may selectively include different types of graphics processing engines within a GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 946 may be a GPU having multiple graphics processing engines 931-932, N, or the graphics processing engines 931-932, N may be individual GPUs integrated on a common package, line card, or chip. As appropriate, the graphics processing engines 931-932, N may be integrated into a common package, line card, or chip. Fig. 9B The above determination of the reconstruction parameters and the reconstruction algorithm is performed in the GPU 931-N.

[0143] In one embodiment, the accelerator integrated circuit 936 includes a memory management unit (MMU) 939 for performing various memory management functions, such as virtual to physical memory translation (also known as effective to real memory translation), and also includes a memory access protocol for accessing the system memory 914. The MMU 939 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 938 may store commands and data for efficient access by the graphics processing engine 931-932, N. In at least one embodiment, the data stored in the cache 938 and the graphics memory 933-934, M may be kept consistent with the kernel cache 962A-962D, 956 and the system memory 914. As before, this task may be accomplished via proxy circuitry 925 acting on behalf of cache 938 and graphics memory 933-934, M (e.g., sending updates related to modifications / accesses of cache lines on processor caches 962A-962D, 956 to cache 938 and receiving updates from cache 938).

[0144] A set of registers 945 stores context data for threads executed by the graphics processing engines 931-932, N, and context management circuitry 948 manages thread contexts. In at least one embodiment, context management circuitry 948 may perform save and restore operations to save and restore contexts for various threads during context switches (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). In at least one embodiment, context management circuitry 948 may store current register values ​​to a specified area in memory (e.g., identified by a context pointer) upon context switching. The register values ​​may then be restored when returning to context. In one embodiment, interrupt management circuitry 947 receives and processes interrupts received from system devices.

[0145] In one implementation, the MMU 939 converts virtual / effective addresses from the graphics processing engine 931 into real / physical addresses in the system memory 914. One embodiment of the accelerator integrated circuit 936 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 946 and / or other accelerator devices. The graphics accelerator module 946 can be dedicated to a single application executed on the processor 907, or can be shared between multiple applications. In one embodiment, a virtualized graphics execution environment is presented in which the resources of the graphics processing engines 931-932, N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.

[0146] In at least one embodiment, the accelerator integrated circuit 936 performs as a bridge to the system of graphics acceleration modules 946 and provides address translation and system memory cache services. In addition, the accelerator integrated circuit 936 can provide virtualization facilities for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 931-932, N.

[0147] Since the hardware resources of the graphics processing engines 931-932, N are explicitly mapped to the real address space seen by the host processor 907, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of the accelerator integrated circuit 936 is to physically separate the graphics processing engines 931-932, N so that they appear to the system as independent units.

[0148] In at least one embodiment, one or more graphics memories 933-934, M are respectively coupled to each graphics processing engine 931-932, N. The graphics memories 933-934, M store instructions and data, which are processed by each graphics processing engine 931-932, N. The graphics memories 933-934, M may be volatile memories, such as DRAM (including stacked DRAM), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memories, such as 3DXPoint or Nano-Ram.

[0149] In one embodiment, to reduce data traffic on link 940, biasing techniques are used to ensure that data stored in graphics memory 933-934, M is data that is most frequently used by graphics processing engines 931-932, N, and that may not be used (at least not frequently) by cores 960A-960D. Similarly, the biasing mechanism attempts to keep data that is needed by a core (and may not be graphics processing engine 931-932, N) in cache 962A-962D, core 956, and system memory 914.

[0150] Fig. 9C At least one embodiment is shown in which an accelerator integrated circuit 936 is integrated within the processor 907 for use, support and / or implementation of the multi-format GPU docking board described herein in accordance with at least one embodiment disclosed herein. In at least this embodiment, the graphics processing engines 931-932, N communicate directly with the accelerator integrated circuit 936 via the interface 937 and the interface 935 (again, any form of bus or interface protocol can be used) through the high-speed link 940. The accelerator integrated circuit 936 can perform operations related to Fig. 9B The operations described above are similar to the operations described above. However, due to its close proximity to the coherence bus 964 and caches 962A-962D, 956, it may have a higher throughput. At least one embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), and the programming models may include a programming model controlled by the accelerator integrated circuit 936 and a programming model controlled by the graphics acceleration module 946.

[0151] In at least one embodiment, the graphics processing engines 931-932, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to the graphics processing engines 931-932, N, thereby providing virtualization within a VM / partition.

[0152] In at least one embodiment, the graphics processing engines 931-932, N can be shared by multiple VM / application partitions. In at least one embodiment, the sharing model can use a system hypervisor to virtualize the graphics processing engines 931-932, N to allow each operating system to access. For a single partition system without a hypervisor, the operating system owns the graphics processing engines 931-932, N. In at least one embodiment, the operating system can virtualize the graphics processing engines 931-932, N to provide access to each process or application.

[0153] In at least one embodiment, the graphics acceleration module 946 or individual graphics processing engines 931-932, N use a process handle to select a process element. In at least one embodiment, the process element is stored in the system memory 914 and can be addressed using the effective address to real address conversion technology of this article. In at least one embodiment, the process handle can be an implementation-specific value that is provided to the host process when registering its context with the graphics processing engine 931-932, N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle can be the offset of the process element in the process element linked list.

[0154] Fig.9D At least one embodiment of an accelerator integrated slice 990 for using, supporting and / or implementing a multi-format GPU docking board described herein according to at least one embodiment disclosed herein is shown. As used herein, a "slice" includes a specified portion of the processing resources of an accelerator integrated circuit 936. An application is an effective address space 982 in the system memory 914, which stores a process element 983. In at least one embodiment, the process element 983 is stored in response to a GPU call 981 from an application 980 executed on a processor 907. The process element 983 contains the process state of the corresponding application 980. The work descriptor (WD) 984 contained in the process element 983 can be a single job requested by the application, or can contain a pointer to a job queue. In at least one embodiment, the WD 984 is a pointer to a job request queue in the address space 982 of the application.

[0155] Graphics acceleration module 946 and / or individual graphics processing engines 931-932, N can be shared by all processes or a subset of processes in the system. In at least one embodiment, infrastructure for setting process state and sending WD 984 to graphics acceleration module 946 to start a job in a virtualized environment can be included.

[0156] In at least one embodiment, the dedicated process programming model is implementation specific. In this model, a single process owns a graphics acceleration module 946 or an individual graphics processing engine 931. When the graphics acceleration module 946 is owned by a single process, the hypervisor initializes the accelerator integrated circuit for the owned partition, and when the graphics acceleration module 946 is assigned, the operating system initializes the accelerator integrated circuit 936 for the owned process.

[0157] In operation, the WD acquisition unit 991 in the accelerator integrated slice 990 acquires the next WD 984, which includes an indication of the work to be completed by one or more graphics processing engines of the graphics acceleration module 946. Data from the WD 984 can be stored in registers 945 and used by the MMU 939, interrupt management circuits 947, and / or context management circuits 948, as shown. In at least one embodiment, an embodiment of the MMU 939 includes a segment / page roaming circuit for accessing a segment / page table 986 within the OS virtual address space 985. The interrupt management circuit 947 can handle interrupt events 992 received from the graphics acceleration module 946. In at least one embodiment, when performing graphics operations, the effective address 993 generated by the graphics processing engine 931-932, N is converted to a real address by the MMU 939.

[0158] In one embodiment, the same register set 945 is replicated for each graphics processing engine 931-932, N and / or graphics acceleration module 946, and the same register set 945 can be initialized by a hypervisor or operating system. Each of these replicated registers can be included in an accelerator integrated slice 990. Registers, such as registers used in at least one embodiment and registers that can be initialized by a hypervisor are shown in Table 1.

[0159] Table 1 – Hypervisor Initialization Registers

[0160]

[0161]

[0162] Table 2 shows registers of at least one embodiment that may be initialized by an operating system.

[0163] Table 2 – Operating System Initialization Registers

[0164] 1 Process and thread identification 2 Effective Address (EA) context save / restore pointer 3 Virtual Address (VA) Accelerator Utilizes Record Pointers 4 Virtual Address (VA) Segment Table Pointer 5 Permission shielding 6 Job Descriptor

[0165] In at least one embodiment, each WD 984 is specific to a particular graphics acceleration module 946 and / or graphics processing engine 931-932, N. It contains all the information needed by the graphics processing engine 931-932, N to complete the work, or it can be a pointer to a memory location where the application has set up a command queue for the work to be done.

[0166] Fig.9E Additional details of at least one embodiment of the sharing model are shown. This embodiment includes a hypervisor real address space 998 in which a process element list 999 is stored. The hypervisor real address space 998 can be accessed via a hypervisor 996, which virtualizes a graphics acceleration module engine for an operating system 995.

[0167] In at least one embodiment, the shared programming model allows all processes or subsets of processes from all partitions or subsets of partitions in the system to use the graphics acceleration module 946. There are two programming models where the graphics acceleration module 946 is shared by multiple processes and partitions, namely, time-sliced ​​sharing and graphics-directed sharing.

[0168] In this model, the hypervisor 996 owns the graphics acceleration module 946 and makes its functionality available to all operating systems 995. For the graphics acceleration module 946 to support virtualization through the hypervisor 996, the graphics acceleration module 946 may comply with the following requirements: (1) the application's job request must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 946 must provide a context save and restore mechanism, (2) the graphics acceleration module 946 guarantees that the application's job request is completed within a specified amount of time, including any transition errors, or the graphics acceleration module 946 provides the ability to preempt job processing, and (3) fairness between graphics acceleration module 946 processes must be ensured when operating in a directed shared programming model.

[0169] In one embodiment, the application 980 is required to use the graphics acceleration module type, work descriptor (WD), permission mask register (AMR) value and context save / restore region pointer (CSRP) to make an operating system 995 system call. The graphics acceleration module type describes the target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system-specific value. In at least one embodiment, the WD is formatted specifically for the graphics acceleration module 946 and can take the form of a graphics acceleration module 946 command, an effective address pointer to a user-defined structure, an effective address pointer to a command queue, or any other data structure describing the work to be completed by the graphics acceleration module 946. In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to the application that sets the AMR. If the implementation of the accelerator integrated circuit 936 and the graphics acceleration module 946 does not support the user permission mask override register (UAMOR), the operating system can apply the current UAMOR value to the AMR value before passing the AMR in the hypervisor call. In at least one embodiment, the hypervisor 996 may apply the current privilege mask overwrite register (AMOR) value before placing the AMR into the process element 983. In at least one embodiment, the CSRP is one of the registers 945 that contains the effective address of an area in the application's effective address space 982 for the graphics acceleration module 946 to save and restore context state. This pointer is used in at least one embodiment and is optional if state does not need to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be fixed system memory.

[0170] Upon receiving the system call, the operating system 995 may verify that the application 980 has been registered and granted permission to use the graphics acceleration module 946. The operating system 995 then calls the hypervisor 996 using the information shown in Table 3.

[0171] Table 3 – OS to hypervisor call parameters

[0172]

[0173]

[0174] Upon receiving the hypervisor call, the hypervisor 996 verifies that the operating system 995 has registered and is granted permission to use the graphics acceleration module 946. The hypervisor 996 then places the process element 983 into a linked list of process elements of the corresponding graphics acceleration module 946 type. The process element may include the information shown in Table 4.

[0175] Table 4 – Process element information

[0176] 1 Work Descriptor (WD) 2 The authority mask register (AMR) value (potentially masked). 3 Effective Address (EA) Context Save / Restore Region Pointer (CSRP) 4 Process ID (PID) and optional thread ID (TID) 5 Virtual Address (VA) Accelerator Usage Record Pointer (AURP) 6 Virtual address of the storage segment table pointer (SSTP) 7 Logical Interrupt Service Number (LISN) 8 Interrupt vector table, derived from the hypervisor call parameters 9 Status Register (SR) Value 10 Logical Partition ID (LPID) 11 Real Address (RA) Hypervisor Accelerator Using Record Pointers 12 Storage Descriptor Register (SDR)

[0177] In at least one embodiment, the hypervisor initializes the plurality of accelerator integrated slice 990 registers 945 .

[0178] like Fig.9F As shown, in at least one embodiment, a unified memory is used, and the unified memory can be addressed via a common virtual memory address space for accessing physical processor memory 901-902 and GPU memory 920-923. In this implementation, operations executed on GPUs 910-913 use the same virtual / effective memory address space to access processor memory 901-902, and vice versa, thereby simplifying programmability. In one embodiment, the first part of the virtual / effective address space is allocated to processor memory 901, the second part is allocated to the second processor memory 902, the third part is allocated to GPU memory 920, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thus distributed in each of processor memory 901-902 and GPU memory 920-923, thereby allowing any processor or GPU to access the memory using a virtual address mapped to any physical memory.

[0179] In one embodiment, bias / coherency management circuits 994A-994E within one or more MMUs 939A-939E ensure cache coherency between caches of one or more host processors (e.g., 905) and GPUs 910-913 and implement biasing techniques that indicate physical memory where certain types of data should be stored. Fig.9F Multiple instances of bias / consistency management circuits 994A- 994E are shown in , but bias / consistency circuits may be implemented within an MMU of one or more host processors 905 and / or within an accelerator integrated circuit 936 .

[0180] One embodiment allows the GPU additional memory 920-923 to be mapped as part of the system memory and accessed using shared virtual memory (SVM) technology, but without suffering from the performance defects associated with full system cache coherence. In at least one embodiment, the ability to access the GPU additional memory 920-923 as system memory without heavy cache coherence overhead provides a favorable operating environment for GPU offloading. This arrangement allows the host processor 905 software to set operands and access calculation results without the overhead of traditional I / O DMA data copying. Such traditional copies include driver calls, interrupts, and memory mapped I / O (MMIO) accesses, which are all less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPU additional memory 920-923 without cache coherence overhead may be critical to the execution time of the offloaded calculation. For example, in the case of a large amount of streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPU 910-913. In at least one embodiment, the efficiency of operand setting, the efficiency of result access, and the efficiency of GPU calculation may play a role in determining the effectiveness of GPU offloading.

[0181] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. For example, a bias table can be used, which can be a page granular structure (e.g., controlled at the granularity of a memory page) that includes 1 or 2 bits per GPU attached memory page. In at least one embodiment, the bias table can be implemented in the stolen memory range of one or more GPU attached memories 920-923 with or without a bias cache in the GPU 910-913 (e.g., to cache frequently / recently used entries of the bias table). Alternatively, the entire bias table can be maintained within the GPU.

[0182] In at least one embodiment, before actually accessing the GPU memory, the bias table entry associated with each access to the GPU attached memory 920-923 is accessed, resulting in the following operations. Local requests from GPUs 910-913 that find their pages in the GPU bias are forwarded directly to the corresponding GPU memory 920-923. Local requests from GPUs that find their pages in the host bias are forwarded to processors 905 (e.g., via a high-speed link as above). In one embodiment, the request from processor 905 to find the requested page in the host processor bias completes a request similar to a normal memory read. Alternatively, a request pointing to a GPU bias page can be forwarded to GPUs 910-913. In at least one embodiment, if the GPU is not currently using the page, the GPU can then migrate the page to the host processor bias. In at least one embodiment, the bias state of the page can be changed by a software-based mechanism, a hardware-assisted software-based mechanism, or in limited cases by a purely hardware-based mechanism.

[0183] One mechanism for changing the bias state employs an API call (e.g., OpenCL) that in turn calls the GPU's device driver, which in turn sends a message (or queues a command descriptor) to the GPU, directing the GPU to change the bias state and, in some migrations, performs a cache flush operation in the host. In at least one embodiment, the cache flush operation is used for migrations from host processor 905 bias to GPU bias, but not for the reverse migration.

[0184] In one embodiment, cache coherency is maintained by temporarily rendering GPU biased pages that cannot be cached by the host processor 905. To access these pages, the processor 905 may request access from the GPU 910, which may or may not immediately grant access. Therefore, in order to reduce communication between the processor 905 and the GPU 910, it is beneficial to ensure that the GPU biased pages are the pages required by the GPU and not the pages required by the host processor 905, and vice versa.

[0185] Reasoning and / or training logic 615 is used to perform one or more embodiments. Figure 6B and / or Figure 6C Details regarding the inference and / or training logic 615 are provided.

[0186] Fig. 10AAn integrated circuit and associated graphics processor according to at least one embodiment of the various embodiments of this document are shown, which can be manufactured using one or more IP cores to use, support and / or implement the multi-format GPU docking board described herein. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general processor cores.

[0187] Fig. 10A 1 is a block diagram illustrating a system on a chip integrated circuit 1000A that may be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1000A includes one or more application processors 1005 (e.g., CPU), at least one graphics processor 1010, and may additionally include an image processor 1015 and / or a video processor 1020, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 1000A includes peripheral or bus logic, which includes a USB controller 1025, a UART controller 1030, an SPI / SDIO controller 1035, and an I 2 S / I 2 C controller 1040. In at least one embodiment, the integrated circuit 1000A may include a display device 1045 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1050 and a mobile industry processor interface (MIPI) display interface 1055. In at least one embodiment, storage may be provided by a flash memory subsystem 1060, including flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1065 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1070.

[0188] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be used in the integrated circuit 1000A to infer or predict operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0189] Figures 10B-10CAn integrated circuit and associated graphics processor according to at least one embodiment of the various embodiments of this document are shown, which can be manufactured using one or more IP cores to use, support and / or implement the multi-format GPU docking board described herein. In addition to the illustrations, other logic and circuits may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general processor cores.

[0190] Figures 10B-10C is a block diagram illustrating a graphics processor for use within a SoC to use, support and / or implement the multi-format GPU docking board described herein according to at least one embodiment described herein. In one example, a graphics processor can be used to implement the multi-format GPU docking board described herein because existing math engines can process multi-level neural networks faster. Fig. 10B A graphics processor 1010 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, in accordance with at least one embodiment. Fig. 10C An additional graphics processor 1040 of a system on a chip integrated circuit is shown, which may be manufactured using one or more IP cores, according to at least one embodiment. In at least one embodiment, Fig. 10B The graphics processor 1010 is a low power graphics processor core. In at least one embodiment, Fig. 10C The graphics processor 1040 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1010, 1040 can be Fig. 10A A variant of graphics processor 1010.

[0191] In at least one embodiment, the graphics processor 1010 includes a vertex processor 1005 and one or more fragment processors 1015A-1015N (e.g., 1015A, 1015B, 1015C, 1015D to 1015N-1 and 1015N). In at least one embodiment, the graphics processor 1010 can execute different shader programs via separate logic, so that the vertex processor 1005 is optimized to perform operations for the vertex shader program, while one or more fragment processors 1015A-1015N perform fragment (e.g., pixel) shading operations for fragments or pixels or shader programs. In at least one embodiment, the vertex processor 1005 performs the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the one or more fragment processors 1015A-1015N use the primitives and vertex data generated by the vertex processor 1005 to generate a frame buffer displayed on a display device. In at least one embodiment, one or more fragment processors 1015A-1015N are optimized to execute fragment shader programs as provided in the OpenGL API, which can be used to perform similar operations as pixel shader programs provided in the Direct 3D API.

[0192] In at least one embodiment, graphics processor 1010 additionally includes one or more memory management units (MMUs) 1020A-1020B, one or more caches 1025A-1025B, and one or more circuit interconnects 1030A-1030B. In at least one embodiment, one or more MMUs 1020A-1020B provide a mapping of virtual to physical addresses for graphics processor 1010, including for vertex processor 1005 and / or fragment processors 1015A-1015N, which may reference vertex or image / texture data stored in memory in addition to vertex or image / texture data stored in one or more caches 1025A-1025B. In at least one embodiment, one or more MMUs 1020A-1020B may synchronize with other MMUs within the system, including with Fig. 10A One or more MMUs associated with one or more application processors 1005, image processor 1015, and / or video processor 1020 enable each processor 1005-1020 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 1030A-1030B enable graphics processor 1010 to connect to other IP cores within the SoC via an internal bus of the SoC or via a direct connection.

[0193] In at least one embodiment, graphics processor 1040 includes Fig. 10AThe graphics processor 1040 may include one or more MMUs 1020A-1020B, caches 1025A-1025B, and circuit interconnects 1030A-1030B of the graphics processor 1010. In at least one embodiment, the graphics processor 1040 includes one or more shader cores 1055A-1055N (e.g., 1055A, 1055B, 1055C, 1055D, 1055E, 1055F to 1055N-1 and 1055N), such as Fig. 10B As shown, it provides a unified shader kernel architecture in which a single kernel or type or kernel can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, multiple shader kernels can vary. In at least one embodiment, the graphics processor 1040 includes an inter-core task manager 1045 that acts as a thread dispatcher to dispatch execution threads to one or more shader kernels 1055A-1055N and a blocking unit 1058 to accelerate tile-based rendering operations, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial consistency within a scene or to optimize the use of internal caches.

[0194] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the inference and / or training logic 615. In at least one embodiment, the inference and / or training logic 615 may be implemented in an integrated circuit. Fig. 10A and / or Fig. 10B for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function or architecture, or a neural network use case herein.

[0195] Figures 10D-10E Additional graphics processor logic is shown according to the embodiments described herein for use, support and / or implementation of the multi-format GPU docking board described herein. In at least one embodiment, Fig. 10D shows that it can be included in Fig. 10A The graphics kernel 1000D within the graphics processor 1010 of FIG. 1000D may be configured as follows: Fig. 10C Unified shader cores 1055A-1055N are shown. Fig. 10B A highly parallel general purpose graphics processing unit ("GPGPU") 1030 suitable for deployment on a multi-chip module in at least one embodiment is shown.

[0196] In at least one embodiment, graphics core 1000D may include multiple slices 1001A-1001N or partitions of each core, and the graphics processor may include multiple instances of graphics core 1000D. In at least one embodiment, slices 1001A-1001N may include support logic including local instruction caches 1004A-1004N, thread schedulers 1006A-1006N, thread dispatchers 1008A-1008N, and a set of registers 1010A-1010N. In at least one embodiment, slices 1001A-1001N may include a set of additional function units (AFUs 1012A-1012N), floating point units (FPUs 1014A-1014N), integer arithmetic logic units (ALUs 109A-109N), address calculation units (ACUs 1013A-1013N), double precision floating point units (DPFPUs 1015A-1015N), and matrix processing units (MPUs 1017A-1017N).

[0197] In at least one embodiment, the FPU 1014A-1014N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 1015A-1015N performs double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU 1016A-1016N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. In at least one embodiment, the MPU 1017A-1017N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU 1017A-1010N can perform various matrix operations to accelerate machine learning application frameworks, including enabling general matrix-to-matrix multiplication (GEMM) to support acceleration. In at least one embodiment, the AFU 1012A-1012N can perform additional logical operations that are not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0198] As discussed elsewhere in this disclosure, the reasoning and / or training logic 615 (at least in Figure 6B , Figure 6C Reference) is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics kernel 1000D to infer or predict operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases herein.

[0199] Fig.11A A block diagram of a computer system 1100A according to at least one embodiment is shown. In at least one embodiment, the computer system 1100A includes a processing subsystem 1101 having one or more processors 1102 and a system memory 1104, which communicates via an interconnect path that may include a memory hub 1105. In at least one embodiment, the memory hub 1105 may be a separate component within a chipset component, or may be integrated within one or more processors 1102. In at least one embodiment, the memory hub 1105 is coupled to an I / O subsystem 1111 via a communication link 1106. In one embodiment, the I / O subsystem 1111 includes an I / O hub 1107, which may enable the computer system 1100A to receive input from one or more input devices 1108. In at least one embodiment, the I / O hub 1107 may enable a display controller, which may be included in one or more processors 1102, to provide output to the one or more display devices 1110A. In at least one embodiment, the one or more display devices 1110A coupled to the I / O hub 1107 may include local, internal, or embedded display devices.

[0200] In at least one embodiment, the processing subsystem 1101 includes one or more parallel processors 1112 coupled to the memory hub 1105 via a bus or other communication link 1113. In at least one embodiment, the communication link 1113 can use any of a number of standard-based communication link technologies or protocols, such as but not limited to PCI Express, or can be a vendor-specific communication interface or communication structure. In at least one embodiment, the one or more parallel processors 1112 form a parallel or vector processing system in a computing concentration, and the system can include a large number of processing cores and / or processing clusters, such as a multi-integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 1112 form a graphics processing subsystem, which can output pixels to one of the one or more display devices 1110A coupled via the I / O hub 1107. In at least one embodiment, the one or more parallel processors 1112 can also include a display controller and a display interface (not shown) to enable direct connection to one or more display devices 1110B.

[0201] In at least one embodiment, a system storage unit 1114 can be connected to the I / O hub 1107 to provide a storage mechanism for the computer system 1100A. In at least one embodiment, an I / O switch 1116 can be used to provide an interface mechanism to enable connections between the I / O hub 1107 and other components, such as a network adapter 1118 and / or a wireless network adapter 1119 that can be integrated into one or more platforms, as well as various other devices that can be added through one or more additional devices 1120. In at least one embodiment, the network adapter 1118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 1119 can include one or more of Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more radio devices.

[0202] In at least one embodiment, the computer system 1100A may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc. Other components may also be connected to the I / O hub 1107. In at least one embodiment, the interconnection may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express) or other bus or point-to-point communication interface and / or protocol. Fig.11A The communication paths between the various components in the NV-Link high-speed interconnect or interconnect protocol.

[0203] In at least one embodiment, one or more parallel processors 1112 include circuits optimized for graphics and video processing, including, for example, video output circuits, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 1112 include circuits optimized for general processing. In at least one embodiment, the components of the computer system 1100A can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 1112, memory hub 1105, one or more processors 1102, and I / O hub 1107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of the computer system 1100A can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of the computer system 1100A can be integrated into a multi-chip module (MCM), and the multi-chip module can be interconnected with other multi-chip modules into a modular computer system.

[0204] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provide details about the reasoning and / or training logic 615. In at least one embodiment, the reasoning and / or training logic 615 may be Fig.11A for use in a system for performing inference or prediction operations based at least in part on weight parameters computed using a neural network training operation, a neural network function and / or architecture, or a neural network use case herein.

[0205] processor

[0206] Fig. 11B 100B according to at least one embodiment. In at least one embodiment, various components of the parallel processor 1100B may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the parallel processor 1100B shown is a processor according to at least one embodiment. Fig. 11B A variation of the one or more parallel processors 1112 is shown.

[0207] In at least one embodiment, parallel processor 1100B includes parallel processing unit 1102. In at least one embodiment, parallel processing unit 1102 includes I / O unit 1104, which enables communication with other devices, including other instances of parallel processing unit 1102. In at least one embodiment, I / O unit 1104 can be directly connected to other devices. In at least one embodiment, I / O unit 1104 is connected to other devices by using a hub or switch interface (e.g., memory hub 1105). In at least one embodiment, the connection between memory hub 1105 and I / O unit 1104 forms communication link 1113. In at least one embodiment, I / O unit 1104 is connected to host interface 1106 and memory crossbar switch 1116, wherein host interface 1106 receives commands for performing processing operations and memory crossbar switch 1116 receives commands for performing memory operations.

[0208] In at least one embodiment, when the host interface 1106 receives the command buffer via the I / O unit 1104, the host interface 1106 can direct work operations to execute those commands to the front end 1108. In at least one embodiment, the front end 1108 is coupled with a scheduler 1110, which is configured to distribute commands or other work items to the processing cluster array 1112. In at least one embodiment, the scheduler 1110 ensures that the processing cluster array 1112 is properly configured and in a valid state before assigning tasks to the processing cluster array 1112. In at least one embodiment, the scheduler 1110 is implemented by firmware logic executed on a microcontroller. In at least one embodiment, the microcontroller-implemented scheduler 1110 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularity, thereby achieving fast preemption and context switching of threads executed on the processing array 1112. In at least one embodiment, the host software can prove the workload for scheduling on the processing array 1112 through one of the multiple graphics processing doorbells. In at least one embodiment, the workload may then be automatically distributed across the processing array 1112 by scheduler 1110 logic within a microcontroller that includes scheduler 1110 .

[0209] In at least one embodiment, the processing cluster array 1112 may include up to "N" processing clusters (e.g., cluster 1114A, cluster 1114B, through cluster 1114N). In at least one embodiment, each cluster 1114A-1114N of the processing cluster array 1112 may execute a large number of concurrent threads. In at least one embodiment, the scheduler 1110 may allocate work to the clusters 1114A-1114N of the processing cluster array 1112 using various scheduling and / or work allocation algorithms, which may vary depending on the workload generated by each program or type of calculation. In at least one embodiment, scheduling may be handled dynamically by the scheduler 1110, or may be partially assisted by compiler logic during the compilation of program logic configured to be executed by the processing cluster array 1112. In at least one embodiment, different clusters 1114A-1114N of the processing cluster array 1112 may be allocated to process different types of programs or to perform different types of calculations.

[0210] In at least one embodiment, processing cluster array 1112 can be configured to perform various types of parallel processing operations. In at least one embodiment, processing cluster array 1112 is configured to perform general-purpose parallel computing operations. In at least one embodiment, processing cluster array 1112 can include logic to perform processing tasks including filtering video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0211] In at least one embodiment, processing cluster array 1112 is configured to perform parallel graphics processing operations. In at least one embodiment, processing cluster array 1112 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic to perform texture operations, as well as inlay logic and other vertex processing logic. In at least one embodiment, processing cluster array 1112 may be configured to execute shader programs related to graphics processing, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, parallel processing units 1102 may transfer data from system memory via I / O units 1104 for processing. In at least one embodiment, during processing, the transferred data may be stored to on-chip memory (e.g., parallel processor memory 1122) during processing and then written back to system memory.

[0212] In at least one embodiment, when parallel processing unit 1102 is used to perform graphics processing, scheduler 1110 can be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 1114A-1114N of processing cluster array 1112. In at least one embodiment, portions of processing cluster array 1112 can be configured to perform different types of processing. In at least one embodiment, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform inlay and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to generate a rendered image for display while using, supporting and / or implementing the multi-format GPU docking board described herein. In at least one embodiment, intermediate data generated by one or more of clusters 1114A-1114N can be stored in a buffer to allow the intermediate data to be transferred between clusters 1114A-1114N for further processing.

[0213] In at least one embodiment, the processing cluster array 1112 can receive processing tasks to be performed via the scheduler 1110, which receives commands defining the processing tasks from the front end 1108. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how to process the data (e.g., what program to execute). In at least one embodiment, the scheduler 1110 can be configured to obtain an index corresponding to a task, or can receive the index from the front end 1108. In at least one embodiment, the front end 1108 can be configured to ensure that the processing cluster array 1112 is configured to a valid state before starting a workload specified by an incoming command buffer (e.g., a batch-buffer, a push buffer, etc.).

[0214] In at least one embodiment, each of the one or more instances of PPU 1102 may be coupled to PPU memory 1122. In at least one embodiment, PPU memory 1122 may be accessed via memory crossbar switch 1116, which may receive memory requests from processing cluster array 1112 and I / O unit 1104. In at least one embodiment, memory crossbar switch 1116 may access PPU memory 1122 via memory interface 1118. In at least one embodiment, memory interface 1118 may include a plurality of partition units (e.g., partition unit 1120A, partition unit 1120B to partition unit 1120N), which may each be coupled to a portion of PPU memory 1122 (e.g., memory units). In at least one embodiment, the plurality of partition units 1120A-1120N are configured to be equal to the number of memory cells, such that the first partition unit 1120A has a corresponding first memory cell 1124A, the second partition unit 1120B has a corresponding memory cell 1124B, and the Nth partition unit 1120N has a corresponding Nth memory cell 1124N. In at least one embodiment, the number of partition units 1120A-1120N may not be equal to the number of memory devices.

[0215] In at least one embodiment, memory units 1124A-1124N may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. In at least one embodiment, memory units 1124A-1124N may also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). In at least one embodiment, rendering targets such as frame buffers or texture maps may be stored across memory units 1124A-1124N, allowing partition units 1120A-1120N to write portions of each rendering target in parallel to efficiently use the available bandwidth of parallel processor memory 1122. In at least one embodiment, local instances of parallel processor memory 1122 may be excluded to facilitate a unified memory design that utilizes system memory in combination with local cache memory.

[0216] In at least one embodiment, any of the clusters 1114A-1114N of the processing cluster array 1112 can process data to be written to any memory unit 1124A-1124N within the parallel processor memory 1122. In at least one embodiment, the memory crossbar switch 1116 can be configured to transmit the output of each cluster 1114A-1114N to any partition unit 1120A-1120N or another cluster 1114A-1114N, and the cluster 1114A-1114N can perform other processing operations on the output. In at least one embodiment, each cluster 1114A-1114N can communicate with the memory interface 1118 through the memory crossbar switch 1116 to read from or write to various external storage devices. In at least one embodiment, memory crossbar switch 1116 has connections to memory interface 1118 to communicate with I / O unit 1104, and connections to local instances of parallel processor memory 1102, thereby enabling processing units within different processing clusters 1114A-1114N to communicate with system memory or other memory that is not local to parallel processing unit 1102. In at least one embodiment, memory crossbar switch 1116 can use virtual channels to separate traffic flows between clusters 1114A-1114N and partition units 1120A-1120N.

[0217] In at least one embodiment, multiple instances of parallel processing unit 1102 may be provided on a single plug-in card, or multiple plug-in cards may be interconnected. In at least one embodiment, different instances of parallel processing unit 1102 may be configured to interoperate, even if different instances have different numbers of processing cores, different numbers of local parallel processor memories, and / or other configuration differences. In at least one embodiment, some instances of parallel processing unit 1102 may include higher precision floating point units relative to other instances. In at least one embodiment, a system incorporating one or more instances of parallel processing unit 1102 or parallel processor 1100B may be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld personal computers, servers, workstations, game consoles, and / or embedded systems.

[0218] Fig. 11C is a block diagram of a partition unit 1120 according to at least one embodiment. In at least one embodiment, the partition unit 1120 is Fig. 11B1120N. In at least one embodiment, partition unit 1120 includes L2 cache 1121, frame buffer interface 1125, and ROP 1126 (raster operation unit). L2 cache 1121 is a read / write cache that is configured to perform load and store operations received from memory crossbar switch 1116 and ROP 1126. In at least one embodiment, L2 cache 1121 outputs read misses and urgent write-back requests to frame buffer interface 1125 for processing. In at least one embodiment, updates may also be sent to the frame buffer via frame buffer interface 1125 for processing. In at least one embodiment, frame buffer interface 1125 communicates with memory units (such as ROPs) in parallel processor memory. Fig. 11B Interact with one of the memory units 1124A-1124N (e.g., within parallel processor memory 1122).

[0219] In at least one embodiment, ROP 1126 is a processing unit that performs raster operations such as stenciling, z-testing, blending, and the like. In at least one embodiment, ROP 1126 then outputs processed graphics data that is stored in graphics memory. In at least one embodiment, ROP 1126 includes compression logic to compress depth or color data written to memory and decompress depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The compression logic performed by ROP 1126 can vary based on the statistical characteristics of the data to be compressed. In at least one embodiment, incremental color compression is performed based on the depth and color data on a per-tile basis.

[0220] In at least one embodiment, ROP 1126 is included within each processing cluster (e.g., Fig. 11B In at least one embodiment, read and write requests for pixel data are transmitted through memory crossbar switch 1116 rather than pixel fragment data transfer. In at least one embodiment, the processed graphics data can be displayed on a display device (such as Fig.11A 110), routed by processor 1102 for further processing, or by Fig. 11B One of the processing entities within parallel processor 1100B is routed for further processing.

[0221] Fig.11D is a block diagram of a processing cluster 1114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, a processing cluster is Fig. 11BIn at least one embodiment, one or more processing clusters 1114 can be configured to execute many threads in parallel, where a "thread" refers to an instance of a specific program executed on a specific set of input data. In at least one embodiment, single instruction multiple data (SIMD) instruction issuance technology is used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In at least one embodiment, single instruction multiple thread (SIMT) technology is used to support the parallel execution of a large number of synchronized threads, which uses a common instruction unit that is configured to issue instructions to a set of processing engines within each processing cluster.

[0222] In at least one embodiment, the operation of the processing cluster 1114 can be controlled by a pipeline manager 1132 that allocates processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 1132 Fig. 11B The scheduler 1110 receives instructions and manages the execution of these instructions through the graphics multiprocessor 1134 and / or the texture unit 1136. In at least one embodiment, the graphics multiprocessor 1134 is at least one instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors of different architectures may be included in the processing cluster 1114. In at least one embodiment, one or more instances of the graphics multiprocessor 1134 may be included in the processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 may process data, and the data crossbar switch 1140 may be used to distribute the processed data to one of a plurality of possible destinations (including other shader units). In at least one embodiment, the pipeline manager 1132 may facilitate the distribution of processed data by specifying the destination of the processed data to be distributed via the data crossbar switch 1140.

[0223] In at least one embodiment, each graphics multiprocessor 1134 within a processing cluster 1114 may include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic may be configured in a pipelined manner, where new instructions may be issued before previous instructions are completed. In at least one embodiment, the function execution logic supports a variety of operations, including integer and floating point arithmetic, comparison operations, Boolean operations, shifts, and calculations of various algebraic functions. In at least one embodiment, the same functional unit hardware may be utilized to perform different operations, and any combination of functional units may be present.

[0224] In at least one embodiment, the instructions transmitted to the processing cluster 1114 constitute threads. In at least one embodiment, a group of threads executed across a group of parallel processing engines is a thread group. In at least one embodiment, the thread group executes the program on different input data. In at least one embodiment, each thread in the thread group can be assigned to a different processing engine in the graphics multiprocessor 1134. In at least one embodiment, the thread group may include fewer threads than the number of processing engines in the graphics multiprocessor 1134. In at least one embodiment, when the number of threads included in the thread group is less than the number of processing engines, one or more processing engines may be idle during the cycle of the thread group being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines in the graphics multiprocessor 1134. In at least one embodiment, when the thread group includes more threads than the processing engines in the graphics multiprocessor 1134, processing can be performed in consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on the graphics multiprocessor 1134.

[0225] In at least one embodiment, graphics multiprocessor 1134 includes internal cache memory to perform load and store operations. In at least one embodiment, graphics multiprocessor 1134 can abandon the internal cache and use cache memory (e.g., L1 cache 1148) within processing cluster 1114. In at least one embodiment, each graphics multiprocessor 1134 can also access partition units (e.g., Fig. 11B 1120N) that are shared between all processing clusters 1114 and can be used to transfer data between threads. In at least one embodiment, graphics multiprocessor 1134 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to parallel processing unit 1102 can be used as global memory. In at least one embodiment, processing cluster 1114 includes multiple instances of graphics multiprocessor 1134, which can share common instructions and data that can be stored in L1 cache 1148.

[0226] In at least one embodiment, each processing cluster 1114 may include a memory management unit ("MMU") 1145 configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of MMU 1145 may reside in Fig. 11B1148. In at least one embodiment, the MMU 1145 includes a set of page table entries (PTEs) that are used to map virtual addresses to physical addresses of tiles and, in at least one embodiment, to cache line indices. In at least one embodiment, the MMU 1145 may include an address translation lookaside buffer (TLB) or a cache that may reside within the graphics multiprocessor 1134 or the L1 cache 1148 or the processing cluster 1114. In at least one embodiment, the physical addresses are processed to assign surface data access locality for efficient request interleaving between partition units. In at least one embodiment, the cache line index may be used to determine whether a request for a cache line is a hit or a miss.

[0227] In at least one embodiment, the processing clusters 1114 can be configured such that each graphics multiprocessor 1134 is coupled to a texture unit 1136 to perform texture mapping operations, operations to determine texture sample locations, read texture data, and filter texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multiprocessor 1134, and texture data is retrieved from an L2 cache, local parallel processor memory, or system memory as needed. In at least one embodiment, each graphics multiprocessor 1134 outputs one or more processed tasks to a data crossbar switch 1140 to provide the processed tasks to another processing cluster 1114 for further processing or to store the processed one or more tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar switch 1116. In at least one embodiment, a preROP 1142 (pre-raster operation unit) is configured to receive data from the graphics multiprocessor 1134, direct the data to a ROP unit, which can communicate with a partition unit (e.g., Fig. 11B In at least one embodiment, the PreROP 1142 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0228] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processing cluster 1114 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions, and / or architectures or neural network use cases described herein.

[0229] Fig.11E A graphics multiprocessor 1134 is shown in accordance with at least one embodiment. In at least one embodiment, the graphics multiprocessor 1134 is coupled to a pipeline manager 1132 of a processing cluster 1114. In at least one embodiment, the graphics multiprocessor 1134 has an execution pipeline that includes, but is not limited to, an instruction cache 1152, an instruction unit 1154, an address mapping unit 1156, a register file 1158, one or more general purpose graphics processing unit (GPGPU) cores 1162, and one or more load / store units 1166. The one or more GPGPU cores 1162 and the one or more load / store units 1166 are coupled to a cache memory 1172 and a shared memory 1170 via a memory and cache interconnect 1168.

[0230] In at least one embodiment, the instruction cache 1152 receives a stream of instructions to be executed from the pipeline manager 1132. In at least one embodiment, the instructions are cached in the instruction cache 1152 and dispatched for execution by the instruction unit 1154. In one embodiment, the instruction unit 1154 can dispatch instructions as thread groups (e.g., warps), each thread group being assigned to a different execution unit within one or more GPGPU cores 1162. In at least one embodiment, the instructions can access any local, shared, or global address space by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 1156 can be used to convert addresses in the unified address space into different memory addresses that can be accessed by one or more load / store units 1166.

[0231] In at least one embodiment, register file 1158 provides a set of registers for the functional units of graphics multiprocessor 1134. In at least one embodiment, register file 1158 provides temporary storage for operands for data paths connected to the functional units (e.g., GPGPU core 1162, load / store unit 1166) of graphics multiprocessor 1134. In at least one embodiment, register file 1158 is divided between each functional unit such that a dedicated portion of register file 1158 is allocated to each functional unit. In at least one embodiment, register file 1158 is divided between different warps being executed by graphics multiprocessor 1134.

[0232] In at least one embodiment, the GPGPU cores 1162 may each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions of the graphics multiprocessor 1134. The GPGPU cores 1162 may be similar in architecture or the architecture may be different. In at least one embodiment, the first portion of the GPGPU core 1162 includes a single-precision FPU and an integer ALU, while the second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. In at least one embodiment, the graphics multiprocessor 1134 may additionally include one or more fixed-function or special-function units to perform specific functions, such as copying rectangles or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores may also include fixed or special-function logic.

[0233] In at least one embodiment, the GPGPU kernel 1162 includes SIMD logic capable of executing a single instruction to multiple sets of data. In one embodiment, the GPGPU kernel 1162 can physically execute SIMD4, SIMD8 and SIMD16 instructions, and logically execute SIMD1, SIMD2 and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU kernel can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for a SIMT execution model can be executed by a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads that perform the same or similar operations can be executed in parallel by a single SIMD8 logic unit.

[0234] In at least one embodiment, the memory and cache interconnect 1168 is an interconnect network that connects each functional unit of the graphics multiprocessor 1134 to the register file 1158 and the shared memory 1170. In at least one embodiment, the memory and cache interconnect 1168 is a crossbar switch interconnect that allows the load / store unit 1166 to implement load and store operations between the shared memory 1170 and the register file 1158. In at least one embodiment, the register file 1158 can operate at the same frequency as the GPGPU core 1162, so that the latency of data transfer between the GPGPU core 1162 and the register file 1158 is very low. In at least one embodiment, the shared memory 1170 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1134. In at least one embodiment, the cache memory 1172 can be used as, for example, a data cache to cache texture data communicated between the functional units and the texture unit 1136. In at least one embodiment, the shared memory 1170 can also be used as a program-managed cache. In at least one embodiment, in addition to automatically cached data stored in cache memory 1172, threads executing on GPGPU core 1162 may programmatically store data in shared memory.

[0235] In at least one embodiment, a parallel processor or GPGPU as described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be coupled to a host / processor core via a bus or other interconnect (e.g., such as The GPU may be integrated with the core in the same package or chip and may be communicatively coupled to the core via an internal processor bus / interconnect (inside the package or chip in at least one embodiment). In at least one embodiment, regardless of how the GPU is connected, the processor core may assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0236] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics multiprocessor 1134 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0237] 12A illustrates a multi-GPU computing system 1200A according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 1200A may include a processor 1202 coupled to a plurality of general purpose graphics processing units (GPGPUs) 1206A-D via a host interface switch 1204. In at least one embodiment, the host interface switch 1204 is a PCI Express switch device that couples the processor 1202 to a PCI Express bus, and the processor 1202 may communicate with the GPGPUs 1206A-D via the PCI Express bus. The GPGPUs 1206A-D may be interconnected via a set of high-speed P2P GPU-to-GPU links 1216. In at least one embodiment, the GPU-to-GPU links 1216 are connected to each of the GPGPUs 1206A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU links 1216 enable direct communication between each GPGPU 1206A-D without communicating through the host interface bus 1204 to which the processor 1202 is connected. In at least one embodiment, host interface bus 1204 remains available for system memory access or communication with other instances of multi-GPU computing system 1200A, e.g., via one or more network devices, with GPU-to-GPU traffic directed to P2P GPU link 1216. While in at least one embodiment, GPGPUs 1206A-D are connected to processor 1202 via host interface switch 1204, in at least one embodiment, processor 1202 includes direct support for P2P GPU link 1216 and can connect directly to GPGPUs 1206A-D.

[0238] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in multi-GPU computing system 1200A to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0239] Fig. 12B 1 is a block diagram of a graphics processor 1200B according to at least one embodiment. In at least one embodiment, graphics processor 1200B includes ring interconnect 1202, pipeline front end 1204, media engine 1237, and graphics cores 1280A-1280N. In at least one embodiment, ring interconnect 1202 couples graphics processor 1200B to other processing units, including other graphics processors or one or more general-purpose processor cores. In at least one embodiment, graphics processor 1200B is one of many processors integrated within a multi-core processing system.

[0240] In at least one embodiment, graphics processor 1200B receives batches of commands via ring interconnect 1202. In at least one embodiment, the incoming commands are interpreted by command streamer 1203 in pipeline front end 1204. In at least one embodiment, graphics processor 1200B includes scalable execution logic for performing 3D geometry processing and media processing via graphics cores 1280A-1280N. In at least one embodiment, for 3D geometry processing commands, command streamer 1203 provides commands to geometry pipeline 1236. In at least one embodiment, for at least some media processing commands, command streamer 1203 provides commands to video front end 1234, which is coupled to media engine 1237. In at least one embodiment, media engine 1237 includes video quality engine (VQE) 1230 for video and image post-processing, and multi-format encoding / decoding (MFX) 1233 engine for providing hardware accelerated media data encoding and decoding. In at least one embodiment, geometry pipeline 1236 and media engine 1237 each generate execution threads for thread execution resources provided by at least one graphics core 1280A.

[0241] In at least one embodiment, the graphics processor 1200B includes scalable thread execution resources featuring modular cores 1280A-1280N (sometimes referred to as core slices), each graphics core having multiple sub-cores 1250A-1250N, 1260A-1260N (sometimes referred to as core sub-slices). In at least one embodiment, the graphics processor 1200B can have any number of graphics cores 1280A. In at least one embodiment, the graphics processor 1200B includes a graphics core 1280A having at least a first sub-core 1250A and a second sub-core 1260A. In at least one embodiment, the graphics processor 1200B is a low-power processor having a single sub-core (e.g., 1250A). In at least one embodiment, the graphics processor 1200B includes multiple graphics cores 1280A-1280N, each graphics core including a group of first sub-cores 1250A-1250N and a group of second sub-cores 1260A-1260N. In at least one embodiment, each of the first sub-cores 1250A-1250N includes at least a first set of execution units 1252A-1252N and a media / texture sampler 1254A-1254N. In at least one embodiment, each of the second sub-cores 1260A-1260N includes at least a second set of execution units 1262A-1262N and a sampler 1264A-1264N. In at least one embodiment, each of the sub-cores 1250A-1250N, 1260A-1260N shares a set of shared resources 1270A-1270N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.

[0242] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, inference and / or training logic 615 may be used in graphics processor 1200B to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0243] Fig.13is a block diagram of a microarchitecture for a processor 1300 that may include logic circuitry for executing instructions, according to an illustration of at least one embodiment. In at least one embodiment, the processor 1300 may execute instructions, including x86 instructions, ARM instructions, special instructions for an application-specific integrated circuit (ASIC), and the like. In at least one embodiment, the processor 1300 may include registers for storing packed data, such as 64-bit wide MMXTM registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, the MMX registers available in integer and floating point form may operate with packed data elements that accompany single instruction multiple data (“SIMD”) and streaming SIMD extension (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers associated with SSE2, SSE3, SSE4, AVX, or higher (generally referred to as “SSEx”) technology may hold such packed data operands. In at least one embodiment, the processor 1300 may execute instructions to accelerate machine learning or deep learning algorithms, training, or reasoning.

[0244] In at least one embodiment, the processor 1300 includes an in-order front end ("front end") 1301 to fetch instructions to be executed and prepare instructions for later use in the processor pipeline. In at least one embodiment, the front end 1301 may include several units. In at least one embodiment, an instruction prefetcher 1326 fetches instructions from memory and provides the instructions to an instruction decoder 1328, which in turn decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 1328 decodes the received instructions into one or more operations of so-called "microinstructions" or "microoperations" (also referred to as "micro-operations" or "microinstructions") that the machine can execute. In at least one embodiment, the instruction decoder 1328 parses the instructions into opcodes and corresponding data and control fields, which can be used by the microarchitecture to perform operations according to at least one embodiment. In at least one embodiment, the trace cache 1330 can assemble the decoded microinstructions into a program-ordered sequence or trace in the microinstruction queue 1334 for execution. In at least one embodiment, when trace cache 1330 encounters a complex instruction, microcode ROM 1332 provides the microinstructions necessary to complete the operation.

[0245] In at least one embodiment, some instructions may be converted into a single micro-operation, while other instructions may require several micro-operations to complete the entire operation. In at least one embodiment, if more than four micro-operations are required to complete an instruction, the instruction decoder 1328 may access the microcode ROM 1332 to execute the instruction. In at least one embodiment, the instruction may be decoded into a small number of micro-operations for processing at the instruction decoder 1328. In at least one embodiment, if multiple micro-operations are required to complete the operation, the instruction may be stored in the microcode ROM 1332. In at least one embodiment, the trace cache 1330 references the entry point programmable logic array ("PLA") to determine the correct micro-instruction pointer for reading the microcode sequence from the microcode ROM 1332 to complete one or more instructions according to at least one embodiment. In at least one embodiment, after the microcode ROM 1332 completes the micro-operation sequencing of the instruction, the front end 1301 of the machine may resume fetching micro-operations from the trace cache 1330.

[0246] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 1303 can prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth and reorder the instruction flow to optimize performance as instructions go down the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 1303 includes, but is not limited to, an allocator / register renamer 1340, a memory microinstruction queue 1342, an integer / floating point microinstruction queue 1344, a memory scheduler 1346, a fast scheduler 1302, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 1304, and a simple floating point scheduler ("simple FP scheduler") 1306. In at least one embodiment, the fast scheduler 1302, the slow / general purpose floating point scheduler 1304, and the simple floating point scheduler 1306 are also collectively referred to as "microinstruction schedulers 1302, 1304, 1306". In at least one embodiment, the allocator / register renamer 1340 allocates machine buffers and resources required for each microinstruction to execute in sequence. In at least one embodiment, the allocator / register renamer 1340 renames logical registers into entries in the register file. In at least one embodiment, the allocator / register renamer 1340 also allocates entries for each microinstruction in one of the two microinstruction queues, the memory microinstruction queue 1342 for memory operations and the integer / floating point microinstruction queue 1344 for non-memory operations, in front of the memory scheduler 1346 and the microinstruction schedulers 1302, 1304, 1306. In at least one embodiment, the microinstruction schedulers 1302, 1304, 1306 determine when the microinstructions are ready to execute based on the readiness of their slave input register operand sources and the availability of the execution resource microinstructions that need to be completed. In at least one embodiment, the fast scheduler 1302 of at least one embodiment can be scheduled on each half of the main clock cycle, while the slow / general floating point scheduler 1304 and the simple floating point scheduler 1306 can be scheduled once per main processor clock cycle. In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 arbitrate dispatch ports to schedule microinstructions for execution.

[0247] In at least one embodiment, execution block 1311 includes, but is not limited to, integer register file / bypass network 1308, floating point register file / bypass network ("FP register file / bypass network") 1310, address generation units ("AGUs") 1312 and 1314, fast arithmetic logic units ("fast ALUs") 1316 and 1318, slow arithmetic logic unit ("slow ALU") 1320, floating point ALU ("FP") 1322, and floating point move unit ("FP move") 1324. In at least one embodiment, integer register file / bypass network 1308 and floating point register file / bypass network 1310 are also referred to herein as "register files 1308, 1310". In at least one embodiment, AGUs 1312 and 1314, fast ALUs 1316 and 1318, slow ALU 1320, floating point ALU 1322, and floating point move unit 1324 are also referred to herein as "execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324." In at least one embodiment, execution block 1311 may include, but is not limited to, any number (including zero) and type of register files, bypass networks, address generation units, and execution units (in any combination).

[0248] In at least one embodiment, register networks 1308, 1310 may be arranged between microinstruction schedulers 1302, 1304, 1306 and execution units 1312, 1314, 1316, 1318, 1320, 1322, and 1324. In at least one embodiment, integer register file / bypass network 1308 performs integer operations. In at least one embodiment, floating point register file / bypass network 1310 performs floating point operations. In at least one embodiment, each of register networks 1308, 1310 may include, but is not limited to, a branch network that may bypass or forward a just completed result that has not yet been written to the register file to a new slave object. In at least one embodiment, register networks 1308, 1310 may communicate data with each other. In at least one embodiment, integer register file / bypass network 1308 may include, but is not limited to, two separate register files, one register file for low-order 32-bit data, and a second register file for high-order 32-bit data. In at least one embodiment, floating point register file / bypass network 1310 may include, but is not limited to, 128-bit wide entries, since floating point instructions typically have operands that are 64 to 128 bits wide.

[0249] In at least one embodiment, execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 can execute instructions. In at least one embodiment, register files 1308, 1310 store integer and floating point data operand values ​​that microinstructions need to execute. In at least one embodiment, processor 1300 may include, but is not limited to, any number of execution units 1312, 1314, 1316, 1318, 1320, 1322, 1324 and combinations thereof. In at least one embodiment, floating point ALU 1322 and floating point move unit 1324 can perform floating point, MMX, SIMD, AVX and SSE or other operations, including specialized machine learning instructions. In at least one embodiment, floating point ALU 1322 may include, but is not limited to, a 64-bit by 64-bit floating point divider to perform division, square root and remainder micro-operations. In at least one embodiment, floating point hardware may be used to process instructions involving floating point values. In at least one embodiment, ALU operations can be passed to fast ALUs 1316, 1318. In at least one embodiment, fast ALUs 1316, 1318 can perform fast operations with an effective delay of half a clock cycle. In at least one embodiment, most complex integer operations go to slow ALU 1320 because slow ALU 1320 can include, but is not limited to, integer execution hardware for long-latency type operations, such as multipliers, shifts, flag logic, and branch processing. In at least one embodiment, memory load / store operations can be performed by AGUs 1312, 1314. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 1316, fast ALU 1318, and slow ALU 1320 can be implemented to support various data bit sizes including sixteen, thirty-two, 128, 256, etc. In at least one embodiment, the floating point ALU 1322 and floating point move unit 1324 can be implemented to support a range of operands having bits of various widths. In at least one embodiment, the floating point ALU 1322 and floating point move unit 1324 can operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0250] In at least one embodiment, microinstruction schedulers 1302, 1304, 1306 schedule dependent operations before the parent load completes execution. In at least one embodiment, since microinstructions can be speculatively scheduled and executed in processor 1300, processor 1300 can also include logic for handling memory misses. In at least one embodiment, if the data load in the data cache misses, there may be dependent operations running in the pipeline, which temporarily prevents the scheduler from having the correct data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use incorrect data. In at least one embodiment, it may be necessary to replay dependent operations and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.

[0251] In at least one embodiment, "register" may refer to an onboard processor storage location that can be used as part of an instruction to identify an operand. In at least one embodiment, registers may be those that can be used from outside the processor (from a programmer's perspective). In at least one embodiment, registers may not be limited to a particular type of circuit. On the contrary, in at least one embodiment, registers can store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein can be implemented using a variety of different techniques by circuits within the processor, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated and dynamically allocated physical registers, etc. In at least one embodiment, integer registers store 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packaging data.

[0252] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, part or all of the inference and / or training logic 615 may be incorporated into the execution block 1311 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs shown in the execution block 1311. In addition, weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of the execution block 1311 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0253] Fig.14A deep learning application processor 1400 is shown in accordance with at least one embodiment. In at least one embodiment, the deep learning application processor 1400 uses instructions that, if executed by the deep learning application processor 1400, cause the deep learning application processor 1400 to perform some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 1400 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 1400 performs matrix multiplication operations or is "hardwired" into hardware as a result of executing one or more instructions, or both. In at least one embodiment, the deep learning application processor 1400 includes, but is not limited to, processing clusters 1410(1)-1410(12), inter-chip links (“ICLs”) 1420(1)-1420(12), inter-chip controllers (“ICCs”) 1430(1)-1430(2), memory controllers (“Mem Ctrlr”) 1442(1)-1442(4), high bandwidth memory physical layer (“HBM PHY”) 1444(1)-1444(4), a management controller central processing unit (“management controller CPU”) 1450, serial peripheral interface, inter-integrated circuit and general purpose input / output blocks (“SPI, I2C, GPIO”), a peripheral component interconnect express controller and direct memory access block (“DMAC”). controller and DMA”) 1470, and a sixteen-channel Peripheral Component Interconnect Express port (“PCI Express x 16”) 1480.

[0254] In at least one embodiment, the processing cluster 1410 may perform deep learning operations, including inference or prediction operations based on weight parameters calculated based on one or more training techniques, including those of this article. In at least one embodiment, each processing cluster 1410 may include, but is not limited to, any number and type of processors. In at least one embodiment, the deep learning application processor 1400 may include any number and type of processing clusters 1400. In at least one embodiment, the inter-chip link 1420 is bidirectional. In at least one embodiment, the inter-chip link 1420 and the inter-chip controller 1430 enable multiple deep learning application processors 1400 to exchange information, including activation information generated from executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 1400 may include any number (including zero) and type of ICL 1420 and ICC 1430.

[0255] In at least one embodiment, HBM2 1440 provides a total of 32GB of memory. HBM2 1440(i) is associated with both memory controller 1442(i) and HBM PHY 1444(i). In at least one embodiment, any number of HBM2 1440 can provide any type and total amount of high bandwidth memory and can be associated with any number (including zero) and type of memory controller 1442 and HBM PHY 1444. In at least one embodiment, SPI, I2C, GPIO 3360, Controller 1460 and DMA 1470 and / or 1480, implementing any number and type of communications standards in any technically feasible manner.

[0256] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding the inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (e.g., a neural network) to predict or infer information provided to the deep learning application processor 1400. In at least one embodiment, the deep learning application processor 1400 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 1400. In at least one embodiment, the processor 1400 can be used to perform one or more of the neural network use cases described herein.

[0257] Fig.151 is a block diagram of a neuromorphic processor 1500 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 1500 may receive one or more inputs from a source external to the neuromorphic processor 1500. In at least one embodiment, these inputs may be transmitted to one or more neurons 1502 within the neuromorphic processor 1500. In at least one embodiment, the neurons 1502 and their components may be implemented using circuits or logic including one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, thousands of instances of neurons 1502, although any suitable number of neurons 1502 may be used. In at least one embodiment, each instance of a neuron 1502 may include a neuron input 1504 and a neuron output 1506. In at least one embodiment, a neuron 1502 may generate an output that may be transmitted to the inputs of other instances of the neuron 1502. In at least one embodiment, the neuron input 1504 and the neuron output 1506 may be interconnected via a synapse 1508.

[0258] In at least one embodiment, the neurons 1502 and synapses 1508 may be interconnected such that the neuromorphic processor 1500 operates to process or analyze information received by the neuromorphic processor 1500. In at least one embodiment, the neuron 1502 may send an output pulse (or "trigger" or "spike") when the input received through the neuron input 1504 exceeds a threshold. In at least one embodiment, the neuron 1502 may sum or integrate the signal received at the neuron input 1504. For example, in at least one embodiment, the neuron 1502 may be implemented as a leaky integrate-trigger neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, the neuron 1502 may generate an output (or "trigger") using a transfer function such as a sigmoid or threshold function. In at least one embodiment, the leaky integrate-trigger neuron may sum the signal received at the neuron input 1504 into a membrane potential, and may apply an application attenuation factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-trigger neuron may trigger if multiple input signals are received at the neuron input 1504 fast enough to exceed a threshold (in at least one embodiment, before the membrane potential decays too low to trigger). In at least one embodiment, the neuron 1502 may be implemented using circuitry or logic that receives inputs, integrates the inputs into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. In addition, in at least one embodiment, the neuron 1502 may include, but is not limited to, a comparator circuit or logic that generates an output spike at the neuron output 1506 when the result of applying the transfer function to the neuron input 1504 exceeds a threshold. In at least one embodiment, once the neuron 1502 triggers, it may ignore previously received input information by, for example, resetting the membrane potential to 0 or another suitable default value. In at least one embodiment, once the membrane potential is reset to 0, the neuron 1502 may resume normal operation after a suitable period of time (or recovery period).

[0259] In at least one embodiment, neurons 1502 may be interconnected via synapses 1508. In at least one embodiment, synapses 1508 may be operable to transmit a signal from an output of a first neuron 1502 to an input of a second neuron 1502. In at least one embodiment, a neuron 1502 may transmit information over more than one instance of synapse 1508. In at least one embodiment, one or more instances of a neuron output 1506 may be connected to an instance of a neuron input 1504 in the same neuron 1502 via an instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that produces an output to be transmitted over an instance of synapse 1508 may be referred to as a "pre-synaptic neuron" relative to that instance of synapse 1508. In at least one embodiment, an instance of a neuron 1502 that receives an input transmitted through an instance of synapse 1508 may be referred to as a "post-synaptic neuron" relative to an instance of synapse 1508. In at least one embodiment, with respect to various instances of synapses 1508, because an instance of neuron 1502 can receive input from one or more instances of synapses 1508 and can also transmit output through one or more instances of synapses 1508, a single instance of neuron 1502 can be both a "pre-synaptic neuron" and a "post-synaptic neuron."

[0260] In at least one embodiment, neurons 1502 may be organized into one or more layers. Each instance of a neuron 1502 may have a neuron output 1506 that may fan out to one or more neuron inputs 1504 through one or more synapses 1508. In at least one embodiment, a neuron output 1506 of a neuron 1502 in a first layer 1510 may be connected to a neuron input 1504 of a neuron 1502 in a second layer 1512. In at least one embodiment, the layers 1510 may be referred to as "feed-forward layers". In at least one embodiment, each instance of a neuron 1502 in an instance of the first layer 1510 may fan out to each instance of a neuron 1502 in a second layer 1512. In at least one embodiment, the first layer 1510 may be referred to as a "fully connected feed-forward layer". In at least one embodiment, each instance of a neuron 1502 in each instance of the second layer 1512 may fan out to less than all instances of a neuron 1502 in a third layer 1514. In at least one embodiment, the second layer 1512 may be referred to as a "sparsely connected feed-forward layer". In at least one embodiment, neurons 1502 in the (same) second layer 1512 may fan out to neurons 1502 in multiple other layers, including also fanning out to neurons 1502 in the second layer 1512. In at least one embodiment, the second layer 1512 may be referred to as a "recurrent layer". In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, any suitable combination of recurrent layers and feed-forward layers, including but not limited to sparsely connected feed-forward layers and fully connected feed-forward layers.

[0261] In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, a reconfigurable interconnect architecture or a dedicated hardwired interconnect to connect the synapses 1508 to the neurons 1502. In at least one embodiment, the neuromorphic processor 1500 may include, but is not limited to, circuitry or logic that allows synapses to be assigned to different neurons 1502 as needed, depending on the neural network topology and neuron fan-in / fan-out. For example, in at least one embodiment, the synapses 1508 may be connected to the neurons 1502 using an interconnect architecture such as a network on a chip or through dedicated connections. In at least one embodiment, the synaptic interconnects and their components may be implemented using circuitry or logic.

[0262] Fig.16AA processing system according to at least one embodiment is shown. In at least one embodiment, system 1600A includes one or more processors 1602 and one or more graphics processors 1608, and can be a single processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1602 or processor cores 1607. In at least one embodiment, system 1600A is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.

[0263] In at least one embodiment, the system 1600A may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console for gaming and media consoles. In at least one embodiment, the system 1600A is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1600A may also include a wearable device coupled to or integrated in a wearable device, such as a smart watch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1600A is a television or set-top box device having one or more processors 1602 and a graphical interface generated by one or more graphics processors 1608.

[0264] In at least one embodiment, one or more processors 1602 each include one or more processor cores 1607 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1607 is configured to process a specific instruction set 1609. In at least one embodiment, the instruction set 1609 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing through very long instruction words (VLIW). In at least one embodiment, the processor cores 1607 can each process a different instruction set 1609, which can include instructions that help emulate other instruction sets. In at least one embodiment, the processor cores 1607 can also include other processing devices, such as digital signal processors (DSPs).

[0265] In at least one embodiment, the processor 1602 includes a cache memory 1604. In at least one embodiment, the processor 1602 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared between various components of the processor 1602. In at least one embodiment, the processor 1602 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which may be shared between the processor cores 1607 using known cache coherence techniques. In at least one embodiment, the processor 1602 additionally includes a register file 1606, which may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 1606 may include general registers or other registers.

[0266] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transmit communication signals, such as address, data, or control signals, between the processor 1602 and other components in the system 1600A. In at least one embodiment, the interface bus 1610 can be a processor bus in one embodiment, such as a version of a direct media interface (DMI) bus. In at least one embodiment, the interface bus 1610 is not limited to a DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1602 includes an integrated memory controller 1616 and a platform controller hub 1630. In at least one embodiment, the memory controller 1616 facilitates communication between memory devices and other components of the processing system 1600A, while the platform controller hub (PCH) 1630 provides connections to input / output (I / O) devices through a local I / O bus.

[0267] In at least one embodiment, the memory device 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have appropriate performance to be used as processor memory. In at least one embodiment, the memory device 1620 may be used as a system memory of the processing system 1600A to store data 1622 and instructions 1621 for use when one or more processors 1602 execute applications or processes. In at least one embodiment, the memory controller 1616 is also coupled to an external graphics processor 1612 of at least one embodiment, which may communicate with one or more graphics processors 1608 in the processor 1602 to perform graphics and media operations. In at least one embodiment, the display device 1611 may be connected to the processor 1602. In at least one embodiment, the display device 1611 may include one or more of the internal display devices, such as in a mobile electronic device or laptop device or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1611 may include a head mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.

[0268] In at least one embodiment, the platform controller hub 1630 enables peripheral devices to be connected to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, a touch sensor 1625, a data storage device 1624 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1625 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1626 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or long-term evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can be, for example, a unified extensible firmware interface (UEFI). In at least one embodiment, the network controller 1634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1610. In at least one embodiment, the audio controller 1646 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1600A includes a legacy I / O controller 1640 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1600A. In at least one embodiment, the platform controller hub 1630 can also be connected to one or more universal serial bus (USB) controllers 1642, which connect input devices such as a keyboard and mouse 1643 combination, a camera 1644, or other USB input devices.

[0269] In at least one embodiment, instances of memory controller 1616 and platform controller hub 1630 may be integrated into a discrete external graphics processor, such as external graphics processor 1612. In at least one embodiment, platform controller hub 1630 and / or memory controller 1616 may be external to one or more processors 1602. In at least one embodiment, system 1600A may include external memory controller 1616 and platform controller hub 1630, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with processor 1602.

[0270] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C1600A. In at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs that are embodied in graphics processor 1612. In addition, in at least one embodiment, the inference and / or training operations described herein may use additional ALUs. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600A to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0271] Fig. 16B is a block diagram of a processor 1600B having one or more processor cores 1602A-1602N, an integrated memory controller 1614, and an integrated graphics processor 1608, according to at least one embodiment. In at least one embodiment, the processor 1600B may include additional cores, up to and including the additional core 1602N represented by the dashed box. In at least one embodiment, each processor core 1602A-1602N includes one or more internal cache units 1604A-1604N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1606.

[0272] In at least one embodiment, the internal cache units 1604A-1604N and the shared cache unit 1606 represent a cache memory hierarchy within the processor 1600B. In at least one embodiment, the cache memory units 1604A-1604N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1606 and 1604A-1604N.

[0273] In at least one embodiment, the processor 1600B may also include a set of one or more bus controller units 1616 and a system agent core 1610. In at least one embodiment, the one or more bus controller units 1616 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1610 provides management functions for various processor components. In at least one embodiment, the system agent core 1610 includes one or more integrated memory controllers 1614 to manage access to various external memory devices (not shown).

[0274] In at least one embodiment, one or more processor cores 1602A-1602N include support for multiple threads simultaneously. In at least one embodiment, system agent core 1610 includes components for coordinating and operating cores 1602A-1602N during multithreaded processing. In at least one embodiment, system agent core 1610 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of processor cores 1602A-1602N and graphics processor 1608.

[0275] In at least one embodiment, the processor 1600B also includes a graphics processor 1608 for performing image processing operations. In at least one embodiment, the graphics processor 1608 is coupled to a shared cache unit 1606 and a system agent core 1610 including one or more integrated memory controllers 1614. In at least one embodiment, the system agent core 1610 also includes a display controller 1611 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1611 may also be a separate module coupled to the graphics processor 1608 via at least one interconnect, or may be integrated within the graphics processor 1608.

[0276] In at least one embodiment, a ring-based interconnect unit 1612 is used to couple the internal components of the processor 1600B. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1608 is coupled to the ring interconnect 1612 via an I / O link 1613.

[0277] In at least one embodiment, I / O link 1613 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 1618 (e.g., eDRAM modules). In at least one embodiment, each of processor cores 1602A-1602N and graphics processor 1608 uses embedded memory modules 1618 as a shared last level cache.

[0278] In at least one embodiment, the processor cores 1602A-1602N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1602A-1602N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1602A-1602N execute a common instruction set, and one or more other processor cores 1602A-1602N execute a subset or a different instruction set of the common instruction set. In at least one embodiment, in terms of microarchitecture, the processor cores 1602A-1602N are heterogeneous, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1600B can be implemented on one or more chips or implemented as a SoC integrated circuit.

[0279] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C 615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the processor 1600B. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more ALUs embodied in Fig.16A In addition, in at least one embodiment, the inference and / or training operations described herein may use the inference and / or training operations described herein. Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 1600B to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0280] Fig. 16Cis a block diagram of the hardware logic of a graphics processor core 1600C according to at least one embodiment of the present invention. In at least one embodiment, the graphics processor core 1600C is included in a graphics core array. In at least one embodiment, the graphics processor core 1600C (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. In at least one embodiment, the graphics processor core 1600C is at least one embodiment of a graphics core slice, and the graphics processor herein can include multiple graphics core slices based on a target power and performance envelope. In at least one embodiment, each graphics core 1600C can include a fixed function block 1630 coupled to multiple sub-cores 1601A-1601F, also referred to as a sub-slice, which includes modular blocks of general and fixed function logic.

[0281] In at least one embodiment, fixed function block 1630 includes a geometry fixed function pipeline 1636, which can be shared by all sub-cores in graphics processor 1600C, for example, in lower performance and / or lower power graphics processor implementations. In at least one embodiment, geometry and fixed function pipeline 1636 includes a 3D fixed function pipeline, a video front end unit, a thread generator and thread dispatcher, and a unified return buffer manager that manages a unified return buffer.

[0282] In at least one embodiment of the fixed function block 1630, the fixed function block 1630 also includes a graphics SoC interface 1637, a graphics microcontroller 1638, and a media pipeline 1639. In at least one embodiment, the fixed graphics SoC interface 1637 provides an interface between the graphics kernel 1600C and other processor kernels in the integrated circuit system on a chip. In at least one embodiment, the graphics microcontroller 1638 is a programmable subprocessor that can be configured to manage various functions of the graphics processor 1600C, including thread dispatching, scheduling, and preemption. In at least one embodiment, the media pipeline 1639 includes logic that helps to decode, encode, pre-process, and / or post-process multimedia data including image and video data. In at least one embodiment, the media pipeline 1639 implements media operations via requests to calculation or sampling logic within the sub-kernels 1601-1601F.

[0283] In at least one embodiment, the SoC interface 1637 enables the graphics core 1600C to communicate with a general application processor core (e.g., a CPU) and / or other components within the SoC, including memory hierarchy elements such as a shared last level cache, system RAM, and / or embedded on-chip or packaged DRAM. In at least one embodiment, the SoC interface 1637 may also enable communication with fixed-function devices within the SoC (e.g., a camera imaging pipeline), and enable the use and / or implementation of global memory atomics that can be shared between the graphics core 1600C and the CPU within the SoC. In at least one embodiment, the SoC interface 1637 may also implement power management controls for the graphics core 1600C and enable interfaces between the clock domain of the graphics core 1600C and other clock domains within the SoC. In at least one embodiment, the SoC interface 1637 enables command buffers to be received from a command stream converter and a global thread dispatcher, which is configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. In at least one embodiment, commands and instructions may be dispatched to media pipeline 1639 when media operations are to be performed, or may be assigned to geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 1636, geometry and fixed function pipeline 1614) when graphics processing operations are to be performed.

[0284] In at least one embodiment, the graphics microcontroller 1638 can be configured to perform various scheduling and management tasks for the graphics core 1600C. In at least one embodiment, the graphics microcontroller 1638 can perform graphics and / or computing workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 1602A-1602F, 1604A-1604F in the sub-cores 1601A-1601F. In at least one embodiment, host software executed on a CPU core of a SoC including the graphics core 1600C can submit a workload to one of a plurality of graphics processor doorbells, which invokes a scheduling operation on the appropriate graphics engine. In at least one embodiment, the scheduling operation includes determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is completed. In at least one embodiment, graphics microcontroller 1638 may also facilitate low power or idle states for graphics core 1600C, thereby providing graphics core 1600C with the ability to save and restore registers across low power state transitions within graphics core 1600C independent of the operating system and / or graphics driver software on the system.

[0285] In at least one embodiment, the graphics core 1600C may have up to N modular sub-cores more or less than the sub-cores 1601A-1601F shown. For each group of N sub-cores, in at least one embodiment, the graphics core 1600C may also include shared function logic 1610, shared and / or cache memory 1612, geometry / fixed function pipelines 1614, and additional fixed function logic 1616 to accelerate various graphics and compute processing operations. In at least one embodiment, the shared function logic 1610 may include logic units (e.g., samplers, math and / or inter-thread communication logic) that may be shared by each of the N sub-cores within the graphics core 1600C. In at least one embodiment, the fixed, shared and / or cache memory 1612 may be the last level cache of the N sub-cores 1601A-1601F within the graphics core 1600C, and may also be used as a shared memory accessible by multiple sub-cores. In at least one embodiment, geometry / fixed function pipeline 1614 may be included in place of geometry / fixed function pipeline 1636 within fixed function block 1630 and may include similar logic units.

[0286] In at least one embodiment, the graphics kernel 1600C includes additional fixed function logic 1616, which may include various fixed function acceleration logic for use by the graphics kernel 1600C. In at least one embodiment, the additional fixed function logic 1616 includes additional geometry pipelines for use in position-only shading. In position-only shading, there are at least two geometry pipelines, and in the complete geometry pipeline and the culling pipeline within the geometry and fixed function pipelines 1614, 1636, it is an additional geometry pipeline that may be included in the additional fixed function logic 1616. In at least one embodiment, the culling pipeline is a trimmed version of the complete geometry pipeline. In at least one embodiment, the complete pipeline and the culling pipeline can execute different instances of an application, each with a separate environment. In at least one embodiment, position-only shading can hide long culling runs for discarded triangles, so that shading can be completed earlier in some cases. In at least one embodiment, the culling pipeline logic in the additional fixed function logic 1616 can execute the position shader in parallel with the main application and generate critical results faster than the full pipeline because the culling pipeline obtains and masks the position attributes of the vertex without performing rasterization and rendering the pixels to the frame buffer. In at least one embodiment, the culling pipeline can use the generated critical results to calculate visibility information for all triangles, regardless of whether they are culled. In at least one embodiment, the full pipeline (which can be called a replay pipeline in this case) can consume visibility information to skip culled triangles to mask only visible triangles that are ultimately passed to the rasterization stage.

[0287] In at least one embodiment, the additional fixed function logic 1616 may also include machine learning acceleration logic, such as fixed function matrix multiplication logic, for implementing optimizations including for machine learning training or inference.

[0288] In at least one embodiment, a set of execution resources is included within each graphics sub-core 1601A-1601F, which can be used to perform graphics, media, and compute operations in response to requests from a graphics pipeline, a media pipeline, or a shader program. In at least one embodiment, the graphics sub-core 1601A-1601F includes multiple EU arrays 1602A-1602F, 1604A-1604F, thread dispatch and inter-thread communication (TD / IC) logic 1603A-1603F, 3D (e.g., texture) samplers 1605A-1605F, media samplers 1606A-1606F, shader processors 1607A-1607F, and shared local memory (SLM) 1608A-1608F. Each of the EU arrays 1602A-1602F, 1604A-1604F includes a plurality of execution units, which are general purpose graphics processing units capable of servicing graphics, media, or compute operations, performing floating point and integer / fixed point logic operations, including graphics, media, or compute shader programs. In at least one embodiment, the TD / IC logic 1603A-1603F performs local thread dispatch and thread control operations for the execution units within the sub-core, and facilitates communication between threads executed on the execution units of the sub-core. In at least one embodiment, the 3D samplers 1605A-1605F can read data associated with textures or other 3D graphics into memory. In at least one embodiment, the 3D samplers can read texture data differently based on the configured sampling state and texture format associated with a given texture. In at least one embodiment, the media samplers 1606A-1606F can perform similar read operations based on the type and format associated with the media data. In at least one embodiment, each graphics sub-core 1601A-1601F may alternatively include a unified 3D and media sampler. In at least one embodiment, threads executing on execution units within each sub-core 1601A-1601F may utilize shared local memory 1608A-1608F within each sub-core to enable threads executing within a thread group to execute using a common pool of on-chip memory.

[0289] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C615. In at least one embodiment, some or all of the inference and / or training logic 615 may be incorporated into the graphics processor 1610. In at least one embodiment, the training and / or inference techniques described herein may be used in Fig. 16B In addition, in at least one embodiment, the inference and / or training operations described herein may use the addition of Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown), which configure the ALU of the graphics processor 1600C to execute one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0290] Figures 16D-16E Thread execution logic 1600D is shown for an array of processing elements including a graphics processor core, according to at least one embodiment. Fig.16D At least one embodiment is shown in which thread execution logic 1600D is used. Fig.16E Internal details of an execution unit according to at least one embodiment are shown.

[0291] like Fig.16DAs shown in , in at least one embodiment, thread execution logic 1600D includes a shader processor 1602, a thread dispatcher 1604, an instruction cache 1606, a scalable execution unit array including a plurality of execution units 1608A-1608N, one or more samplers 1610, a data cache 1612, and a data port 1614. In at least one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 1608A, 1608B, 1608C, 1608D, 1608N-1, and 1608N), for example, based on the computational requirements of the workload. In at least one embodiment, the scalable execution units are interconnected by an interconnect structure that links to each execution unit. In at least one embodiment, thread execution logic 1600D includes one or more connections to a memory (such as system memory or cache memory) through instruction cache 1606, data port 1614, sampler 1610, and one or more of execution units 1608A-1608N. In at least one embodiment, each execution unit (e.g., 1608A) is an independent programmable general-purpose computing unit that is capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In at least one embodiment, the array of execution units 1608A-1608N is scalable to include any number of separate execution units.

[0292] In at least one embodiment, execution units 1608A-1608N are primarily used to execute shader programs. In at least one embodiment, shader processor 1602 can process various shader programs and dispatch execution threads associated with shader programs via thread dispatcher 1604. In at least one embodiment, thread dispatcher 1604 includes logic for arbitrating thread initialization celebrations from graphics and media pipelines and instantiating requested threads on one or more execution units in execution units 1608A-1608N. In at least one embodiment, in at least one embodiment, the geometry pipeline can dispatch vertices, inlays, or geometry shaders to thread execution logic for processing. In at least one embodiment, thread dispatcher 1604 can also process runtime thread generation requests from executing shader programs.

[0293] In at least one embodiment, the execution units 1608A-1608N support an instruction set that includes native support for many standard 3D graphics shader instructions, allowing shader programs in graphics libraries (such as Direct 3D and OpenGL) to execute with minimal conversion. In at least one embodiment, the execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general processing (e.g., compute and media shaders). In at least one embodiment, each execution unit 1608A-1608N includes one or more arithmetic logic units (ALUs) capable of performing multiple-issue single instruction multiple data (SIMD), and multi-threaded operations enable an efficient execution environment despite higher latency memory access. In at least one embodiment, each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. In at least one embodiment, execution is multiple issues per clock to the pipeline, which is capable of integer, single-precision and double-precision floating point operations, SIMD branch functions, logical operations, prior operations, and other other operations. In at least one embodiment, while waiting for data from memory or one of the shared functions, dependency logic within execution units 1608A-1608N causes the waiting thread to sleep until the requested data is returned. In at least one embodiment, while the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. In at least one embodiment, during the delay associated with the vertex shader operation, the execution unit can perform operations on a pixel shader, a fragment shader, or another type of shader program (including a different vertex shader).

[0294] In at least one embodiment, each of the execution units 1608A-1608N operates on an array of data elements. In at least one embodiment, the number of data elements is the "execution size" or number of lanes of an instruction. In at least one embodiment, an execution lane is a logical unit used for the execution of data element access, masking, and flow control within an instruction. In at least one embodiment, the multiple lanes may be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) for a particular graphics processor. In at least one embodiment, the execution units 1608A-1608N support integer and floating point data types.

[0295] In at least one embodiment, the execution unit instruction set includes SIMD instructions. In at least one embodiment, various data elements can be stored in registers as packed data types, and the execution unit will process various elements based on the data size of those elements. In at least one embodiment, in at least one embodiment, when operating on a 256-bit wide vector, 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four separate 64-bit packed data elements (quadword (QW) size data elements), eight separate 32-bit packed data elements (doubleword (DW) size data elements), sixteen separate 16-bit packed data elements (word (W) size data elements) or thirty-two separate 8-bit data elements (byte (B) size data elements). However, in at least one embodiment, different vector widths and register sizes are possible.

[0296] In at least one embodiment, one or more execution units may be combined into fused execution units 1609A-1609N having thread control logic (1607A-1607N) for executing fused EUs. In at least one embodiment, multiple EUs may be merged into one EU group. In at least one embodiment, each EU in the fused EU group may be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group may vary according to various embodiments. In at least one embodiment, each EU may execute various SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32. In at least one embodiment, each fused graphics execution unit 1609A-1609N includes at least two execution units. In at least one embodiment, in at least one embodiment, the fused execution unit 1609A includes a first EU 1608A, a second EU 1608B, and a thread control logic 1607A shared by the first EU 1608A and the second EU 1608B. In at least one embodiment, thread control logic 1607A controls threads executing on fused graphics execution unit 1609A, allowing each EU within fused execution units 1609A-1609N to execute using a common instruction pointer register.

[0297] In at least one embodiment, one or more internal instruction caches (e.g., 1606) are included in the thread execution logic 1600D to cache thread instructions for the execution unit. In at least one embodiment, one or more data caches (e.g., 1612) are included to cache thread data during thread execution. In at least one embodiment, a sampler 1610 is included to provide texture sampling for 3D operations and media sampling for media operations. In at least one embodiment, the sampler 1610 includes a specialized texture or media sampling function to process the texture or media data during the sampling process before providing the sampled data to the execution unit.

[0298] During execution, in at least one embodiment, the graphics and media pipeline sends thread initiation requests to the thread execution logic 1600D through the thread generation and dispatch logic. In at least one embodiment, once a set of geometric objects have been processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 1602 is called to further calculate output information and cause the results to be written to the output surface (e.g., color buffer, depth buffer, template buffer, etc.). In at least one embodiment, the pixel shader or fragment shader calculates the values ​​of various vertex attributes to be interpolated on the rasterized object. In at least one embodiment, the pixel processor logic within the shader processor 1602 then executes the pixel or fragment shader program provided by the application program interface (API). In at least one embodiment, in order to execute the shader program, the shader processor 1602 dispatches the thread to the execution unit (e.g., 1608A) via the thread dispatcher 1604. In at least one embodiment, the shader processor 1602 uses the texture sampling logic in the sampler 1610 to access the texture data in the texture map stored in the memory. In at least one embodiment, arithmetic operations on texture data and input geometry data calculate pixel color data for each geometry fragment, or discard one or more pixels for further processing.

[0299] In at least one embodiment, data port 1614 provides a memory access mechanism for thread execution logic 1600D to output processed data to memory for further processing on the graphics processor output pipeline. In at least one embodiment, data port 1614 includes or is coupled to one or more cache memories (e.g., data cache 1612) to cache data for memory access via the data port.

[0300] like Fig.16EAs shown, in at least one embodiment, graphics execution unit 1608 may include instruction fetch unit 1637, general register file array (GRF) 1624, architectural register file array (ARF) 1626, thread arbiter 1622, issue unit 1630, branch unit 1632, a set of SIMD floating point units (FPUs) 1634, and in at least one embodiment, a set of dedicated integer SIMD ALUs 1635. In at least one embodiment, GRF 1624 and ARF 1626 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in graphics execution unit 1608. In at least one embodiment, per-thread architectural state is maintained in ARF 1626, while data used during thread execution is stored in GRF 1624. In at least one embodiment, the execution state of each thread, including the instruction pointer of each thread, may be saved in thread-specific registers in ARF 1626.

[0301] In at least one embodiment, graphics execution unit 1608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). In at least one embodiment, the architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where execution unit resources are logically allocated for executing multiple simultaneous threads.

[0302] In at least one embodiment, the graphics execution unit 1608 may issue multiple instructions together, each of which may be a different instruction. In at least one embodiment, the thread arbiter 1622 of the graphics execution unit thread 1608 may dispatch the instruction to one of the issue unit 1630, the branch unit 1632, or the SIMD FPU 1632 for execution. In at least one embodiment, each execution thread may access 128 general registers in the GRF 1624, each of which may store 32 bytes and may be accessed as a SIMD 8-element vector of 32-bit data elements. In at least one embodiment, each execution unit thread may access 4KB in the GRF 1624, although the embodiment is not limited thereto, and more or less register resources may be provided in other embodiments. In at least one embodiment, although the number of threads per execution unit may also vary depending on the embodiment, up to seven threads may be executed simultaneously. In at least one embodiment in which seven threads may access 4KB, the GRF 1624 may store a total of 28KB. In at least one embodiment, flexible addressing modes may allow registers to be addressed together to efficiently build wider registers or rectangular block data structures representing strides.

[0303] In at least one embodiment, memory operations, sampler operations, and other longer latency system communications are scheduled via "send" instructions executed by message passing send unit 1630. In at least one embodiment, dispatching branch instructions to a dedicated branch unit 1632 facilitates SIMD divergence and eventual convergence.

[0304] In at least one embodiment, the graphics execution unit 1608 includes one or more SIMD floating point units (FPUs) 1634 to perform floating point operations. In at least one embodiment, one or more FPUs 1634 also support integer calculations. In at least one embodiment, one or more FPUs 1634 can SIMD perform up to M 32-bit floating point (or integer) operations, or SIMD perform up to 2M 16-bit integer or 16-bit floating point operations. In at least one embodiment, at least one of the one or more FPUs provides extended math capabilities to support high throughput a priori math functions and double-precision 64-bit floating points. In at least one embodiment, there is also a set of 8-bit integer SIMD ALUs 1635, and can be specifically optimized to perform operations related to machine learning calculations.

[0305] In at least one embodiment, an array of multiple instances of graphics execution unit 1608 may be instantiated in graphics sub-kernel groupings (e.g., sub-slices). In at least one embodiment, execution unit 1608 may execute instructions across multiple execution lanes. In at least one embodiment, each thread executing on graphics execution unit 1608 executes on a different lane.

[0306] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Provides details about the reasoning and / or training logic 615. In at least one embodiment, some or all of the reasoning and / or training logic 615 can be incorporated into the execution logic 1600D. In addition, in at least one embodiment, other than Figure 6B and / or Figure 6C In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALUs of execution logic 1600D to execute one or more machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0307] Fig.17AA parallel processing unit ("PPU") 1700A is shown according to at least one embodiment. In at least one embodiment, the PPU 1700A is configured with machine-readable code that, if executed by the PPU 1700A, causes the PPU 1700A to perform some or all of the processes and techniques described throughout the present disclosure. In at least one embodiment, the PPU 1700A is a multithreaded processor implemented on one or more integrated circuit devices, and utilizes multithreading as a latency hiding technique designed to process computer-readable instructions (also referred to as machine-readable instructions or simply instructions) executed in parallel on multiple threads. In at least one embodiment, a thread refers to an execution thread and is an instance of a group of instructions configured to be executed by the PPU 1700A. In at least one embodiment, the PPU 1700A is a graphics processing unit ("GPU") configured to implement a graphics rendering pipeline for processing three-dimensional ("3D") graphics data to generate two-dimensional ("2D") image data for display on a display device (such as a liquid crystal display ("LCD") device). In at least one embodiment, PPU 1700A is used to perform computations, such as linear algebra operations and machine learning operations. Fig.17A The example parallel processor is shown for illustrative purposes only, and should be construed as a non-limiting example of a processor architecture contemplated within the scope of the present disclosure, and any suitable processor may be employed in addition and / or in place thereof.

[0308] In at least one embodiment, one or more PPUs 1700A are configured to accelerate high performance computing ("HPC"), data center, and machine learning applications. In at least one embodiment, PPU 1700A is configured to accelerate deep learning systems and applications, including the following non-limiting examples: autonomous vehicle platforms, deep learning, high-precision speech, image, text recognition systems, intelligent video analysis, molecular simulation, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulation, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, etc.

[0309] In at least one embodiment, the PPU 1700A includes, but is not limited to, an input / output ("I / O") unit 1706, a front end unit 1710, a scheduler unit 1712, a work distribution unit 1714, a hub 1716, a crossbar switch ("Xbar") 1720, one or more general processing clusters ("GPCs") 1718, and one or more partitioning units ("memory partitioning units") 1722. In at least one embodiment, the PPU 1700A is connected to a host processor or other PPUs 1700A via one or more high-speed GPU interconnects ("GPU interconnects") 1708. In at least one embodiment, the PPU 1700A is connected to a host processor or other peripheral devices via interconnect 1702. In one embodiment, the PPU 1700A is connected to a local memory including one or more memory devices ("memory") 1704. In at least one embodiment, the memory devices 1704 include, but are not limited to, one or more dynamic random access memory ("DRAM") devices. In at least one embodiment, one or more DRAM devices are configured and / or configurable as a high bandwidth memory ("HBM") subsystem with multiple DRAM dies stacked within each device.

[0310] In at least one embodiment, high-speed GPU interconnect 1708 may refer to a wire-based multi-channel communication link that a system uses to scale and includes one or more PPUs 1700A in conjunction with one or more central processing units ("CPUs"), supporting cache coherence between PPU 1700A and CPUs and CPU mastering. In at least one embodiment, high-speed GPU interconnect 1708 transmits data and / or commands to other units of PPU 1700A, such as one or more copy engines, video encoders, video decoders, power management units, and / or other units in PPU 1700A via hub 1716. Fig.17A Other components that may not be explicitly shown.

[0311] In at least one embodiment, I / O unit 1706 is configured to receive data from a host processor ( Fig.17A 1700A). In at least one embodiment, I / O unit 1706 communicates with a host processor directly through interconnect 1702 or through one or more intermediate devices (e.g., a memory bridge). In at least one embodiment, I / O unit 1706 can communicate with one or more other processors (e.g., one or more PPUs 1700A) via interconnect 1702. In at least one embodiment, I / O unit 1706 implements Peripheral Component Interconnect Express. Interface for In at least one embodiment, I / O unit 1706 implements an interface for communicating with external devices.

[0312] In at least one embodiment, I / O unit 1706 decodes packets received via interconnect 1702. In at least one embodiment, at least some of the packets represent commands configured to cause PPU 1700A to perform various operations. In at least one embodiment, I / O unit 1706 sends the decoded commands to various other units of PPU 1700A as specified by the commands. In at least one embodiment, the commands are sent to front end unit 1710 and / or to hub 1716 or other units of PPU 1700A, such as one or more replication engines, video encoders, video decoders, power management units, etc. ( Fig.17A In at least one embodiment, I / O unit 1706 is configured to route communications between various logical units of PPU 1700A.

[0313] In at least one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 1700A for processing. In at least one embodiment, the workload includes instructions and data to be processed by those instructions. In at least one embodiment, the buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and the PPU 1700A—the host interface unit can be configured to access the buffer in the system memory connected to the interconnect 1702 via memory requests transmitted via the I / O unit 1706 through the interconnect 1702. In at least one embodiment, the host processor writes the command stream to the buffer and then sends a pointer indicating the beginning of the command stream to the PPU 1700A, so that the front end unit 1710 receives pointers to one or more command streams and manages one or more command streams, reads commands from the command streams and forwards the commands to the various units of the PPU 1700A.

[0314] In at least one embodiment, the front end unit 1710 is coupled to a scheduler unit 1712 that configures the various GPCs 1718 to process tasks defined by one or more command streams. In at least one embodiment, the scheduler unit 1712 is configured to track state information related to the various tasks managed by the scheduler unit 1712, where the state information may indicate which GPC 1718 the task is assigned to, whether the task is active or inactive, a priority associated with the task, etc. In at least one embodiment, the scheduler unit 1712 manages multiple tasks that are executed on one or more GPCs 1718.

[0315] In at least one embodiment, the scheduler unit 1712 is coupled to a work distribution unit 1714, which is configured to dispatch tasks for execution on the GPCs 1718. In at least one embodiment, the work distribution unit 1714 tracks a plurality of scheduled tasks received from the scheduler unit 1712 and the work distribution unit 1714 manages a pending task pool and an active task pool for each GPC 1718. In at least one embodiment, the pending task pool includes a plurality of time slots (e.g., 16 time slots) containing tasks assigned to be processed by a particular GPC 1718; the active task pool may include a plurality of time slots (e.g., 4 time slots) for tasks actively processed by the GPC 1718, such that as one of the tasks in the GPC 1718 completes execution, the task is evicted from the active task pool of the GPC 1718 and one of the other tasks is selected from the pending task pool and scheduled for execution on the GPC 1718. In at least one embodiment, if the active task is idle on GPC 1718, such as while waiting for data dependencies to be resolved, the active task is evicted from GPC 1718 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 1718.

[0316] In at least one embodiment, work distribution unit 1714 communicates with one or more GPCs 1718 via XBar 1720. In at least one embodiment, XBar 1720 is an interconnect network that couples many units of PPU 1700A to other units of PPU 1700A and can be configured to couple work distribution unit 1714 to a specific GPC 1718. In at least one embodiment, one or more other units of PPU 1700A can also be connected to XBar 1716 via hub 1716.

[0317] In at least one embodiment, tasks are managed by a scheduler unit 1712 and assigned to one of the GPCs 1718 by a work distribution unit 1714. The GPC 1718 is configured to process tasks and produce results. In at least one embodiment, the results can be consumed by other tasks in the GPC 1718, routed to a different GPC 1718 via an XBar 1716, or stored in memory 1704. In at least one embodiment, the results can be written to the memory 1704 via a partition unit 1722, which implements a memory interface for writing data to or reading data from the memory 1704. In at least one embodiment, the results can be transmitted to another PPU 1704 or a CPU via a high-speed GPU interconnect 1708. In at least one embodiment, the PPU 1700A includes, but is not limited to, U partition units 1722, which are equal to the number of separate and distinct memory devices 1704 coupled to the PPU 1700A. In at least one embodiment, the following will be combined with Fig. 17C The partition unit 1722 is described in more detail.

[0318] In at least one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on the PPU 1700A. In one embodiment, multiple computing applications are executed simultaneously by the PPU 1700A, and the PPU 1700A provides isolation, quality of service (“QoS”), and independent address spaces for the multiple computing applications. In at least one embodiment, the application generates instructions (e.g., in the form of API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 1700A, and the driver kernel outputs the tasks to one or more streams processed by the PPU 1700A. In at least one embodiment, each task includes one or more related thread groups, which may be referred to as warps. In at least one embodiment, a warp includes multiple related threads (e.g., 32 threads) that can be executed in parallel. In at least one embodiment, a cooperative thread may refer to multiple threads, including instructions for executing tasks and exchanging data through shared memory, combined with Fig. 17C Threads and cooperating threads are described in greater detail according to at least one embodiment.

[0319] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6CDetails are provided regarding the inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the PPU 1700A. In at least one embodiment, the PPU 1700A is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or the PPU 1700A. In at least one embodiment, the PPU 1700A can be used to perform one or more of the neural network use cases described herein.

[0320] Fig. 17B A general processing cluster ("GPC") 1700B is shown in accordance with at least one embodiment. In at least one embodiment, GPC 1700B is Fig.17A 1700B. In at least one embodiment, each GPC 1700B includes, but is not limited to, a plurality of hardware units for processing tasks, and each GPC 1700B includes, but is not limited to, a pipeline manager 1702, a pre-raster operations unit ("PROP") 1704, a raster engine 1708, a work distribution crossbar switch ("WDX") 1716, a memory management unit ("MMU") 1718, one or more data processing clusters ("DPCs") 1706, and any suitable combination of components.

[0321] In at least one embodiment, the operation of the GPC 1700B is controlled by a pipeline manager 1702. In at least one embodiment, the pipeline manager 1702 manages the configuration of one or more DPCs 1706 to process tasks assigned to the GPC 1700B. In at least one embodiment, the pipeline manager 1702 configures at least one of the one or more DPCs 1706 to implement at least a portion of a graphics rendering pipeline. In at least one embodiment, the DPC 1706 is configured to execute vertex shader programs on a programmable streaming multiprocessor (“SM”) 1714. In at least one embodiment, the pipeline manager 1702 is configured to route packets received from the work distribution unit to appropriate logic units within the GPC 1700B, and in at least one embodiment, some packets may be routed to fixed function hardware units in the PROP 1704 and / or the raster engine 1708, while other packets may be routed to the DPC 1706 for processing by the primitive engine 1712 or the SM 1714. In at least one embodiment, pipeline manager 1702 configures at least one of DPCs 1706 to implement a neural network model and / or a computational pipeline.

[0322] In at least one embodiment, PROP unit 1704 is configured to route data generated by raster engine 1708 and DPC 1706 to a raster operations ("ROP") unit in partition unit 1722 in at least one embodiment, in conjunction with Fig.17A Described in more detail. In at least one embodiment, the PROP unit 1704 is configured to perform optimizations for color blending, organize pixel data, perform address translation, and the like. In at least one embodiment, the raster engine 1708 includes, but is not limited to, a plurality of fixed-function hardware units configured to perform various raster operations, and in at least one embodiment, the raster engine 1708 includes, but is not limited to, a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, a tile aggregation engine, and any suitable combination thereof. In at least one embodiment, the setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices; the plane equations are transmitted to the coarse raster engine to generate coverage information for the base primitives (e.g., an x, y coverage mask for the tile); the output of the coarse raster engine is transmitted to the culling engine, where fragments associated with primitives that fail the z test are culled, and to the clipping engine, where the fragments that are outside the frustum are clipped. In at least one embodiment, the clipped and culled fragments are passed to the fine raster engine to generate properties for the pixel fragments based on the plane equations generated by the setup engine. In at least one embodiment, the output of raster engine 1708 includes fragments to be processed by any appropriate entity (e.g., by a fragment shader implemented within DPC 1706).

[0323] In at least one embodiment, each DPC 1706 included in GPC 1700B includes, but is not limited to, an M-pipeline controller ("MPC") 1710; a primitive engine 1712; one or more SMs 1714; and any suitable combination thereof. In at least one embodiment, MPC 1710 controls the operation of DPC 1706, routing packets received from pipeline manager 1702 to appropriate units in DPC 1706. In at least one embodiment, packets associated with vertices are routed to primitive engine 1712, which is configured to fetch vertex attributes associated with the vertices from memory; conversely, packets associated with shader programs may be sent to SM 1714.

[0324] In at least one embodiment, SM 1714 includes, but is not limited to, a programmable stream processor configured to process tasks represented by multiple threads. In at least one embodiment, SM 1714 is multithreaded and configured to execute multiple threads (e.g., 32 threads) from a particular thread group simultaneously, and implements a single instruction, multiple data ("SIMD") architecture, wherein each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same instruction set. In at least one embodiment, all threads in a thread group execute the same instruction. In at least one embodiment, SM 1714 implements a single instruction, multiple thread ("SIMT") architecture, wherein each thread in a group of threads is configured to process different data sets based on the same instruction set, but wherein individual threads in a thread group are allowed to diverge during execution. In at least one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby achieving concurrency between warps and serial execution within the warp when threads in the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads within and between warps. In at least one embodiment, execution state is maintained for each individual thread, and threads of the same instruction can be converged and executed in parallel to improve efficiency. At least one embodiment of SM 1714 is described in more detail below.

[0325] In at least one embodiment, MMU 1718 is used between GPC 1700B and memory partition unit (e.g., Fig.17A The MMU 1718 provides an interface between the partition unit 1722 of the memory, and provides virtual address to physical address translation, memory protection, and arbitration of memory requests. In at least one embodiment, the MMU 1718 provides one or more translation lookaside buffers ("TLBs") for performing translation of virtual addresses to physical addresses in memory.

[0326] The reasoning and / or training logic 615 is used to perform reasoning and / or training operations associated with one or more embodiments. Figure 6B and / or Figure 6C Details are provided regarding inference and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or infer information provided to the GPC 1700B. In at least one embodiment, the GPC 1700B is used to infer or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or the GPC 1700B. In at least one embodiment, the GPC 1700B can be used to perform one or more of the neural network use cases described herein.

[0327] Fig. 17C A memory partition unit 1700C of a parallel processing unit ("PPU") according to at least one embodiment is shown. In at least one embodiment, the memory partition unit 1700C includes, but is not limited to, a raster operation ("ROP") unit 1702; a level 2 ("L2") cache 1704; a memory interface 1706; and any suitable combination thereof. In at least one embodiment, the memory interface 1706 is coupled to a memory. In at least one embodiment, the memory interface 1706 may implement a 32, 64, 128, 1024 bit data bus, or a similar implementation for high speed data transfer. In at least one embodiment, the PPU includes U memory interfaces 1706, one memory interface 1706 for each pair of partition units 1700C, wherein each pair of partition units 1700C is connected to a corresponding memory device. In at least one embodiment, in at least one embodiment, the PPU may be connected to up to Y memory devices, such as a high bandwidth memory stack or a graphics double data rate version 5 synchronous dynamic random access memory ("GDDR5 SDRAM").

[0328] In at least one embodiment, the memory interface 1706 implements a high bandwidth memory second generation ("HBM2") memory interface, and Y is equal to half of U. In at least one embodiment, the HBM2 memory stack is located on the same physical package as the PPU, providing significant power and area savings compared to a GDDR5 SDRAM system. In at least one embodiment, each HBM2 stack includes, but is not limited to, four memory dies, and Y=4, each HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits. In at least one embodiment, the memory supports single error correction double error detection ("SECDED") error correction code ("ECC") to protect data. In at least one embodiment, ECC provides higher reliability for computing applications that are sensitive to data corruption.

[0329] In at least one embodiment, the PPU implements a multi-level memory hierarchy. In at least one embodiment, the memory partition unit 1700C supports unified memory to provide a single unified virtual address space for the central processing unit ("CPU") and the PPU memory, thereby enabling data sharing between virtual memory systems. In at least one embodiment, the frequency of accesses by the PPU to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the page more frequently. In at least one embodiment, the high-speed GPU interconnect 1708 supports address translation services that allow the PPU to directly access the CPU's page tables and provide full access to the CPU memory through the PPU.

[0330] In at least one embodiment, the copy engine transfers data between multiple PPUs or between a PPU and a CPU. In at least one embodiment, the copy engine can generate a page fault for an address that is not mapped in a page table, and the memory partition unit 1700C then services the page fault, maps the address into a page table, and then the copy engine performs the transfer. In at least one embodiment, fixed (in at least one embodiment, non-pageable) memory is used for multiple copy engine operations between multiple processors, thereby substantially reducing the available memory. In at least one embodiment, in the event of a hardware page fault, the address can be passed to the copy engine without regard to whether it resides in a memory page, and the copy process is transparent.

[0331] According to at least one embodiment, from Fig.17A Data from memory 1704 or other system memory is retrieved by memory partition unit 1700C and stored in L2 cache 1704, which is located on the chip and shared between various GPCs. In at least one embodiment, each memory partition unit 1700C includes, but is not limited to, at least a portion of the L2 cache associated with the corresponding memory device. In at least one embodiment, lower-level caches are implemented in various units within the GPC. In at least one embodiment, each SM 1714 can implement a level 1 ("L1") cache, where the L1 cache is private memory dedicated to a particular SM 1714, and data is retrieved from the L2 cache 1704 and stored in each L1 cache for processing in the functional units of the SM 1714. In at least one embodiment, the L2 cache 1704 is coupled to the memory interface 1706 and the XBar 1720.

[0332] In at least one embodiment, ROP unit 1702 performs graphics raster operations related to pixel color, such as color compression, pixel blending, etc. In at least one embodiment, ROP unit 1702 performs depth testing in conjunction with raster engine 1708, receiving the depth of a sample position associated with a pixel fragment from a culling engine of raster engine 1708. In at least one embodiment, the depth is tested against a corresponding depth in a depth buffer of the sample position associated with the fragment. In at least one embodiment, if the fragment passes the depth test for the sample position, ROP unit 1702 updates the depth buffer and sends the result of the depth test to raster engine 1708. It will be appreciated that the number of partition units 1700C can be different than the number of GPCs, and therefore, each ROP unit 1702 can be coupled to each GPC in at least one embodiment. In at least one embodiment, ROP unit 1702 tracks packets received from different GPCs and determines to which result the result generated by ROP unit 1702 is routed via XBar 1720.

[0333] Fig.17D Streaming multiprocessor ("SM") 1700D is shown in accordance with at least one embodiment. In at least one embodiment, SM 1700D is Fig. 17BSMs. In at least one embodiment, SM 1700D includes, but is not limited to, instruction cache 1702; one or more scheduler units 1704; register file 1708; one or more processing cores ("cores") 1710; one or more special function units ("SFUs") 1712; one or more load / store units ("LSUs") 1714; interconnect network 1716; shared memory / level 1 ("L1") cache 1718; and any suitable combination thereof. In at least one embodiment, a work distribution unit schedules tasks for execution on a general processing cluster ("GPC") of a parallel processing unit ("PPU"), and each task is assigned to a specific data processing cluster ("DPC") within the GPC, and if the task is associated with a shader program, the task is assigned to one of SMs 1700D. In at least one embodiment, scheduler unit 1704 receives tasks from the work distribution unit and manages instruction scheduling for one or more thread blocks assigned to SMs 1700D. In at least one embodiment, the scheduler unit 1704 schedules thread blocks to execute as warps of parallel threads, where each thread block is assigned at least one warp. In at least one embodiment, each warp executes a thread. In at least one embodiment, the scheduler unit 1704 manages a plurality of different thread blocks, assigns warps to different thread blocks, and then dispatches instructions from a plurality of different cooperative groups to various functional units (e.g., processing cores 1710, SFUs 1712, and LSUs 1714) in each clock cycle.

[0334] In at least one embodiment, a cooperative group may refer to a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, thereby enabling the expression of richer and more efficient parallel decompositions. In at least one embodiment, a cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. In at least one embodiment, the application of the programming model provides a single, simple construct for synchronizing cooperative threads: a barrier across all threads of a thread block (e.g., a syncthreads() function). However, in at least one embodiment, programmers can define thread groups at a granularity smaller than a thread block and synchronize within the defined group to achieve higher performance, design flexibility, and software reuse in the form of a collective group-wide functional interface. In at least one embodiment, cooperative groups enable programmers to explicitly define thread groups at sub-block (in at least one embodiment, as small as a single thread) and multi-block granularity, and perform collective operations, such as synchronizing threads in a cooperative group. In at least one embodiment, the programming model supports clean composition across software boundaries, so that libraries and utility functions can be safely synchronized in their local environment without having to make assumptions about convergence. In at least one embodiment, the cooperation group primitives enable new patterns of cooperative parallelism, including but not limited to producer-consumer parallelism, opportunistic parallelism, and global synchronization over an entire grid of thread blocks.

[0335] In at least one embodiment, the scheduling unit 1706 is configured to send instructions to one or more of the functional units, and the scheduler unit 1704 includes, but is not limited to, two scheduling units 1706 that enable two different instructions from the same thread warp to be scheduled per clock cycle. In at least one embodiment, each scheduler unit 1704 includes a single scheduling unit 1706 or additional scheduling units 1706.

[0336] In at least one embodiment, each SM 1700D includes, in at least one embodiment, but is not limited to, a register file 1708 that provides a set of registers for the functional units of the SM 1700D. In at least one embodiment, the register file 1708 is divided between each functional unit, so that a dedicated portion of the register file 1708 is allocated to each functional unit. In at least one embodiment, the register file 1708 is divided between different thread warps executed by the SM 1700D, and the register file 1708 provides temporary storage for operands of the data paths connected to the functional units. In at least one embodiment, each SM 1700D includes, in at least one embodiment, but is not limited to, a plurality of L processing cores 1710. In at least one embodiment, the SM 1700D includes, but is not limited to, a large number (e.g., 128 or more) of different processing cores 1710. In at least one embodiment, each processing core 1710 includes, but is not limited to, a full pipeline, single-precision, double-precision, and / or mixed-precision processing unit, which includes, but is not limited to, a floating-point arithmetic logic unit and an integer arithmetic logic unit. In at least one embodiment, the floating point arithmetic logic unit implements the IEEE 754-2008 standard for floating point arithmetic. In at least one embodiment, the processing cores 1710 include, but are not limited to, 64 single precision (32-bit) floating point cores, 64 integer cores, 32 double precision (64-bit) floating point cores, and 8 tensor cores.

[0337] According to at least one embodiment, the tensor core is configured to perform matrix operations. In at least one embodiment, one or more tensor cores are included in the processing core 1710. In at least one embodiment, the tensor core is configured to perform deep learning matrix arithmetic, such as convolution operations for neural network training and inference. In at least one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.

[0338] In at least one embodiment, the matrix multiplication inputs A and B are 16-bit floating point matrices, and the accumulation matrices C and D are 16-bit floating point or 32-bit floating point matrices. In at least one embodiment, the tensor core performs a 32-bit floating point accumulation operation on the 16-bit floating point input data. In at least one embodiment, the 16-bit floating point multiplication uses 64 operations and obtains a full-precision product, which is then accumulated with other intermediate products using 32-bit floating point addition to perform a 4x4x4 matrix multiplication. In at least one embodiment, the tensor core is used to perform larger two-dimensional or higher dimensional matrix operations composed of these smaller elements. In at least one embodiment, an API (such as a CUDA 9C++ API) exposes specialized matrix loads, matrix multiplications and accumulations, and matrix storage operations to efficiently use tensor cores from CUDA-C++ programs. In at least one embodiment, at the CUDA level, the warp level interface assumes a 16×16 size matrix across all 32 warp threads.

[0339] In at least one embodiment, each SM 1700D includes, but is not limited to, M SFUs 1712 that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In at least one embodiment, the SFUs 1712 include, but are not limited to, tree traversal units configured to traverse a hierarchical tree data structure. In at least one embodiment, the SFUs 1712 include, but are not limited to, texture units configured to perform texture map filtering operations. In at least one embodiment, the texture units are configured to load texture maps (e.g., 2D arrays of texture pixels) from memory and sample the texture maps to generate sampled texture values ​​for use by shader programs executed by the SM 1700D. In at least one embodiment, the texture maps are stored in a shared memory / L1 cache 1718. In at least one embodiment, according to at least one embodiment, the texture units use mip-maps (e.g., texture maps with different levels of detail) to implement texture operations (such as filtering operations). In at least one embodiment, each SM 1700D includes, but is not limited to, two texture units.

[0340] In at least one embodiment, each SM 1700D includes, but is not limited to, N LSUs 1714 that implement load and store operations between shared memory / L1 cache 1718 and register file 1708. In at least one embodiment, an interconnect network 1716 connects each functional unit to register file 1708, and the LSUs 1714 connect to register file 1708 and shared memory / L1 cache 1718. In at least one embodiment, the interconnect network 1716 is a crossbar switch that can be configured to connect any functional unit to any register in register file 1708, and to connect the LSUs 1714 to memory locations in register file 1708 and shared memory / L1 cache 1718.

[0341] In at least one embodiment, shared memory / L1 cache 1718 is an array of on-chip memory that allows data storage and communication between SM 1700D and primitive engines and between threads in SM 1700D in at least one embodiment. In at least one embodiment, shared memory / L1 cache 1718 includes, but is not limited to, 128KB of storage capacity and is located in the path from SM 1700D to the partition unit. In at least one embodiment, shared memory / L1 cache 1718 is used in at least one embodiment for caching reads and writes. In at least one embodiment, one or more of shared memory / L1 cache 1718, L2 cache, and memory is a backing store.

[0342] In at least one embodiment, data cache and shared memory functionality are combined into a single memory block, providing improved performance for both types of memory accesses. In at least one embodiment, the capacity is used by programs that do not use the shared memory or use it as a cache, for example if the shared memory is configured to use half of the capacity, and texture and load / store operations can use the remaining capacity. According to at least one embodiment, the integration within the shared memory / L1 cache 1718 enables the shared memory / L1 cache 1718 to be used as a high throughput pipeline for streaming data while providing high bandwidth and low latency access to frequently reused data. In at least one embodiment, when configured for general parallel computing, a simpler configuration can be used compared to graphics processing. In at least one embodiment, the fixed-function graphics processing unit is bypassed, thereby creating a simpler programming model. In at least one embodiment, in a general parallel computing configuration, the work distribution unit directly allocates and distributes blocks of threads to DPCs. In at least one embodiment, the threads in the block execute a common program, use unique thread IDs in computations to ensure that each thread generates unique results, use SM 1700D to execute the program and perform computations, use shared memory / L1 cache 1718 to communicate between threads, and use LSU 1714 to read and write global memory through shared memory / L1 cache 1718 and memory partitioning units. In at least one embodiment, when configured for general parallel computation, SM 1700D writes commands to scheduler unit 1704 that can be used to start new work on the DPC.

[0343] In at least one embodiment, the PPU is included in or coupled to a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head mounted display, a handheld electronic device, etc. In at least one embodiment, the PPU is implemented on a single semiconductor substrate. In at least one embodiment, the PPU is included in a system on a chip ("SoC") along with one or more other devices (e.g., additional PPUs, memory, a reduced instruction set computer ("RISC") CPU, one or more memory management units ("MMU"), a digital-to-analog converter ("DAC"), etc.).

[0344] In at least one embodiment, the PPU may be included on a graphics card that includes one or more storage devices. The graphics card may be configured to communicate with a computer on a desktop computer motherboard. In at least one embodiment, the PPU may be an integrated graphics processing unit ("iGPU") included in a chipset of a motherboard.

[0345] Reasoning and / or training logic 615 is used to perform reasoning and / or training operations related to one or more embodiments. Figure 6B and / or Figure 6C Provides details about reasoning and / or training logic 615. In at least one embodiment, the deep learning application processor is used to train a machine learning model (such as a neural network) to predict or reason about information provided to the SM 1700D. In at least one embodiment, the SM 1700D is used to reason or predict information based on a machine learning model (e.g., a neural network) that has been trained by another processor or system or by the SM 1700D. In at least one embodiment, the SM 1700D can be used to perform one or more of the neural network use cases herein.

[0346] In at least one embodiment, a single semiconductor platform may refer to a unique single semiconductor-based integrated circuit or chip. In at least one embodiment, a multi-chip module with increased connectivity may be used that emulates on-chip operations and provides substantial improvements over implementations utilizing a central processing unit ("CPU") and bus. In at least one embodiment, the various modules may also be placed separately or in various combinations of semiconductor platforms, depending on the needs of the user.

[0347] In at least one embodiment, a computer program in the form of a machine-readable executable code or a computer control logic algorithm is stored in the main memory 4ee04 and / or the auxiliary storage. According to at least one embodiment, if executed by one or more processors, the computer program enables the system 4ee00 to perform various functions. In at least one embodiment, the memory 4ee04, storage and / or any other storage are possible examples of computer-readable media. In at least one embodiment, the auxiliary storage can refer to any suitable storage device or system, such as a hard disk drive and / or a removable storage drive, which represents a floppy disk drive, a tape drive, an optical disk drive, a digital versatile disk ("DVD") drive, a recording device, a universal serial bus ("USB") flash memory, etc. In at least one embodiment, the architecture and / or functions of each of the previous figures are implemented in the environment of CPU 4ee02; parallel processing system 4ee12; an integrated circuit capable of having at least part of the capabilities of two CPUs 4ee02; parallel processing system 4ee12; a chipset (e.g., a group of integrated circuits designed to work and sell as a unit that performs related functions, etc.); and any appropriate combination of integrated circuits.

[0348] In at least one embodiment, the architecture and / or functionality of the various previous figures are implemented in the context of a general purpose computer system, a circuit board system, a game console system dedicated to entertainment purposes, a dedicated system, etc. In at least one embodiment, the computer system 4ee00 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., wireless, handheld device), a personal digital assistant ("PDA"), a digital camera, a vehicle, a head mounted display, a handheld electronic device, a mobile telephone device, a television, a workstation, a game console, an embedded system, and / or any other type of logic.

[0349] In at least one embodiment, the parallel processing system 4ee12 includes, but is not limited to, a plurality of parallel processing units ("PPUs") 4ee14 and associated memory 4ee16. In at least one embodiment, the PPUs 4ee14 are connected to a host processor or other peripheral device via an interconnect 4ee18 and a switch 4ee20 or multiplexer. In at least one embodiment, the parallel processing system 4ee12 distributes computational tasks across parallelizable PPUs 4ee14, for example, as part of a distribution of computational tasks across multiple graphics processing unit ("GPU") thread blocks. In at least one embodiment, memory (e.g., for read and / or write access) is shared and accessed between some or all of the PPUs 4ee14, although such shared memory may incur a performance penalty relative to the use of local memory and registers resident on the PPUs 4ee4ee. In at least one embodiment, the operation of the PPUs 4ee14 is synchronized by using a command (such as __syncthreads()), where all threads in a block (e.g., executing across multiple PPUs 4ee14) reach a certain code execution point before proceeding.

[0350] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of the present disclosure as defined by the appended claims.

[0351] Unless otherwise noted or clearly contradictory to the context, in the context of describing the disclosed embodiments (particularly in the context of the appended claims), the use of the terms "one" and "an" and "the" and similar references should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include", "have", "include" and "contain" should be interpreted as open terms (meaning "including but not limited to") unless otherwise noted. The term "connected" (which refers to a physical connection when unmodified) should be interpreted as partially or completely included, attached to or connected together, even if there are some interventions. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, respectively, and each individual value is incorporated into the specification as if it were individually described herein. Unless otherwise noted or contradictory to the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set including one or more members. Furthermore, unless otherwise indicated or contradicted by context, a "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equal.

[0352] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B and C" are understood in context to be generally used to indicate an item, clause, or the like that may be A or B or C, or any non-empty subset of the set A and B and C. For example, in the illustrative example of a set having three members, the conjunction phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunction language may not be intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless expressly indicated otherwise or contradicted by context, "plurality" indicates a plural state (e.g., "plurality of items" means a plurality of items). The plural is at least two items, but may be more if expressly indicated or indicated by context. Further, unless stated otherwise or clear from the context, “based on” means “based at least in part on” rather than “based solely on”.

[0353] Unless otherwise indicated herein or clearly contradictory to the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are jointly executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of, for example, a computer program that includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (in at least one embodiment, as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all the code, but the plurality of non-transitory computer-readable storage media stores all the code together. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.

[0354] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. In addition, a computer system that implements at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices that operate in different ways, such that the distributed computer system performs the operations herein, and such that a single device does not perform all operations.

[0355] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential to practicing the disclosure.

[0356] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.

[0357] In the specification and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. On the contrary, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0358] Unless expressly stated otherwise, it is understood that throughout the specification, “processing references,” “computing,” “calculating,” “determining,” and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or transforms data represented as physical quantities (e.g., electronic) in registers and / or memories of the computing system into other data similarly represented as physical quantities in the computing system's memories, registers, or other such information storage, transmission, or display devices.

[0359] In a similar manner, a "processor" may refer to any device or part of a memory that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.

[0360] In this document, reference may be made to obtaining, acquiring, receiving or inputting analog or digital data into a subsystem, a computer system or a computer-implemented machine. Acquisition, acquisition, reception or inputting analog and digital data may be accomplished in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving or inputting analog or digital data may be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving or inputting analog or digital data may be accomplished by transmitting data from a providing entity to an acquisition entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending or presenting analog or digital data may be accomplished by transmitting data as an input or output parameter of a function call, an application programming interface or an interprocess communication mechanism.

[0361] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. In addition, although specific responsibilities are defined above for discussion purposes, various functions and responsibilities may be allocated and divided in different ways, depending on the circumstances.

[0362] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.

Claims

1. A multi-format docking plate, comprising: a first switch for enabling communication between a first format graphics processing unit (GPU) board and a second format GPU board; a second switch, configured to implement communication between a central processing unit (CPU) and the first format GPU board and between the CPU and the second format GPU board; a ribbon cable disposed between a bridge connector of the first format GPU board and a second bridge connector of the first switch; as well as A board line is located between the inlay connector of the first switch and the second inlay connector of the second format GPU board.

2. The multi-format docking plate of claim 1, further comprising: a bridge connector for coupling between the first format GPU board and the first switch; as well as An inlay connector is used to couple between the first switch and the second format GPU board.

3. The multi-format docking plate of claim 1, further comprising: an inlay connector for coupling between the first format GPU board and the second switch; a second inlay connector for coupling between the second format GPU board and the second switch; as well as The third inlay connector is used for coupling between the second switch and the CPU.

4. The multi-format docking plate of claim 1, further comprising: One or more clock circuits are used to provide a first clock frequency for the first format GPU board and a second clock frequency for the second format GPU board.

5. The multi-format docking plate of claim 1, further comprising: a first communication channel operating at a first bandwidth for a first communication between the first format GPU board and the first switch; as well as A second communication channel, operating at a second bandwidth, is used for second communication between the second format GPU board and the first switch.

6. The multi-format docking plate of claim 5, further comprising: A common communication channel, operating at a common clock frequency, is used for the first communication and the second communication.

7. The multi-format docking plate of claim 1, further comprising: at least one processor associated with the first switch; as well as A memory having instructions for execution on at least one processor to cause the at least one processor to allocate at least communication resources for the first format GPU board and the second format GPU board.

8. The multi-format docking plate of claim 1, further comprising: A bridge connector between the first format GPU board and the first switch is used to provide at least a clock signal referenced from the second format GPU board.

9. The multi-format docking plate of claim 1, further comprising: Rapid peripheral component interconnection between the first format GPU board and the multi-format docking board connectors; and NVLINK between the first format GPU board and the first switch TM format bridge connector.

10. The multi-format docking plate of claim 1, further comprising: SXM TM Connector that supports NVLINK between the second format GPU board and the multi-format docking board TM interconnection standard and supports the second format GPU board and the multi-format docking board Communication standards.

11. The multi-format docking plate of claim 1 , further comprising: NVLINK forming the first switch TM Interconnect standard switches; as well as The second switch is formed switch.

12. A docking plate, comprising: an inlay connector for enabling communication between a first format graphics processing unit (GPU) board and a central processing unit (CPU) on the docking board and between a second format GPU board and the CPU on the docking board, wherein the inlay connector is connected to another inlay connector of another switch through a board line and enables third communication between the first format GPU board and the second format GPU board; and A bridge connector is connected to another bridge connector of the first format GPU board via a ribbon cable.

13. The docking plate of claim 12, further comprising: The clock receiver circuit is used to receive a clock signal from a clock generator of the docking board.

14. The docking plate of claim 12, further comprising: at least one processor; as well as A memory having instructions for execution on at least one processor to cause the at least one processor to: receive identifiers for the first format GPU board and the second format GPU board, allocate resources for the first format GPU board and the second format GPU board, and transmit the identifiers to the CPU.

15. The docking plate of claim 12, further comprising: The inlay connector is adapted to be coupled with a second inlay connector of the docking board to achieve communication between the first format GPU board and the CPU and between the second format GPU board and the CPU through board wires in the docking board.

16. The docking plate of claim 12, further comprising: The inlay connector forms a peripheral component interconnect for rapid Connector.

17. The docking plate of claim 12, further comprising: at least one processor; as well as A memory having instructions for execution on at least one processor to cause the at least one processor to communicate with an operating system associated with the CPU at least through a GPU driver associated with the CPU.

18. A method for implementing a multi-format docking board, the method comprising: a bridge connector for enabling communication between a first format graphics processing unit (GPU) board and a switch; as well as one or more inlay connectors for enabling communication between the switch and a second format GPU board to enable communication between the CPU and the first format GPU board, wherein at least one of the one or more inlay connectors is connected to the first or second format GPU board via a board line; and The bridge connector complies with NVLINK TM An interconnection standard is used to support communication between the first format GPU board, the switch and the second format GPU board.

19. The method for implementing a multi-format docking board as claimed in claim 18, further comprising: at least one processor; as well as A memory having instructions for execution on at least one processor to cause the at least one processor to allocate at least communication resources for the first format GPU board and the second format GPU board to use NVLINK TM Interconnect standards communicate with each other.

20. A method for operating a multi-format docking plate, comprising: implementing a first switch for communication between a first format graphics processing unit (GPU) board and a second format GPU board; implementing a second switch for communication between a central processing unit (CPU) and the first format GPU board and between the CPU and the second format GPU board; a ribbon cable enabling communication between the bridge connector of the first format GPU board and the second bridge connector of the first switch; as well as A board line is provided to enable communication between the inlay connector of the first switch and the second inlay connector of the second format GPU board.

21. The method of claim 20, further comprising: coupling the first format GPU board to the first switch using at least a bridge connector; as well as The first switch is coupled to the second format GPU board using at least an inlay connector.

22. The method of claim 20, further comprising: providing an inlay connector for coupling between the first format GPU board and the second switch; providing a second inlay connector for coupling between the second format GPU board and the second switch; and A third inlay connector is provided for coupling between the second switch and the CPU.

23. The method of claim 20, further comprising: One or more clock circuits are used to provide a first clock frequency for the first format GPU board and a second clock frequency for the second format GPU board.

24. The method of claim 20, further comprising: operating within a first communication channel at a first bandwidth to implement a first communication between the first format GPU board and the first switch; as well as A second communication between the second format GPU board and the first switch is implemented within a second communication channel operating at a second bandwidth.

25. The method of claim 24, further comprising: The first communication and the second communication are accomplished using a common communication channel operating at a common clock frequency.

26. The method of claim 20, further comprising: At least communication resources are allocated to the first format GPU board and the second format GPU board by at least a processor associated with the first switch.

27. The method of claim 20, further comprising: At least a clock signal referenced from the second format GPU board is provided using a bridge connector between the first format GPU board and the first switch.

Citation Information

Patent Citations

  • Daughter card approach to employing multiple graphics cards within a system

    US20050270298A1

  • Electronic device and graphics processing unit card

    WO2016061794A1