Switchable hybrid graphics

By using the mixed states of MUX 32, 34, 36, and 38 and the logic control of information provider 54, analyzer 24, and trigger 26, the latency and error problems caused by memory copying during graphics processor switching are solved, achieving seamless switching between integrated and discrete graphics processors, improving user experience and battery efficiency.

CN109584141BActive Publication Date: 2025-12-30INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201811134140.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-29
Filing Date
2018-09-27
Publication Date
2025-12-30
Estimated Expiration
2038-09-27

AI Technical Summary

Technical Problem

In existing technologies, there are delays and errors caused by memory copying operations during the switching process of graphics processors, especially when switching between integrated graphics processors and discrete graphics processors, which may lead to blue screens and increased battery consumption.

Method used

The mixed states or modes of MUX 32, 34, 36, and 38 allow selective switching between integrated and discrete graphics processors. The logic control of information provider 54, analyzer 24, and trigger 26 ensures stable communication between the display device and the graphics processor, reduces the total motion-to-photon latency, and avoids memory errors.

Benefits of technology

It enables seamless switching between integrated and discrete graphics processors, reduces total motion-to-photon latency, reduces jitter, improves user experience, and optimizes battery consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109584141B_ABST
    Figure CN109584141B_ABST
Patent Text Reader

Abstract

Systems, devices, and methods can provide a technique that forms a determination of whether to connect a discrete graphics processor or an integrated graphics processor to a connected display device based on information from the connected display device that corresponds to whether the connected display device is to be driven by an integrated graphics processor or a discrete graphics processor.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] The embodiments generally relate to graphics processors, such as, for example, integrated graphics processors and / or discrete graphics processors. Different display devices may require different graphics processors as needed. For example, head-mounted display (HMD) systems can be used in virtual reality (VR) and augmented reality (AR) applications to present visual content to a wearer in various settings, such as immersive gaming and / or entertainment. A typical HMD may include a display that visually presents images. Image data can be processed to present content to the display. More specifically, game applications may use hardware-accelerated graphics application programming interfaces (APIs) to utilize the capabilities of discrete graphics processors, where such utilization may include offloading graphics and non-graphics computations to the discrete graphics processor in order to maintain an interactive frame rate. Attached Figure Description

[0002] Various advantages of the embodiments will become apparent to those skilled in the art from reading the following specification and appended claims and by referring to the following drawings, wherein:

[0003] Figure 1 This is a diagram illustrating an example of a computing architecture according to an embodiment;

[0004] Figure 2-3 This is a flowchart illustrating an example of a method for operating a computing architecture according to an embodiment;

[0005] Figure 4 yes Figure 3 Continuing from the flowchart;

[0006] Figure 5-10 An additional flowchart is an example of a method for operating a computing architecture according to an embodiment;

[0007] Figure 11 This is a block diagram of an example of semiconductor packaging;

[0008] Figure 12 This is a block diagram illustrating an example overview of a processing system according to an embodiment;

[0009] Figure 13 This is an example block diagram providing an overview of a processor according to an embodiment;

[0010] Figure 14 This is an example block diagram providing an overview of a graphics processor according to an embodiment;

[0011] Figure 15 This is a block diagram of an example of a graphics processing engine according to an embodiment;

[0012] Figure 16 This is a block diagram of an example graphics processor core according to an embodiment;

[0013] Figures 17A-17B The execution logic of the embodiment is shown;

[0014] Figure 18 This is a block diagram illustrating an example of a graphics processor instruction format according to an embodiment;

[0015] Figure 19 This is a block diagram of an example of a graphics processor according to an embodiment;

[0016] Figures 20A-20B This is a block diagram illustrating an example of graphics processor programming according to an embodiment;

[0017] Figure 21 This is a block diagram illustrating an example of a graphical software architecture according to an embodiment;

[0018] Figure 22A This is a block diagram of an example intellectual property (IP) core development system according to an embodiment;

[0019] Figure 22B This is a block diagram illustrating an example of an integrated circuit package according to an embodiment; and

[0020] Figure 23-25B This is a block diagram of an example of an integrated circuit and a related graphics processor according to an embodiment. Detailed Implementation

[0021] Figure 1 A computing architecture 52 is illustrated. The illustrated computing architecture 52 includes an integrated graphics processor 28 and a discrete graphics processor 30, which may also be referred to as a dedicated graphics card or dedicated graphics processor. The integrated graphics processor 28 and the discrete graphics processor 30 may be part of a computing system or computing device, such as a server, desktop computer, laptop computer, tablet computer, convertible tablet computer, smartphone, personal digital assistant (PDA), mobile internet device (MID), wearable device, media player, etc., or any combination thereof.

[0022] The computing architecture 52 may include MUX 32, 34, 36, and 38. Each of the MUX 32, 34, 36, and 38 may have signal lines (e.g., inputs) connected to the integrated graphics processor 28 and the discrete graphics processor 30. The MUX 32, 34, 36, and 38 may be controlled to output signals from either the integrated graphics processor 28 or the discrete graphics processor 30. The illustrated MUX 32, 34, 36, and 38 are connected to display devices 42, 44, and 46 via display interfaces 48, 50, and 52, such that the outputs of the MUX 32, 34, 36, and 38 are provided to the display devices 42, 44, and 46 via the display interfaces 48, 50, and 52. Depending on the selection of the MUX 32, 34, 36, and 38, this connection may allow bidirectional communication between the integrated graphics processor 28 and the display devices 42, 44, and 46, and between the discrete graphics processor 30 and the display devices 42, 44, and 46. Therefore, MUX 32, 34, 36, 38 can selectively electrically connect the integrated graphics processor 28 and the discrete graphics processor 30 to the display devices 42, 44, 46 so that information can be transferred between the integrated graphics processor 28 and the display devices 42, 44, 46, and between the discrete graphics processor 30 and the display devices 42, 44, 46.

[0023] For example, if discrete graphics processor 30 is electrically connected to the output of MUX 36, discrete graphics processor 30 can receive information from and provide information to display device 44, while integrated graphics processor 28 is electrically disconnected from display device 44. Host controller 40 (e.g., chipset) can also control connections to, for example, a Universal Serial Bus (USB) Type-C connector. Host controller 40 can be connected to display interface 48, which is connected to the illustrated display device 42. Display devices 42, 44, and 46 can be different display devices connected to a computing system or a part of a computing system. For example, display device 42 can be an HMD, display device 44 can be a high-definition main display, and display device 46 can be an internal monitor of the computing system (e.g., a laptop computer monitor).

[0024] The illustrated computing architecture 52 includes an information provider 54 (e.g., logic instructions, configurable logic, fixed-function hardware logic, etc., or any combination thereof), an analyzer 24 (e.g., logic instructions, configurable logic, fixed-function hardware logic, etc., or any combination thereof), and a trigger 26 (e.g., logic instructions, configurable logic, fixed-function hardware logic, etc., or any combination thereof), which can be collectively referred to as "logic". The information provider 54, analyzer 24, and trigger 26 can determine whether each of the display devices 42, 44, 46 is driven by the integrated graphics processor 28 or the discrete graphics processor 30, and individually control the MUX 32, 34, 36, 38 to output a corresponding one of the outputs of the integrated graphics processor 28 and the discrete graphics processor 30, based on the determination. This reduces the overall motion-to-photon (M2P) latency and reduces memory errors that can cause blue screens to be displayed by the display devices 42, 44, 46. Trigger 26 and / or information provider 54 can suppress information related to the modification of MUX 32, 34, 36, 38 (e.g., ASL information) so that MUX 32, 34, 36, 38 are not modified by other components.

[0025] For example, during the boot sequence of the computing system (or when the display device 42 is connected to the computing system), and based on information from the display device 42, the information provider 54 (e.g., logic instructions, configurable logic, fixed-function hardware logic, etc., or any combination thereof), analyzer 24, and trigger 26 can operate together (described below) to determine and selectively control the MUX 32, 34 to electrically connect the integrated graphics processor 28 or the discrete graphics processor 30 to the display device 42. The information may include, for example, the Extended Display Identification Data (EDID) of the display device 42. Similarly, the information provider 54, analyzer 24, and trigger 26 can control the MUX 36, 38 to electrically connect the integrated graphics processor 28 or the discrete graphics processor 30 to the display devices 44, 46.

[0026] Conversely, M2P is larger when discrete graphics processors provide information (e.g., frames) to an integrated graphics processor, which in turn provides information to the display device. Reducing the total M2P latency to below 20 milliseconds can enhance the user experience by reducing jitter and providing an immersive experience. However, such memory copying operations can increase M2P latency to unacceptable levels.

[0027] Furthermore, switching between the integrated and discrete graphics processors (GPUs) after an application begins operating with the display device can lead to errors. For example, the integrated GPU may use different memory than the discrete GPU. However, the application might continue writing to the integrated GPU's memory after such a switch, causing a memory error and potentially displaying a blue screen. Therefore, errors can occur when determining whether an application's content (e.g., whether the application can utilize a large number of graphics uses) can be provided to the display device using either the discrete GPU's output or the integrated GPU's output, as such processes can involve memory copying and switching.

[0028] Some digital access control (DAC) media can assign the use of a single controller (e.g., integrated graphics controller 28) to all display devices for seamless playback. Therefore, DAC media may not operate seamlessly if different display interfaces are permanently hardwired to different ones of integrated graphics processor 28 and discrete graphics processor 30. Instead, the hybrid switching described above allows each of display devices 42, 44, and 46 to be mapped to the integrated graphics driver of integrated graphics processor 28, and also reduces battery consumption by enabling switching between discrete graphics processor 30 and integrated graphics processor 28.

[0029] Discrete graphics processor 30 can have a higher performance graphics processor than integrated graphics processor 28. For example, discrete graphics processor 30 can have dedicated random access memory (RAM) and may not require the use of the central processing unit's RAM. Furthermore, discrete graphics processor 30 may include a dedicated cooling system and has higher parallel processing capabilities than integrated graphics processor 28. Conversely, integrated graphics processor 28 may share resources (e.g., RAM) with the central processing unit and has less parallel processing capability than discrete graphics processor 30. However, integrated graphics processor 28 can use less power than discrete graphics processor 30. Integrated graphics processor 28 can be integrated into the motherboard or central processing unit. Conversely, discrete graphics processor 30 can be connected to the motherboard, but may not be integrated with the motherboard and can be separate from the motherboard. Therefore, integrated graphics processor 28 and discrete graphics processor 30 can be used in different situations.

[0030] MUX 32, 34, 36, and 38 can be configured into a hybrid state or mode. The hybrid state allows for individual switching of each of MUX 32, 34, 36, and 38, enabling each of them to switch between outputting the integrated graphics processor 28 and the discrete graphics processor 30. Therefore, MUX 32 can output the first output of the discrete graphics processor 30, MUX 34 can output the second output of the discrete graphics processor 30, MUX 36 can output the first output of the integrated graphics processor 28, and MUX 38 can output the second output of the integrated graphics processor 28. As mentioned, MUX 32, 34, 36, and 38 allow bidirectional communication. In the hybrid state, unless otherwise changed, MUX 32, 34, 36, and 38 can by default connect the integrated graphics processor 30 to display devices 42, 44, and 46. For example, MUX 32, 34, 36, 38 can electrically connect the integrated graphics processor 30 to each of the display devices 42, 44, 46.

[0031] However, the hybrid state can be overridden by the user. For example, the user can override the hybrid state to control at least one of the MUX 32, 34, 36, and 38 to always electrically connect the discrete graphics processor 30 or the integrated graphics processor 28 to the corresponding display device 42, 44, or 46. For example, the user can override the hybrid state through BIOS boot options so that the MUX 32 and 34 always output the output of the discrete graphics processor 30.

[0032] As already mentioned, the information provider 54, analyzer 24, and flip-flop 26 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), flash memory, etc., implemented as configurable logic such as, for example, a programmable logic array (PLA), a field-programmable gate array (FPGA), a complex programmable logic device (CPLD), implemented as fixed-function logic hardware using, for example, application-specific integrated circuit (ASIC), complementary metal-oxide-semiconductor (CMOS), or transistor-transistor (TTL) technology, or any combination thereof.

[0033] For example, the computer program code for implementing the outputs of information provider 54, analyzer 24, and flip-flop 26 can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as C++, and traditional procedural programming languages ​​such as "C" or similar languages. Furthermore, information provider 54, analyzer 24, and flip-flop 26 can be implemented using any of the circuit techniques mentioned herein. Additionally, the logic instructions may include assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, state setting data, configuration data for integrated circuits, and state information that personalizes native electronic circuits and / or other structural components (e.g., host processor, central processing unit / CPU, microcontroller, etc.).

[0034] Information provider 54 (which may be firmware such as a Basic Input / Output System (BIOS)) may include provider information. For example, provider information may include information about which of one or more devices can preferably operate with discrete graphics processor 30, computing architecture 52 itself, and how MUX 32, 34, 36, 38 are configured to operate with it, user-configured information corresponding to whether the user prefers to use discrete graphics processor 30 or integrated graphics processor 28, or information about which of one or more devices can preferably operate with integrated graphics processor 28. The provider information may also include information about the correspondence between the type of connected display device and the outputs of MUX 32, 34, 36, 38. For example, if the display device is an HMD or a graphics-intensive display device, the provider information may include data instructing discrete graphics processor 30 to be electrically connected to the display device, and data on how MUX 32, 34, 36, 38 are configured to achieve the required electrical connection. Information provider 54 may provide provider information to analyzer 24. For example, information provider 54 (e.g., BIOS) can communicate with analyzer 24 via the Advanced Configuration and Power Interface (ACPI) specification.

[0035] Analyzer 24 can detect whether the display is connected to the computing system and receive display information (e.g., extended display identification data) from the display to determine whether each of the display devices 42, 44, 46 is to be electrically connected to the integrated graphics processor 28 or the discrete graphics processor 30. In some embodiments, analyzer 24 can receive display information and detect the display based on information from the microcontroller or integrated graphics processor driver of the integrated graphics processor 28.

[0036] In some embodiments, the integrated graphics processor driver can detect a display that may initially be connected to the integrated graphics processor 28. For example, analyzer 24 can determine that display device 42 is connected, and then analyzer 24 can determine whether MUX 32, 34 outputs the output of discrete graphics processor 30 or the output of integrated graphics processor 28 based on provider information and display information (e.g., EDID) of display device 42. For example, when display device 42 is a graphics-intensive display device (such as an HMD), analyzer 24 can determine that MUX 32, 34 will be configured to connect discrete graphics processor 30 to display device 42. In some embodiments, analyzer 24 can compare display information with provider information to determine whether MUX 32, 34 will be configured to output the output of discrete graphics processor 30 or the output of integrated graphics processor 28. In some embodiments, analyzer 24 may be an integrated graphics processor driver.

[0037] Analyzer 24 can provide a determination to trigger 26 (e.g., a BIOS, microcontroller, integrated graphics driver of integrated graphics processor 28, device driver, or firmware). For example, trigger 26 can control MUX 32, 34 to reflect the determination of analyzer 24. In some embodiments, trigger 26 can control MUX 32, 34 based on provider information from information provider 54 and the determination of analyzer 24. For example, provider information may include information about how MUX 32, 34 can be controlled via, for example, general purpose input / output (GPIO) pins. Trigger 26 can set the GPIO pins to provide appropriate voltages to MUX 32, 34 to reflect the determination. Provider information may include different ways in which trigger 26 controls the GPIO pins, such as by writing to specific memory or another mechanism, or by utilizing firmware. Although display device 42 and MUX 32, 24 have been discussed above, MUX 36, 38 can be similarly configured based on the configuration of display devices 44, 46, respectively.

[0038] In some embodiments, the integrated graphics driver of the integrated graphics processor 28 may include at least one of an information provider 54 and an analyzer 24. In other words, the integrated graphics driver may include both an information provider 54 and an analyzer 24.

[0039] For example, to operate as an information provider 54, the integrated graphics driver may include a “whitelist” of display devices that will operate with the discrete graphics processor 30. Upon detecting that a display device 42 is connected to the integrated graphics processor 30 and directly in response to this connection, the integrated graphics driver will detect and receive information (e.g., EDID) from the display device 42. The integrated graphics driver may operate as an analyzer 24 to compare the information with the whitelist. If the display device 42 is on the list, the integrated graphics driver will decide that the display device 42 should be connected to the discrete graphics processor 30. The integrated graphics driver may then provide this decision to trigger 26. Trigger 26 may be, for example, firmware such as a BIOS. The firmware may control MUX 32, 34, for example, by writing to memory to provide the output of the discrete graphics processor 30 to the display device 42. In some embodiments, trigger 26 (e.g., firmware) may control another device, such as the integrated graphics processor 28, which in turn controls MUX 32, 34 via select lines. In some embodiments, trigger 26 can control GPIO to control MUX 32, 34. For example, trigger 26 can write to specific memory to change GPIO to control MUX 32, 34. As described above, other MUX 36, 38 and display devices 44, 46 can be driven and controlled similarly.

[0040] In some embodiments, the microcontroller of the integrated graphics processor 28 may be operated by a trigger 26, wherein the output of the integrated graphics processor 28 is provided to the selection lines of the MUX 32, 34, 36, 38. In some embodiments, the drivers of the host controller 40 may be an information provider 54 and an analyzer 24. Therefore, various implementations of the information provider 54, the analyzer 24, and the trigger 26 are possible.

[0041] In some embodiments, MUX 32, 34 can provide different outputs depending on whether more than one display device is connected to the computing system via host controller 40. For example, MUX 32 can output the output of discrete graphics processor 30, and MUX 34 can output the output of integrated graphics processor 28.

[0042] In some embodiments, the discrete graphics driver of the discrete graphics processor 30 may be an analyzer 24. For example, if a display device 42 is connected to the discrete graphics processor 30 via the hybrid switching described above, the discrete graphics driver may monitor the display interface 48 (e.g., a port of a computing system) to which the display device 42 is connected. The discrete graphics driver may detect whether the display device 42 has been disconnected from the display interface 48. The discrete graphics driver may then determine that MUX 32, 34 should be reset to connect the integrated graphics processor 28 to the output of MUX 32, 34. Trigger 26 (which may be a BIOS or the discrete graphics driver) may then reset MUX 32, 34 to connect the integrated graphics processor 28 to the output of MUX 32, 34. The integrated graphics driver may then monitor the display interface 48 to determine whether another display device is connected to the display interface 48.

[0043] In some embodiments, the hybrid state of MUX 32, 34, 36, and 38 can be overridden by the user. For example, the user can overridden the hybrid state to ensure that the output of discrete graphics processor 30 or integrated graphics processor 28 is always output through at least one of MUX 32, 34, 36, and 38. For example, the user can overridden the hybrid state via BIOS boot options so that MUX 32 and 34 always output the output of discrete graphics processor 30. In this case, analyzer 24 (e.g., BIOS) can suppress communication to the drivers (e.g., integrated graphics driver and discrete graphics driver) (e.g., ASL communication of MUX 32 and 34 port information) to effectively disable the ability of other components to switch the output of MUX 32 and 34. Trigger 26 (e.g., BIOS) will also control MUX 32 and 34 to output the output of discrete graphics processor 30.

[0044] Figure 2 A method 70 for operating a semiconductor packaged device to achieve hybrid switching is illustrated. Method 70 may be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc.; as configurable logic such as, for example, PLA, FPGA, CPLD; as fixed-function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL; or any combination thereof.

[0045] In processing block 72, the logic coupled to the substrate can obtain information from the connected display device. This information may correspond to whether the connected display device operates with an integrated graphics processor or a discrete graphics processor. In block 74, logic is formed to determine whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device. The logic may make this determination based on the information. The logic may include an information provider, an analyzer, and triggers, as described above.

[0046] Figure 3 A method 1300 for hybrid switching is illustrated. As indicated by box 1326, method 1300 can occur during the boot-up process (or pre-operating system initialization) of a computing system including an integrated graphics processor. That is, each of steps 1302-1326 can occur during the boot or wake-up process of the computing system. A connected display device can be connected to the computing system.

[0047] Furthermore, method 1300 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc., as configurable logic such as PLA, FPGA, CPLD, etc., as fixed-function logic hardware using circuit technologies such as ASIC, CMOS, or TTL, etc., or any combination thereof.

[0048] In box 1326, the startup process is initialized. In box 1302, the display interface is detected. If the computing system has several display interfaces, they are detected sequentially. In the example shown, box 1304 can determine whether a display device is detected at the display interface. For example, if a voltage is detected on the hot-plug detection (HPD) pin connected to the display interface, a display device is connected to the display interface. If no display device is detected, box 1316 can determine whether the display interface is the last display interface of the computing system. If not, box 1302 detects the next display interface.

[0049] If a display device is detected in box 1304, method 1300 can proceed to box 1306. In box 1306, it can be determined whether multiple display devices are connected together and connected to a display interface. For example, multi-streaming can be used in a daisy-chained display setup, where one device is connected to the display interface. If multiple devices are detected, box 1308 can determine whether the display device detected in box 1304 is the first of the multiple display devices to be detected, or the first device in the daisy chain. If not, display devices can be enumerated in box 1318 based on the first display device among the multiple display devices. For example, the display device can be displayed based on the settings of the first display device among the multiple display devices (e.g., whether it receives output from a discrete graphics processor or an integrated graphics processor). This can cause the integrated graphics processor to "see" the plug being unplugged even when the discrete graphics processor sees an "insert event". Furthermore, the associated MUX can remain unchanged. Enumeration can also include allowing the device driver (e.g., an integrated graphics processor driver or a discrete graphics processor driver) to drive the display device to display an image.

[0050] If it is determined in box 1308 that the display device is the first display device, then in box 1328, the configuration information of the display device is determined. For example, information about the display device (e.g., EDID) can be obtained and retrieved from the display device. In box 1312, it is determined from the configuration information whether the display device will utilize a discrete graphics processor. For example, if the configuration information indicates that the display device is an HMD, it can be determined that the display device will utilize a discrete graphics processor. The information about the display device can be compared with other information (e.g., a whitelist) to determine whether the display device should utilize a discrete graphics processor. If the display device does utilize a discrete graphics processor, then in box 1314, enumeration (such as displaying the display device and allowing some applications of the operating system to access the display device) is deferred, and information identifying the display interface can be stored.

[0051] Enumeration of the display device is postponed if the discrete graphics driver for the discrete graphics processor has not yet been initialized. For example, during startup, the driver, including the discrete graphics driver, may not have been initialized yet. Therefore, enumeration of the display device can be delayed until it is determined that the discrete graphics driver is available. Method 1300 can proceed to block 1316, where it can determine whether the display interface is the final display interface of the computing system. If so, then in circle A shown, it can be referenced as... Figure 4 The described continuation method 1300.

[0052] If, in box 1312, the display device does not utilize a discrete graphics processor, then in box 1322, the display device is normally enumerated. For example, by operating the display using an integrated graphics processor, the MUX is modified to electrically connect the integrated graphics processor to the display device, notifying the operating system that the display device is available upon initialization, or by operating the display using the display device driver of the display device. Method 1300 proceeds from box 1322 to box 1316, and is similar to that described above.

[0053] Figure 4 This illustrates a continuation of method 1300. Each of blocks 1452-1472 can occur after the boot process is complete and when the operating system is initialized.

[0054] In box 1450, the operating system is initialized. In box 1452, it is determined whether the discrete graphics processor (GPU) is ready. For example, if the GPU driver is available, the GPU is also available. If the GPU is not yet ready, a timer is incremented in box 1454 (illustrated). ASL information regarding how to control the MUX can also be suppressed until the GPU is available. Box 1456 determines whether the timer has reached a timeout value. The timeout value can be set to reflect the probability that the GPU will become available. For example, if the GPU was available in a previous instance (e.g., another time when the user used the computing system), the timeout value may be higher because the confidence that the GPU will become available is likely to be high. However, if the GPU was previously unavailable in a previous instance (e.g., a previous operation of the computing device), the timeout value may be lower because the GPU is more likely to remain unavailable. Therefore, the timer can be adjusted by setting the timer value high to reflect whether the software associated with the GPU (e.g., the GPU driver) is still being initialized. Conversely, if a discrete graphics processor itself might be unavailable, the timer value can be set to a lower value. If the timeout value has not yet been reached, repeat box 1402 as shown and determine again whether a discrete graphics processor is available.

[0055] If the timer has reached its timeout value in box 1456, it is determined in box 1458 that the graphics processor is unavailable. Method 1300 can then proceed to box 1460. In some embodiments, in box 1458, a prompt may also be displayed to the user indicating that the discrete graphics processor is unavailable.

[0056] If a discrete graphics processor is available in block 1452, then in block 1460, a delayed display interface corresponding to one of the delayed display devices is detected. The delayed display device corresponds to the display device utilizing the discrete graphics processor identified in block 1312. The delayed display interface can be determined based on the identification information of the display interface stored in block 1314. Block 1460 determines whether the delayed display device is still connected to the delayed display interface. For example, if the delayed display device is still connected to the display interface, a voltage on the HPD can be received. If no display device is detected, then the illustrated block 1464 determines whether the current delayed display interface is the last delayed display interface. If so, then block 1470 can monitor the display interface. If, in block 1464, the current delayed display interface is not the last delayed display interface, then the illustrated block 1410 detects the next delayed display interface.

[0057] If a delayed display device is detected in box 1462, box 1466 determines whether the discrete graphics processor was ever available at box 1452. If the discrete graphics processor was ever available, box 1468 can electrically connect the discrete graphics processor to the display device. For example, a MUX that receives the outputs of both the discrete graphics processor and the integrated graphics processor can be controlled to provide the discrete graphics processor output to the display device.

[0058] If it is determined in box 1466 that the discrete graphics processor was previously unavailable, then in box 1472, the integrated graphics processor can be electrically connected to the display device. For example, a MUX that receives the outputs of both the discrete and integrated graphics processors can be controlled to provide the output of the integrated graphics processor to the display device. If the discrete graphics processor is unavailable, these delayed display devices can be connected to the output of the integrated graphics processor to avoid an error state.

[0059] Following boxes 1468 and 1472, method 1300 moves to box 1464, as described above. Method 1300 can be implemented using an integrated graphics driver. Method 1300 can also be implemented using logic (e.g., an analyzer, an information provider, and triggers).

[0060] Figure 5 A hybrid switching method 1500 is illustrated. Method 1500 can occur after the startup or wake-up sequence is completed and when the operating system is initialized.

[0061] Method 1500 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc., as configurable logic such as PLA, FPGA, CPLD, etc., as fixed-function logic hardware using circuit technologies such as ASIC, CMOS, or TTL, etc., or any combination thereof.

[0062] In box 1502, a new connection event is detected. This event can indicate that a previously unconnected device has been connected to the computing system via the display interface. In box 1504, the display device can be verified as connected. For example, if a voltage is detected on the HPD pin connected to the display interface, the display device may have been connected to the display interface.

[0063] Box 1506 determines whether multiple display devices are connected together and connected to the display interface. For example, multi-stream transport can be used in a daisy-chained display setup where one device is connected to the display interface. If no multiple devices are detected, the method proceeds to box 1510, which will be discussed below.

[0064] If multiple devices are detected, box 1508 can determine whether the display device detected in box 1504 is the first display device among the multiple display devices to be detected. If not, the display devices can be enumerated or displayed in box 1518 based on the first display device and based on the settings of the first display device among the multiple display devices (e.g., whether it receives the output of a discrete processing processor or an integrated processing processor).

[0065] If box 1508 determines that the display device is the first display device among a plurality of display devices, then box 1510 reads the display device configuration of the display device. For example, configuration information (e.g., EDID) of the display device can be obtained and retrieved from the display device.

[0066] In box 1512, it is determined whether the display device should utilize a discrete graphics processor based on configuration information. For example, if the configuration information indicates that the display device is an HMD, it can be determined that the display device will utilize a discrete graphics processor. The configuration information can be compared with a whitelist as described above to determine whether the display device should operate with a discrete graphics processor. If the display device does utilize a discrete graphics processor, then in box 1514, the discrete graphics processor can be electrically connected to the display device. For example, a MUX that receives the outputs of both the discrete graphics processor and the integrated graphics processor can be controlled to provide the output of the discrete graphics processor to the display device.

[0067] If the display device in box 1512 does not utilize a discrete graphics processor, then box 1516 can properly enumerate the display device. For example, the display device can be driven by an integrated graphics processor. In box 1518, changes to the display interface are monitored.

[0068] The method 1500 described above can be implemented by an integrated graphics driver. The method 1500 described above can be implemented by logic (e.g., an analyzer, an information provider, and a trigger).

[0069] Figure 6 A method 1600 for a discrete graphics processor and a discrete graphics processor driver is illustrated. Method 1600 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc.; as configurable logic such as, for example, PLA, FPGA, CPLD; as fixed-function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL; or any combination thereof.

[0070] In box 1602, the discrete graphics driver for the discrete graphics processor is initialized. In box 1604, the system (e.g., BIOS) is notified that the discrete graphics processor and discrete graphics driver are available. Method 1600 may correspond to box 1402 of method 1300, where it is determined whether the discrete graphics processor is available. In box 1606, each delayed display port may be enumerated.

[0071] Figure 7 A hybrid switching method 1700 is illustrated. Method 1700 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc., implemented as configurable logic such as PLA, FPGA, CPLD, etc., implemented as fixed-function logic hardware using circuit technologies such as ASIC, CMOS, or TTL, etc., or any combination thereof.

[0072] In box 1702, the display device utilizes a discrete graphics processor. For example, a MUX that receives the outputs of both the discrete graphics processor and the integrated graphics processor can be controlled to provide the output of the discrete graphics processor to the display device via a display interface, as discussed herein.

[0073] At box 1704, the display device is detected as unplugged. For example, the display interface may not have HPD voltage. In box 1706, the MUX is reset to hybrid mode, causing the output of the integrated graphics processor, rather than the output of the discrete graphics processor, to be connected to the display interface. In box 1708, the unplugging event can be processed. The above method 1300 can be implemented by a discrete graphics driver.

[0074] Figure 8 Method 1800 for a power-down process is illustrated. Method 1800 may be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc., implemented as configurable logic such as, for example, PLA, FPGA, CPLD, implemented as fixed-function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL technology, or any combination thereof.

[0075] In box 1802, the discrete graphics processor is powered down. In box 1804, each MUX is reset to hybrid mode, causing the output of the integrated graphics processor, instead of the output of the discrete graphics processor, to be connected to the corresponding display interface of the MUX. In box 1806, the system (e.g., BIOS) is notified that the discrete graphics processor is no longer available.

[0076] Figure 9 A method 1900 for setting the operating mode is illustrated. Method 1900 can be implemented by the BIOS and occurs during the startup, wake-up, or initialization of the computing system. Method 1900 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc., as configurable logic such as, for example, PLA, FPGA, CPLD, or as fixed-function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL, or any combination thereof.

[0077] In block 1902, user preferences can be read. For example, memory can be read to determine preferences. In some embodiments, a BIOS switch is read. In block 1904, it is determined whether to set the discrete graphics processor as the primary graphics processor based on information from the preferences. If the preferences indicate that the discrete graphics processor should operate with the display device, then block 1908 can configure the discrete graphics processor to be electrically connected to the display device via, for example, a MUX as described above. Furthermore, ASL configuration information used to change the MUX to connect either the integrated graphics processor or the discrete graphics processor can be suppressed so that other components (e.g., the integrated graphics driver) cannot change the MUX.

[0078] If, in box 1904, the preference is determined not to indicate that the discrete graphics processor will operate together with the discrete graphics processor, the MUX can be set to a hybrid mode to allow selective electrical connection of either the integrated graphics processor or the discrete graphics processor to the display device, as described above. For example, ASL configuration information for changing the MUX to connect either the integrated graphics processor or the discrete graphics processor can be shared, allowing other components (e.g., the integrated graphics driver) to modify the MUX.

[0079] The user preferences mentioned above can be set by the user through, for example, the BIOS. Preferences can also be set by the display device manufacturer.

[0080] Figure 10 A method 2000 for setting the operating mode is illustrated. Method 2000 can be implemented by the BIOS and occurs during the startup, wake-up, or initialization of the computing system. Method 2000 can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium such as RAM, ROM, PROM, flash memory, etc.; as configurable logic such as, for example, PLA, FPGA, CPLD; as fixed-function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL; or any combination thereof.

[0081] Box 2002 may determine whether a discrete graphics processor is present. For example, box 2002 may determine whether a discrete graphics processor was detected during a previous startup of the computing system that occurred prior to the current startup, or another method for detecting a discrete graphics processor. If no discrete graphics processor is detected, box 2004 may utilize the integrated graphics processor to operate each connected display device. For example, the MUX may be configured to electrically connect the integrated graphics processor to the display device. In box 2006, a hybrid mode may be disabled for each MUX, preventing changes to each MUX to prevent the MUX from electrically disconnecting the integrated graphics processor from the display device. For example, ASL configuration information used to change the MUX to connect either the integrated graphics processor or the discrete graphics processor may be suppressed, preventing other components (e.g., the integrated graphics driver) from changing the MUX. In some embodiments, if no discrete graphics processor is detected, a prompt may also be displayed to the user indicating that the discrete graphics processor is unavailable.

[0082] If, in box 2002, it is determined that a discrete graphics driver was detected in a previous operation of the computing system, then box 2008 reads the preference. Box 2010 may determine whether the preference indicates that the discrete graphics processor is selected to operate with the display device, such that the discrete graphics processor is the primary driver of the display device. If, in box 2010, the preference indicates that a discrete graphics processor should be selected, then box 2014 sets the discrete graphics processor to operate with the display device. For example, the MUX may be set to electrically connect the discrete graphics processor to the display device. Then, method 2000 may proceed to box 2016, where it is determined whether the discrete graphics driver of the discrete graphics processor is powered on. If the discrete graphics driver is not powered on, then box 2012 may enable the MUX's hybrid mode, allowing the MUX to be controlled to allow the display device to be connected to the integrated graphics processor. If the discrete graphics driver is indeed powered on, then box 2018 maintains the operation of the display device and the discrete graphics processor, such that the display device is driven by the discrete graphics processor. Method 2000 can proceed to block 2006, where the blending mode is disabled so that the discrete graphics processor is not electrically disconnected from the display device. For example, ASL configuration information (which may be in ACPI source language) used to change the MUX to connect either the integrated graphics processor or the discrete graphics processor can be suppressed so that other components (e.g., the integrated graphics driver) cannot change the MUX.

[0083] If in box 2010 the preference does not indicate that a discrete graphics processor will be connected to the display device, then method 2000 can proceed to box 2012. In box 2012, a hybrid mode can be enabled to allow the MUX to switch between electrically connecting the display device to the integrated graphics processor and the discrete graphics processor. For example, ASL configuration information for changing the MUX to connect to either the integrated graphics processor or the discrete graphics processor can be shared, allowing other components (e.g., the integrated graphics driver) to modify the MUX.

[0084] The preferences described above can be user preferences. Preferences can also be set by the display device manufacturer. In some embodiments, the same user preferences can be applied to all displays, causing all displays to be mapped to the same integrated or discrete graphics processor. Furthermore, user preferences can indicate whether the display device should be connected to an integrated graphics processor instead of a discrete graphics processor.

[0085] The above method 200 can also be applied to every display device.

[0086] Figure 11 Semiconductor package 2102 is shown. Semiconductor package 2102 can implement method 70 ( Figure 2 ), 1300 Figure 3-4 ), 1500 Figure 5 ), 1600 Figure 6 ), 1700 Figure 7 ), 1800 Figure 8 ), 1900 Figure 9 ) and 2000 Figure 10 One or more aspects of ) and can easily replace the information provider 54, analyzer 24 and trigger 26 already discussed. Figure 1 The illustrated device 2102 includes a substrate 2106 (e.g., silicon, sapphire, gallium arsenide) and logic 2104 (e.g., transistor arrays and other integrated circuit / IC components) coupled to the substrate 2106. The logic 2104 may be at least partially implemented in configurable logic or fixed-function logic hardware. Furthermore, the logic 2104 can detect the successful startup of a discrete graphics processor in a computing system, detect a display device, determine whether to operate the display device using a discrete graphics processor or an integrated graphics processor, and modify associated components (e.g., a MUX) to operate the display device using either a discrete graphics processor or an integrated graphics processor as determined.

[0087] System Overview

[0088] Figure 12 This is a block diagram of a processing system 100 according to an embodiment. In various embodiments, system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, system 100 is a processing platform incorporated within a system-on-a-chip (SoC) integrated circuit for use in a mobile device, handheld device, or embedded device.

[0089] In one embodiment, system 100 may include or incorporate a server-based gaming platform, a game console, including a game and media console, a mobile game console, a handheld game console, or an online game console. In some embodiments, system 100 is a mobile phone, smartphone, tablet computing device, or mobile internet device. Processing system 100 may also include a wearable device (such as a smartwatch, smart glasses, augmented reality, or virtual reality device), coupled to or integrated into the wearable device. In some embodiments, processing system 100 is a television or set-top box device having one or more processors 102 and a graphical interface generated by one or more graphics processors 108.

[0090] In some embodiments, each of the one or more processors 102 includes one or more processor cores 107 for processing instructions that, when executed, perform operations on the system and user software. In some embodiments, each of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). Multiple processor cores 107 may each process different instruction sets 109, which may include instructions for facilitating emulation of other instruction sets. Processor cores 107 may also include other processing means, such as digital signal processors (DSPs).

[0091] In some embodiments, processor 102 includes cache memory 104. Depending on the architecture, processor 102 may have a single internal cache or multiple levels of internal caches. In some embodiments, cache memory is shared among components of processor 102. In some embodiments, processor 102 also uses an external cache (e.g., a Level 3 (L3) cache or a Last Level Cache (LLC)) (not shown), which can be shared among processor core 107 using known cache coherence techniques. Additionally, register file 106 is included in processor 102, which may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while others may be specific to the design of processor 102.

[0092] In some embodiments, one or more processors 102 are coupled to one or more interface buses 110 for transmitting communication signals, such as address, data, or control signals, between the processors 102 and other components in the system 100. In one embodiment, the interface bus 110 may be a processor bus, such as a version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In one embodiment, the processor(s) 102 include an integrated memory controller 116 and a platform controller hub (PCH) 130. The memory controller 116 facilitates communication between memory devices and other components of the system 100, while the platform controller hub (PCH) 130 provides connectivity to I / O devices via a local I / O bus.

[0093] Memory device 120 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or some other memory device with suitable performance for use as processing memory. In one embodiment, memory device 120 may operate as system memory of system 100 for storing data 122 and instructions 121 for use when the one or more processors 102 execute an application or process. Memory controller 116 is also coupled to an optional external graphics processor 112, which may communicate with the one or more graphics processors 108 in processor 102 to perform graphics and media operations. In some embodiments, display device 111 may be connected to processor(s) 102. Display device 111 may be one or more of the following: an internal display device, such as in a mobile electronic device or a laptop device; or an external display device attached via a display interface (e.g., a display port, etc.). In one embodiment, display device 111 may be a head-mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) or augmented reality (AR) applications.

[0094] In some embodiments, the platform controller hub 130 enables peripheral devices to connect to the memory device 120 and the processor 102 via a high-speed I / O bus. I / O peripheral devices include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., a hard disk drive, flash memory, etc.). The data storage device 124 may be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a peripheral component interconnect bus (e.g., PCI, PCI Express). The touch sensor 125 may include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. The firmware interface 128 enables communication with system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). The network controller 134 enables network connectivity to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. In one embodiment, the audio controller 146 is a multi-channel high-definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., a Personal System 2 (PS / 2)) devices to the system. The platform controller hub 130 may also be connected to one or more Universal Serial Bus (USB) controllers 142 to connect input devices, such as a keyboard and mouse combination 143, a camera 144, or other USB input devices.

[0095] It will be appreciated that the illustrated system 100 is exemplary and not limiting, as other types of data processing systems configured differently may also be used. For example, instances of the memory controller 116 and platform controller hub 130 may be integrated into a discrete external graphics processor, such as external graphics processor 112. In one embodiment, the platform controller hub 130 and / or memory controller 160 may be external to the one or more processors 102. For example, system 100 may include an external memory controller 116 and a platform controller hub 130, which may be configured as a memory controller hub and a peripheral controller hub within a system chipset communicating with the processor(s) 102.

[0096] Figure 13 This is a block diagram of an embodiment of processor 200, which has one or more processor cores 202A to 202N, an integrated memory controller 214, and an integrated graphics processor 208. Figure 13 Those elements having the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. Processor 200 may include, and include, additional cores 202N, indicated by dashed boxes. Each processor core 202A to 202N includes one or more internal cache units 204A to 204N. In some embodiments, each processor core may also access one or more shared cache units 206.

[0097] Internal cache units 204A to 204N and shared cache unit 206 represent the cache memory hierarchy within processor 200. The cache memory hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of shared intermediate cache, such as Level 2 (L2), Level 3 (L3), Level 4 (L4), or other levels of cache, wherein the highest-level cache is classified as LLC before external memory. In some embodiments, cache coherence logic maintains coherence between each cache unit 206 and 204A to 204N.

[0098] In some embodiments, the processor 200 may further include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI or PCI Fast buses. The system agent core 210 provides management functions for each processor component. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).

[0099] In some embodiments, one or more of processor cores 202A to 202N include support for simultaneous multithreading. In this embodiment, system agent core 210 includes components for coordinating and operating cores 202A to 202N during multithreaded processing. Additionally, system agent core 210 may also include a power control unit (PCU) including logic and components for regulating the power states of processor cores 202A to 202N and a graphics processor 208.

[0100] In some embodiments, processor 200 further includes a graphics processor 208 for performing graphics processing operations. In some embodiments, graphics processor 208 is coupled to a shared cache unit 206 and a system proxy core 210, the system proxy core including one or more integrated memory controllers 214. In some embodiments, system proxy core 210 further includes a display controller 211 to drive graphics processor output to one or more coupled displays. In some embodiments, display controller 211 may also be a separate module coupled to the graphics processor via at least one interconnect, or it may be integrated within graphics processor 208.

[0101] In some embodiments, ring-based interconnect units 212 are used to couple internal components of processor 200. However, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, including those well known in the art, may be used. In some embodiments, graphics processor 208 is coupled to ring interconnect 212 via I / O link 213.

[0102] Exemplary I / O link 213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and a high-performance embedded memory module 218 (such as an eDRAM module). In some embodiments, each of the processor cores 202A to 202N and the graphics processor 208 uses the embedded memory module 218 as a shared final-level cache.

[0103] In some embodiments, processor cores 202A to 202N are homogeneous cores executing the same instruction set architecture. In another embodiment, processor cores 202A to 202N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more of processor cores 202A to 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, processor cores 202A to 202N are homogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled to one or more power cores with lower power consumption. Additionally, processor 200 can be implemented on one or more chips or implemented as a SoC integrated circuit having, among other components, the components shown.

[0104] Figure 14 This is a block diagram of a graphics processing unit 300, which may be a discrete graphics processing unit or a graphics processing unit integrated with multiple processing cores. In some embodiments, the graphics processing unit communicates with memory via a mapped I / O interface to registers on the graphics processing unit and using commands placed in processor memory. In some embodiments, the graphics processing unit 300 includes a memory interface 314 for accessing memory. The memory interface 314 may be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.

[0105] In some embodiments, the graphics processor 300 further includes a display controller 302 for driving display output data to the display device 320. The display controller 302 includes hardware for one or more overlapping planes of the display and a multilayer video or user interface element. The display device 320 may be an internal or external display device. In one embodiment, the display device 320 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding, decoding, or converting media codes to, from, or between one or more media encoding formats, including but not limited to: Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Decoding (AVC) formats (such as H.264 / MPEG-4 AVC), and Society of Motion Picture & Television Engineers (SMPTE) 421M / VC-1, and Joint Group of Picture Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG)).

[0106] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterizer operations, including, for example, bit boundary block transfer. However, in one embodiment, 2D graphics operations are performed using one or more components of a graphics processing engine (GPE) 310. In some embodiments, the GPE 310 is a computational engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.

[0107] In some embodiments, GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering 3D images and scenes using processing functions acting on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed functional elements that perform various tasks within elements and / or generated execution threads of the 3D / media subsystem 315. While the 3D pipeline 312 can be used to perform media operations, embodiments of GPE 310 also include a media pipeline 316 specifically for performing media operations such as video post-processing and image enhancement.

[0108] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units to perform one or more specialized media operations, such as video decoding acceleration, video deinterleaving, and video encoding acceleration, in place of or on behalf of the video codec engine 306. In some embodiments, the media pipeline 316 further includes a thread generation unit to generate threads for execution on the 3D / media subsystem 315. The generated threads perform calculations on media operations on one or more graphics execution units included in the 3D / media subsystem 315.

[0109] In some embodiments, the 3D / media subsystem 315 includes logic for executing threads generated by the 3D pipeline 312 and the media pipeline 316. In one embodiment, the pipelines send thread execution requests to the 3D / media subsystem 315, the 3D / media subsystem including thread dispatch logic for arbitrating and dispatching requests to available thread execution resources. Execution resources include an array of graphics execution units for processing 3D and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory (including registers and addressable memory) for sharing data between threads and for storing output data.

[0110] Graphics processing engine

[0111] Figure 15This is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is... Figure 14 The image shows a version of GPE 310. Figure 15 Those elements having the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. For example, shown Figure 14 The 3D pipeline 312 and media pipeline 316 are included. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example, and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.

[0112] In some embodiments, GPE 410 is coupled to or includes command stream converter 403, which provides command streams to 3D pipeline 312 and / or media pipeline 316. In some embodiments, command stream converter 403 is coupled to memory, which may be system memory, or one or more of internal cache memory and shared cache memory. In some embodiments, command stream converter 403 receives commands from memory and sends these commands to 3D pipeline 312 and / or media pipeline 316. The commands are instructions obtained from a ring buffer storing instructions for 3D pipeline 312 and media pipeline 316. In one embodiment, the ring buffer may additionally include a batch command buffer storing multiple batches of commands. Commands for 3D pipeline 312 may also include references to data stored in memory, such as, but not limited to, vertex data and geometry data for 3D pipeline 312 and / or image data and memory objects for media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process the commands and data by performing operations via logic within their respective pipelines or by dispatching one or more execution threads to the execution graphics core array 414. In one embodiment, the graphics core array 414 includes one or more graphics core blocks (e.g., multiple graphics cores 415A, multiple graphics cores 415B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources, which includes: general-purpose execution logic and graphics-specific execution logic for performing graphics and computational operations; and fixed-function texture processing logic and / or machine learning and artificial intelligence acceleration logic.

[0113] In various embodiments, the 3D pipeline 312 includes fixed-function logic and programmable logic for processing one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 414. The graphics core array 414 provides a unified block of execution resources for use in processing these shader programs. The multipurpose execution logic (e.g., execution units) within the graphics core(s)(s)415A to 414B of the graphics core array 414 includes support for various 3D API shader languages ​​and can execute multiple synchronous execution threads associated with multiple shaders.

[0114] In some embodiments, the graphics core array 414 further includes execution logic for performing media functions such as video and / or image processing. In one embodiment, in addition to graphics processing operations, the execution unit also includes general-purpose logic programmable to perform parallel general-purpose computing operations. The general-purpose logic can be... Figure 12 (Multiple) processor cores 107 or Figure 13 The general logic within the cores 202A to 202N performs processing operations in parallel or in combination.

[0115] Output data generated by threads executing on the graphics core array 414 can be output to memory in a uniform return buffer (URB) 418. URB 418 can store data from multiple threads. In some embodiments, URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, URB 418 can also be used for synchronization between threads on the graphics core array and fixed-function logic within shared-function logic 420.

[0116] In some embodiments, the graphics core array 414 is scalable, such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance level of the GPE 410. In one embodiment, the execution resources are dynamically scalable, allowing them to be enabled or disabled as needed.

[0117] The graphics core array 414 is coupled to shared function logic 420, which includes multiple resources shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide dedicated supplementary functions to the graphics core array 414. In various embodiments, the shared function logic 420 includes, but is not limited to, sampler 421, math 422, and inter-thread communication (ITC) 423 logic. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.

[0118] Shared functionality is implemented where the demand for a given dedicated function is insufficient to be contained within the graphics core array 414. Instead, a single instance of the dedicated function is implemented as a separate entity within shared function logic 420 and shared among execution resources within the graphics core array 414. The exact set of functions shared and included within the graphics core array 414 varies across embodiments. In some embodiments, a specific shared function widely used by the graphics core array 414 within shared function logic 420 may be included within shared function logic 416 within the graphics core array 414. In various embodiments, shared function logic 416 within the graphics core array 414 may include some or all of the logic within shared function logic 420. In one embodiment, all logic elements within shared function logic 420 may be repeated within shared function logic 416 of the graphics core array 414. In one embodiment, shared function logic 420 is executed to support shared function logic 416 within the graphics core array 414.

[0119] Figure 16 This is a block diagram of the hardware logic of a graphics processor core 500 according to some embodiments described herein. Figure 16 Those elements having the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. In some embodiments, the illustrated graphics processor core 500 includes Figure 15 Within the graphics core array 414. A graphics processor core 500—sometimes referred to as a core slice—can be one or more graphics cores within a modular graphics processor. An example of a graphics processor core 500 is a graphics core slice, and, based on target power envelopes and performance envelopes, a graphics processor as described herein may include multiple graphics core slices. Each graphics core 500 may include a fixed-function block 530 coupled to multiple sub-cores 501A to 501F (also referred to as sub-slices), which include modular general-purpose logic blocks and fixed-function logic blocks.

[0120] In some embodiments, the fixed-function block 530 includes a geometry / fixed-function pipeline 536, which may be shared by all sub-cores of the graphics processor 500, for example, in low-performance and / or low-power graphics processor implementations. In various embodiments, the geometry / fixed-function pipeline 536 includes a 3D fixed-function pipeline (e.g., as in...). Figure 14 and Figure 15 The 3D pipeline (312), video front-end unit, thread deriver and thread dispatcher, and management such as Figure 15 The unified return buffer manager includes unified return buffers such as the unified return buffer 418.

[0121] In one embodiment, fixed function block 530 further includes a graphics SoC interface 537, a graphics microcontroller 538, and a media pipeline 539. The graphics SoC interface 537 provides an interface between the graphics core 500 and other processor cores within the system-on-a-chip integrated circuit. The graphics microcontroller 538 is a programmable subprocessor configurable to manage various functions of the graphics processor 500, including thread dispatch, scheduling, and pre-emption. The media pipeline 539 (e.g., Figure 14 and Figure 15 The media pipeline 316 includes logic for facilitating the decoding, encoding, preprocessing, and / or post-processing of multimedia data, including image and video data. The media pipeline 539 performs media operations via requests for computation or sampling logic within subcores 501 to 501F.

[0122] In one embodiment, SoC interface 537 enables graphics core 500 to communicate with a general-purpose application processor core (e.g., CPU) and / or other components within the SoC, including memory-level architecture elements such as shared final-level cache memory, system RAM, and / or embedded on-chip or package-based DRAM. SoC interface 537 may also enable communication with fixed-function devices within the SoC, such as camera imaging pipelines, and enable the use and / or implementation of global memory atoms that can be shared between graphics core 500 and the CPU within the SoC. SoC interface 537 may also implement power management control for graphics core 500 and enable interfacing between the clock domain of graphics core 500 and other clock domains within the SoC. In one embodiment, SoC interface 537 enables the receipt of command buffers from a command stream converter and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. When a media operation is about to be executed, these commands and instructions can be dispatched to the media pipeline 539, or when a graphics processing operation is about to be executed, these commands and instructions can be dispatched to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 536, geometry and fixed-function pipeline 514).

[0123] The graphics microcontroller 538 can be configured to perform various scheduling and management tasks for the graphics core 500. In one embodiment, the graphics microcontroller 538 can perform graphics and / or computational workload scheduling for the various parallel graphics engines within the execution unit (EU) arrays 502A to 502F and 504A to 504F of the sub-cores 501A to 501F. In this scheduling model, host software executing on the CPU core of the SoC including the graphics core 500 can submit workloads via one of a plurality of graphics processor doorbells, which invokes scheduling operations for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 538 may also facilitate a low-power or idle state of the graphics core 500, thereby providing the graphics core 500 with the ability to save and restore registers within the graphics core 500 across low-power state transitions, independent of the operating system and / or the graphics driver software on the system.

[0124] The graphics core 500 may have more or fewer sub-cores 501A to 501F shown, up to N modular sub-cores. For each group of N sub-cores, the graphics core 500 may also include shared function logic 510, shared memory and / or cache memory 512, geometry / fixed function pipeline 514, and additional fixed function logic 516 for accelerating various graphics and computational processing operations. Figure 15 The shared functional logic 420 is associated with logic units (e.g., sampler logic, mathematical logic, and / or inter-thread communication logic). Shared memory and / or cache memory 512 can be the final-level cache for the set of N sub-cores 501A to 501F within the graphics core 500, and can also act as shared memory accessible by multiple sub-cores. A geometry / fixed-function pipeline 514 can be included within the fixed-function block 530 instead of the geometry / fixed-function pipeline 536, and can include the same or similar logic units.

[0125] In one embodiment, the graphics core 500 includes additional fixed-function logic 516, which may include various fixed-function acceleration logics for use by the graphics core 500. In one embodiment, the additional fixed-function logic 516 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are two geometry pipelines: a full geometry pipeline within geometry / fixed-function pipelines 516 and 536; and a picking pipeline, which is an additional geometry pipeline that may be included within the additional fixed-function logic 516. In one embodiment, the picking pipeline is a simplified version of the full geometry pipeline. The full pipeline and the picking pipeline can execute different instances of the same application, each with a separate context. Position-only shading can hide longer picking runs of discarded triangles, thereby enabling earlier shading completion in some instances. For example, and in one embodiment, the picking pipeline logic within the attached fixed-function logic 516 can execute the position shader in parallel with the main application and typically generates key results faster than a full pipeline, because a full pipeline only extracts and shades the position attributes of vertices without performing rasterization and rendering of pixels to the frame buffer. The picking pipeline can use the generated key results to compute visibility information for all triangles, regardless of whether those triangles were picked. A full pipeline (which may be referred to as the replay pipeline in this example) can consume visibility information to skip picked triangles and shade only the visible triangles that are ultimately passed to the rasterization stage.

[0126] In one embodiment, the additional fixed-function logic 516 may also include machine learning acceleration logic, such as fixed-function matrix multiplication logic, for implementations including machine learning training or inference.

[0127] Each graphics subcore 501A to 501F includes a set of execution resources that can be used to perform graphics operations, media operations, and computational operations in response to requests from the graphics pipeline, media pipeline, or shader program. The graphics subcores 501A to 501F include: multiple EU arrays 502A to 502F, 504A to 504F; thread dispatch and inter-thread communication (TD / IC) logic 503A to 503F; 3D (e.g., texture) samplers 505A to 505F; media samplers 506A to 506F; shader processors 507A to 507F; and shared local memory (SLM) 508A to 508F. EU arrays 502A to 502F and 504A to 504F each include multiple execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logic operations to serve graphics operations, media operations, or computational operations, including graphics programs, media programs, or computational shader programs. TD / IC logic 503A to 503F performs local thread dispatch and thread control operations for execution units within the subcore and facilitates communication between threads executing on the execution units of the subcore. 3D samplers 505A to 505F can read textures or other 3D graphics-related data into memory. The 3D samplers can read texture data in different ways based on the configured sample state and the texture format associated with a given texture. Media samplers 506A to 506F can perform similar read operations based on the type and format associated with media data. In one embodiment, each graphics subcore 501A to 501F may alternately include unified 3D and media samplers. Threads executing on execution units within each of subcores 501A to 501F can utilize shared local memory 508A to 508F within each subcore, so that threads executing within a thread group can use a common on-chip memory pool for execution.

[0128] Execution unit

[0129] Figures 17A to 17B Thread execution logic 600, including an array of processing elements employed in a graphics processor core, is illustrated according to embodiments described herein. Figures 17A to 17B Those elements having the same reference numerals (or names) as those in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein. Figure 17A An overview of thread execution logic 600 is shown, which may include what is shown as having Figure 16 Variants of the hardware logic for each of the 501A to 501F sub-cores. Figure 17B Exemplary internal details of the execution unit are shown.

[0130] like Figure 17A As shown, in some embodiments, thread execution logic 600 includes a shader processor 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array including multiple execution units 608A to 608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 608A, 608B, 608C, 608D, up to 608N-1 and 608N) based on workload computational needs. In one embodiment, the included components are interconnected via an interconnect structure linking to each component. In some embodiments, thread execution logic 600 includes one or more connections to memory (such as system memory or cache memory) via one or more of the instruction cache 606, data port 614, sampler 610, and execution unit arrays 608A to 608N. In some embodiments, each execution unit (e.g., 608A) is an independent programmable general-purpose computing unit capable of executing multiple synchronous hardware threads while processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 608A to 608N is scalable to include any number of individual execution units.

[0131] In some embodiments, execution units 608A to 608N are primarily used to execute shader programs. Shader processor 602 can handle various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 604. In one embodiment, the thread dispatcher includes logic for arbitrating thread initiation requests from the graphics and media pipeline and instantiating the requested threads on one or more execution units 608A to 608N. For example, a geometry pipeline can dispatch vertex processing, tessellation, or geometry processing threads to thread execution logic for processing. In some embodiments, thread dispatcher 604 can also handle runtime thread generation requests from executing shader programs.

[0132] In some embodiments, execution units 608A to 608N support instruction sets that include native support for many standard 3D graphics shader instructions, enabling minimal conversion to execute shader programs from graphics libraries (e.g., Direct3D and OpenGL). These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., computation and media shaders). Each of the execution units 608A to 608N is capable of executing multiple-issue single-instruction multiple-data (SIMD), and multithreaded operation enables an efficient execution environment in the face of high-latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. For pipelines with integer, single-precision floating-point and double-precision floating-point operations, SIMD branching capabilities, logical operations, transcendental operations, and other hybrid operations, execution is multiple-issue per clock cycle. While waiting for data from memory or a shared function, dependency logic within execution units 608A to 608N causes the waiting thread to sleep until the requested data has been returned. While the waiting thread is sleeping, hardware resources may be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution unit may perform operations on a pixel shader, a fragment shader, or another type of shader program that includes different vertex shaders.

[0133] Each execution unit in the execution units 608A to 608N operates on an array of data elements. The number of data elements is the "execution size," or the number of instruction channels. An execution channel is a logical unit that performs data element access, masking, and flow control within instructions. The number of channels may be independent of the number of physical arithmetic logic units (ALUs) or floating-point units (FPUs) for a particular graphics processor. In some embodiments, the execution units 608A to 608N support both integer and floating-point data types.

[0134] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as compressed data types, and the execution unit will process these elements based on their data size. For example, when operating on a 256-bit wide vector, the 256-bit vector is stored in registers, and the execution unit operates on the vector as four individual 64-bit compressed data elements (four times the word length (QW) size), eight individual 32-bit compressed data elements (double the word length (DW) size), sixteen individual 16-bit compressed data elements (word length (W) size), or thirty-two individual 8-bit data elements (byte (B) size). However, different vector widths and register sizes are possible.

[0135] In one embodiment, one or more execution units can be combined into fused execution units 609A to 609N, which have common thread control logic (607A to 607N) for fused EUs. Multiple EUs can be fused into a group of EUs. Each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary depending on the embodiment. Additionally, different SIMD widths, including but not limited to SIMD8, SIMD16, and SIMD32, can be executed for each EU. Each fused graphics execution unit 609A to 609N includes at least two execution units. For example, fused execution unit 609A includes a first EU 608A, a second EU 608B, and common thread control logic 607A for the first EU 608A and the second EU 608B. Thread control logic 607A controls the threads executing on the fused graphics execution unit 609A, thereby allowing each EU within the fused execution units 609A to 609N to execute using a common instruction pointer register.

[0136] One or more internal instruction caches (e.g., 606) are included in the thread execution logic 600 to cache thread instructions of the execution unit. In some embodiments, one or more data caches (e.g., 612) are included for caching thread data during thread execution. In some embodiments, sampler 610 is included for providing texture sampling for 3D operations and media sampling for media operations. In some embodiments, sampler 610 includes dedicated texture or media sampling functions to process texture or media data during the sampling process before providing sampled data to the execution unit.

[0137] During execution, the graphics and media pipeline sends thread initiation requests to thread execution logic 600 via thread generation and dispatch logic. Once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic within shader processor 602 (e.g., pixel shader logic, fragment shader logic, etc.) is invoked to further compute output information and write the results to output surfaces (e.g., color buffers, depth buffers, stencil buffers, etc.). In some embodiments, the pixel shader or fragment shader computes values ​​for vertex attributes interpolated across the rasterized object. In some embodiments, the pixel processor logic within shader processor 602 then executes a pixel or fragment shader program provided by an application programming interface (API). To execute the shader program, shader processor 602 dispatches threads to execution units (e.g., 608A) via thread dispatcher 604. In some embodiments, shader processor 602 uses texture sampling logic in sampler 610 to access texture data in a texture map stored in memory. Arithmetic operations are performed on the texture data and the input geometry data to calculate the pixel color data of each geometric fragment, or to discard one or more pixels without further processing.

[0138] In some embodiments, data port 614 provides a memory access mechanism for thread execution logic 600 to output processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, data port 614 includes or is coupled to one or more cache memories (e.g., data cache 612) to cache data via the data port for memory access.

[0139] like Figure 17B As shown, the graphics execution unit 608 may include an instruction fetch unit 637, a general-purpose register file array (GRF) 624, an architecture register file array (ARF) 626, a thread arbiter 622, a send unit 630, a branch unit 632, a set of SIMD floating-point units (FPUs) 634, and a set of dedicated integer SIMD ALUs 635 in one embodiment. The GRF 624 and ARF 626 include the set of general-purpose register files and architecture register files associated with each synchronized hardware thread that may be active in the graphics execution unit 608. In one embodiment, per-thread architecture state is maintained in the ARF 626, while data used during thread execution is stored in the GRF 624. The execution state of each thread, including the instruction pointer of each thread, may be maintained in thread-specific registers in the ARF 626.

[0140] In one embodiment, the graphics execution unit 608 has an architecture that is a combination of synchronous multithreading (SMT) and fine-grained interleaved multithreading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronous threads and the target number of registers per execution unit, in which execution unit resources are partitioned across logic used to execute multiple synchronous threads.

[0141] In one embodiment, the graphics execution unit 608 can issue multiple instructions, which can each be different instructions. The thread arbiter 622 of the graphics execution unit thread 608 can dispatch instructions to one of the following for execution: sending unit 630, branching unit 642, or (multiple) SIMD FPUs 634. Each execution thread can access 128 general-purpose registers within the GRF 624, where each register can store 32 bytes accessible as a SIMD 8-element vector with 32-bit data elements. In one embodiment, each execution unit thread accesses 4 kilobytes within the GRF 624, but the embodiment is not limited to this, and more or fewer register resources may be provided in other embodiments. In one embodiment, up to seven threads can execute synchronously, but the number of threads per execution unit may also vary depending on the embodiment. In an embodiment where seven threads can access 4 kilobytes, the GRF 624 can store a total of 28 kilobytes. Flexible addressing modes allow multiple registers to be addressed simultaneously, thereby efficiently constructing wider registers or representing straddle rectangular block data structures.

[0142] In one embodiment, memory operations, sampler operations, and other long-latency system communications are dispatched via a "send" instruction executed by message sending unit 630. In one embodiment, branch instructions are dispatched to dedicated branch unit 632 to facilitate SIMD divergence and eventual convergence.

[0143] In one embodiment, the graphics execution unit 608 includes one or more SIMD floating-point units (FPUs) 634 for performing floating-point operations. In one embodiment, the FPU(s) 634 also support integer computation. In one embodiment, the FPU(s) 634 can perform up to M 32-bit floating-point (or integer) operations in SIMD, or up to 2M 16-bit integer or 16-bit floating-point operations in SIMD. In one embodiment, at least one of the FPUs provides extended mathematical capabilities that support high throughput beyond mathematical functions and double-precision 64-bit floating-point. In some embodiments, a set of 8-bit integer SIMD ALUs 635 also represents and can be specifically optimized to perform operations associated with machine learning computations.

[0144] In one embodiment, an array of multiple instances of the graphics execution unit 608 can be instantiated when graphics subcores are grouped (e.g., sub-slices). For scalability, the product architecture can select the exact number of execution units per subcore group. In one embodiment, the execution unit 608 can execute instructions across multiple execution channels. In a further embodiment, each thread executed on the graphics execution unit 608 is executed on a different channel.

[0145] Figure 18 This is a block diagram illustrating a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set having multiple instruction formats. Solid lines represent components typically included in the execution unit instructions, while dashed lines represent optional components or components included only in subsets of the instructions. In some embodiments, the instruction format 700 described and illustrated are macro instructions, as they are instructions supplied to the execution unit, as opposed to micro-operations generated from instruction decoding (once the instruction is processed).

[0146] In some embodiments, the graphics processor execution unit natively supports instructions using a 128-bit instruction format 710. A 64-bit compact instruction format 730 can be used for some instructions based on the selected instruction, multiple instruction options, and the number of operands. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted to the 64-bit format 730. The native instructions available in the 64-bit format 730 vary depending on the embodiment. In some embodiments, instructions are partially compressed using a set of index values ​​in an index field 713. The execution unit hardware references a set of compression tables based on the index values ​​and uses the output of the compression tables to reconstruct the native instructions using the 128-bit instruction format 710.

[0147] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel, which represents a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control over certain execution options, such as channel selection (e.g., prediction) and data channel ordering (e.g., blending). For instructions using the 128-bit instruction format 710, the execution size field 716 limits the number of data channels that will be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.

[0148] Some execution unit instructions have up to three operands, including two source operands (src0 720, src1 722) and a destination 718. In some embodiments, the execution unit supports dual-destination instructions, where one of these destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction may be an on-the-fly (e.g., hard-coded) value passed using the instruction.

[0149] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726, which specifies, for example, whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits in the instruction.

[0150] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726 that specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, wherein the byte alignment of the access mode determines the access alignment of the instruction operands. For example, in a first mode, the instruction can use byte-aligned addressing for both source and destination operands, and in a second mode, the instruction can use 16-byte aligned addressing for both source and destination operands.

[0151] In one embodiment, the address mode portion of the access / address mode field 726 determines whether the instruction uses direct or indirect addressing. When using direct register addressing mode, bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate number field in the instruction.

[0152] In some embodiments, instructions are grouped based on the 712-bit opcode field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The precise opcode grouping shown is merely exemplary. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares five most significant bits (MSB), where move (mov) instructions are in the form of 0000xxxxb, and logic instructions are in the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jump (jmp)) includes instructions in the form of 0010xxxxb (e.g., 0x20). The promiscuous instruction group 746 includes a mixture of instructions, including synchronous instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). Parallel math instruction set 748 includes component-based arithmetic instructions (e.g., add, multiply) in the form 0100xxxxb (e.g., 0x40). Parallel math set 748 performs arithmetic operations in parallel across data channels. Vector math set 750 includes arithmetic instructions (e.g., dp4) in the form 0101xxxxb (e.g., 0x50). Vector math set performs arithmetic operations on vector operands, such as dot product.

[0153] Graphics Pipeline

[0154] Figure 19 This is a block diagram of another embodiment of the graphics processor 800. Figure 19 Those elements having the same reference numerals (or names) as those in any other figure herein may operate or function in any manner similar to, but not limited to, those described elsewhere herein.

[0155] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a rendering output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system including one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or by commands issued to the graphics processor 800 via a ring interconnect 802. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command stream converter 803, which supplies instructions to individual components of the geometry pipeline 820 or the media pipeline 830.

[0156] In some embodiments, a command stream converter 803 directs the operation of a vertex acquirer 805, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 803. In some embodiments, the vertex acquirer 805 provides vertex data to a vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, the vertex acquirer 805 and the vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A to 852B via a thread dispatcher 831.

[0157] In some embodiments, execution units 852A to 852B are vector processor arrays having an instruction set for performing graphics and media operations. In some embodiments, execution units 852A to 852B have an attached L1 cache 851, which is dedicated to each array or shared between arrays. The cache may be configured as a data cache, an instruction cache, or a single cache partitioned to contain data and instructions in different partitions.

[0158] In some embodiments, the geometry pipeline 820 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some embodiments, a programmable shell shader 811 configures the tessellation operation. A programmable domain shader 817 provides back-end evaluation of the tessellation output. A tessellation unit 813 operates in the direction of the shell shader 811 and includes dedicated logic for generating a detailed set of geometric objects based on a rough geometry model that is provided as input to the geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation components (e.g., shell shader 811, tessellation unit 813, domain shader 817) can be bypassed.

[0159] In some embodiments, the complete geometry object may be processed by the geometry shader 819 via one or more threads dispatched to the execution units 852A to 852B, or it may proceed directly to the clipper 829. In some embodiments, the geometry shader operates on the entire geometry object (rather than vertices or vertex patches such as those in previous stages of the graphics pipeline). If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 may be programmed by a geometry shader program to perform geometric tessellation when the tessellation unit is disabled.

[0160] Prior to rasterization, clipper 829 processes vertex data. Clipper 829 can be a fixed-function clipper or a programmable clipper with clipping and geometry shader capabilities. In some embodiments, the rasterizer and depth testing unit 873 in the render output pipeline 870 dispatch pixel shaders to convert geometry objects into a per-pixel representation. In some embodiments, pixel shader logic is included in thread execution logic 850. In some embodiments, the application can bypass the rasterizer and depth testing unit 873 and access the unrasterized vertex data via outgoing unit 823.

[0161] The graphics processor 800 has an interconnect bus, interconnect structure, or some other interconnect mechanism that allows data and messages to be transferred among the main components of the graphics processor. In some embodiments, execution units 852A to 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via data port 856 to perform memory accesses and communicate with the processor's rendering output pipeline components. In some embodiments, sampler 854, caches 851, 858, and execution units 852A to 852B each have a separate memory access path. In one embodiment, texture cache 858 may also be configured as a sampler cache.

[0162] In some embodiments, the rendering output pipeline 870 includes a rasterizer and a depth testing unit 873 that converts vertex-based objects into associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / mask unit for performing fixed-function triangle and line rasterization. Associated rendering cache 878 and depth cache 879 are also available in some embodiments. Pixel manipulation unit 877 performs pixel-based operations on the data; however, in some instances, pixel operations associated with 2D operations (e.g., using mixed bit-block image passing) are performed by the 2D engine 841, or alternatively by the display controller 843 using an overlay display plane at display time. In some embodiments, a shared L3 cache 875 is available for all graphics components, allowing data to be shared without using main system memory.

[0163] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front-end 834. In some embodiments, the video front-end 834 receives pipeline commands from a command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front-end 834 processes media commands before sending them to the media engine 837. In some embodiments, the media engine 837 includes a thread generation function for generating threads for dispatch to thread execution logic 850 via a thread dispatcher 831.

[0164] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and coupled to the graphics processor via a ring interconnect 802, or some other interconnect bus or mechanism. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 includes dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system-integrated display device (such as in a laptop computer) or an external display device attached via a display device connector.

[0165] In some embodiments, the geometry pipeline 820 and media pipeline 830 may be configured to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some embodiments, the graphics processor's driver software translates API schedules specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and computing APIs from the Khronos Group. In some embodiments, support may also be provided for Microsoft's Direct3D library. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the open-source computer vision library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a pipeline mapping from future APIs to the graphics processor's pipeline can be made.

[0166] Graphical Pipeline Programming

[0167] Figure 20A This is a block diagram illustrating a graphics processor command format 900 according to some embodiments. Figure 20B This is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 20ASolid lines in the diagram represent components that are typically included in the drawing command, while dashed lines represent components that are optional or included only in a subset of the drawing command. Figure 20A An exemplary graphics processor command format 900 includes data fields for identifying the client 902, a command operation code (opcode) 904, and data 906 for the command. Some commands also include a sub-opcode 905 and a command size 908.

[0168] In some embodiments, client 902 defines a client unit of a graphics device that processes command data. In some embodiments, a graphics processor command parser examines the client field of each command to adjust further processing of the command and route command data to the appropriate client unit. In some embodiments, the graphics processor client unit includes a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads opcode 904 and sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses information within data field 906 to execute the command. For some commands, an explicit command size 908 is desired to define the size of the command. In some embodiments, the command parser automatically determines the size of at least some commands in the command based on the command opcode. In some embodiments, commands are aligned via multiples of double word length.

[0169] Figure 20B The flowchart illustrates an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of a graphics processor uses a version of the illustrated command sequence to initiate, execute, and terminate a set of graphics operations. Sample command sequences are shown and described for illustrative purposes only, and embodiments are not limited to these specific commands or this command sequence. Moreover, the commands may be issued as a batch of commands in a command sequence, such that the graphics processor will process the command sequence in a manner that is at least partially simultaneous.

[0170] In some embodiments, the graphics processor command sequence 910 may begin with a pipeline dump clearing command 912 to cause any active graphics pipeline to complete its current pending commands. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate simultaneously. Pipeline dump clearing is performed to cause the active graphics pipeline to complete any pending commands. In response to pipeline dump clearing, the command parser for the graphics processor will stop command processing until the active rendering engine completes its pending operations and invalidates the associated read cache. Optionally, any data marked as 'dirty' in the render cache may be dumped and cleared into memory. In some embodiments, pipeline dump clearing command 912 may be used for pipeline synchronization or before placing the graphics processor into a low-power state.

[0171] In some embodiments, a pipeline selection command 913 is used when a sequence of commands requires the graphics processor to explicitly switch between pipelines. In some embodiments, only one pipeline selection command 913 is required in an execution context before a pipeline command is issued, unless the context requires issuing commands for two pipelines. In some embodiments, a pipeline dump clearing command 912 is required exactly before the pipeline switch via pipeline selection command 913.

[0172] In some embodiments, pipeline control command 914 configures a graphics pipeline for operation and programs the 3D pipeline 922 and the media pipeline 924. In some embodiments, pipeline control command 914 configures the pipeline state of an active pipeline. In one embodiment, pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within an active pipeline before processing a batch of commands.

[0173] In some embodiments, the return buffer state command 916 is used to configure a set of return buffers for corresponding pipelined write data. Some pipelined operations require allocating, selecting, or configuring one or more return buffers, in which intermediate data is written during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer state 916 includes selecting the size and number of return buffers for a set of pipelined operations.

[0174] The remaining commands in the command sequence vary based on the active pipeline used for the operation. Based on pipeline determination 920, the command sequence is tailored for either the 3D pipeline 922 starting at 3D pipeline state 930, or the media pipeline 924 starting at media pipeline state 940.

[0175] Commands for configuring 3D pipeline states 930 include 3D state setting commands for vertex buffer states, vertex element states, constant color states, depth buffer states, and other state variables to be configured before processing 3D primitive commands. The values ​​of these commands are determined at least in part based on the specific 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass specific pipeline components (if those components will not be used).

[0176] In some embodiments, the 3D primitive 932 command is used to submit 3D primitives to be processed by the 3D pipeline. The command and associated parameters passed to the graphics processor via the 3D primitive 932 command are forwarded to the vertex acquisition function in the graphics pipeline. The vertex acquisition function uses the 3D primitive 932 command data to generate multiple vertex data structures. These vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 command is used to perform vertex operations on the 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution unit.

[0177] In some embodiments, the 3D pipeline 922 is triggered by executing command 934 or an event. In some embodiments, register writing triggers command execution. In some embodiments, execution is triggered via a 'go' or 'kick' command in a command sequence. In one embodiment, pipeline synchronization commands are used to trigger command execution so that the command sequence is cleared via a graphics pipeline dump. The 3D pipeline performs geometry processing on 3D primitives. Once the operation is complete, the resulting geometry is rasterized, and the pixel engine shades the resulting pixels. Additional commands for controlling pixel shading and pixel backend operations may also be included for these operations.

[0178] In some embodiments, when performing media operations, a sequence of graphics processor commands 910 follows the media pipeline 924 path. Generally, the specific purpose and manner of programming the media pipeline 924 depends on the media or computational operation to be performed. During media decoding, specific media decoding operations can be offloaded to the media pipeline. In some embodiments, the media pipeline can also be bypassed, and media decoding can be performed wholly or partially using resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, wherein the graphics processor is used to perform SIMD vector operations using computation shader programs that are not explicitly associated with rendering graphics primitives.

[0179] In some embodiments, the media pipeline 924 is configured in a manner similar to that of the 3D pipeline 922. A set of commands for configuring media pipeline states 940 is dispatched or placed in a command queue before the media object commands 942. In some embodiments, the commands 940 for media pipeline states include data for configuring media pipeline elements that will be used to process media objects. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands 940 for media pipeline states also support the use of one or more pointers to "indirect" state elements that contain a batch of state settings.

[0180] In some embodiments, media object command 942 supplies pointers to a media object for processing by the media pipeline. The media object includes a memory buffer containing video data to be processed. In some embodiments, all media pipeline states must be valid before issuing media object command 942. Once the pipeline states are configured and media object command 942 is queued, media pipeline 924 is triggered via execution command 944 or an equivalent execution event (e.g., register write). The output from media pipeline 924 can then be post-processed by operations provided by 3D pipeline 922 or media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.

[0181] Graphical software architecture

[0182] Figure 21 An exemplary graphics software architecture of a data processing system 1000 according to some embodiments is illustrated. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.

[0183] In some embodiments, the 3D graphics application 1010 includes one or more shader programs, which include shader instructions 1012. The shader language instructions may employ a high-level shader language, such as High-Level Shading Language (HLSL) or OpenGL Shading Language (GLSL). The application also includes executable instructions 1014, which employ a machine language suitable for execution by a general-purpose processor core 1034. The application also includes graphics objects 1016 defined by vertex data.

[0184] In some embodiments, the operating system 1020 is from Microsoft Corporation. The operating system 1020 may be a dedicated UNIX-like operating system or an open-source UNIX-like operating system using a variant of the Linux kernel. The operating system 1020 may support graphics APIs 1022, such as the Direct3D API, OpenGL API, or Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. This compilation may be just-in-time (JIT) compilation or pre-compilation of the application-executable shaders. In some embodiments, high-level shaders are compiled into low-level shaders during the compilation of the 3D graphics application 1010. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the standard Portable Intermediate Representation (SPIR) used by the Vulkan API.

[0185] In some embodiments, the user-mode graphics driver 1026 includes a back-end shader compiler 1027 for translating shader instructions 1012 into a hardware-specific representation. When using the OpenGL API, shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses operating system kernel-mode functionality 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.

[0186] IP core implementation

[0187] One or more aspects of at least one embodiment can be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions representing various logic within a processor. When read by a machine, these instructions can cause the machine to manufacture logic for performing the techniques described herein. Such representations (referred to as “IP cores”) are reusable units of logic for an integrated circuit, which can be stored on a tangible, machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model can be supplied to various consumers or manufacturing facilities that load the hardware model onto manufacturing machines that manufacture integrated circuits. Integrated circuits can be manufactured such that the circuits perform the operations described in association with any of the embodiments described herein.

[0188] Figure 22A This is a block diagram illustrating an IP core development system 1100 that can be used to manufacture integrated circuits to perform operations, according to an embodiment. The IP core development system 1100 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build entire integrated circuits (e.g., SOC integrated circuits). Design facility 1130 can generate software simulations 1110 of the IP core design using a high-level programming language (e.g., C / C++). Software simulation 1110 can be used to design, test, and verify the behavior of the IP core using simulation model 1112. Simulation model 1112 can include functional, behavioral, and / or timing simulations. Register transfer level (RTL) designs 1115 can then be created or synthesized from simulation model 1112. RTL design 1115 is an abstraction of the behavior of an integrated circuit (including associated logic performed using the modeled digital signals) that models the flow of digital signals between hardware registers. In addition to RTL design 1115, lower-level designs at logic or transistor levels can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0189] The RTL design 1115 or an equivalent can be further synthesized into a hardware model 1120 by the design facility. This hardware model may employ a Hardware Description Language (HDL) or some other representation of the physical design data. The HDL can be further simulated or tested to validate the IP core design. The IP core design can be stored in non-volatile memory 1140 (e.g., hard disk, flash memory, or any non-volatile storage medium) for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted (e.g., via the Internet) through a wired connection 1150 or a wireless connection 1160. The manufacturing facility 1165 can then fabricate an integrated circuit at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0190] Figure 22BA cross-sectional side view of an integrated circuit package assembly 1170 according to some embodiments described herein is shown. The integrated circuit package assembly 1170 illustrates an implementation of one or more processor or accelerator devices as described herein. The package assembly 1170 includes a plurality of hardware logic units 1172, 1174 connected to a substrate 1180. The logic units 1172, 1174 may be implemented at least partially in configurable logic or fixed-function logic hardware and may include one or more portions of a processor core(s), a graphics processor(s), or any other accelerator device described herein. Each logic unit 1172, 1174 may be implemented within a semiconductor die and coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic units 1172, 1174 and the substrate 1180 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, interconnect structure 1173 may be configured to route electrical signals, such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of logic 1172, 1174. In some embodiments, substrate 1180 is an epoxy-based laminated substrate. In other embodiments, package substrate 1180 may include other suitable types of substrates. Package assembly 1170 may be connected to other electrical devices via package interconnect 1183. Package interconnect 1183 may be coupled to the surface of substrate 1180 to route electrical signals to other electrical devices, such as a motherboard, other chipsets, or multi-chip modules.

[0191] In some embodiments, logic cells 1172, 1174 are electrically coupled to bridge 1182, which is configured to route electrical signals between logic cells 1172, 1174. Bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. Bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide chip-to-chip connections between logic cells 1172, 1174.

[0192] Although two logic units 1172 and 1174 and a bridge 1182 are shown, the embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges, since bridge 1182 can be excluded when logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Furthermore, multiple logic units, dies, and bridges may be connected together in other possible configurations, including a three-dimensional configuration.

[0193] Exemplary System-on-Chip Integrated Circuit

[0194] Figures 23 to 2 5 illustrates exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to those shown, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0195] Figure 23 This is a block diagram illustrating an exemplary system-on-a-chip integrated circuit 1200 that can be fabricated using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., CPU), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, any of which can be modular IP cores from the same or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I / O controller 1220. 2 S / I 2 C controller 1240. Additionally, the integrated circuit may include a display device 1245 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1250 and a Mobile Industry Processor Interface (MIPI) display interface 1255. Storage may be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface may be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Furthermore, some integrated circuits also include an embedded security engine 1270.

[0196] Figures 24A to 24B This is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. Figure 24A An exemplary graphics processor 1310, which can be fabricated using one or more IP cores according to an embodiment, is shown. Figure 24B An additional exemplary graphics processor 1340 of a system-on-a-chip integrated circuit, which can be fabricated using one or more IP cores according to an embodiment, is shown. Figure 24A The graphics processor 1310 is an example of a low-power graphics processor core. Figure 24B The graphics processor 1340 is an example of a higher-performance graphics processor core. Each of the graphics processors 1310 and 1340 can be... Figure 23 A variant of the 1210 graphics processor.

[0197] like Figure 24AAs shown, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A to 1315N (e.g., 1315A, 1315B, 1315C, 1315D, up to 1315N-1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic, such that the vertex processor 1305 is optimized to perform vertex shader program operations, while the one or more fragment processors 1315A to 1315N perform fragment (e.g., pixel) shading operations for use in fragment or pixel shader programs. The vertex processor 1305 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. The fragment processors (multiple) 1315A to 1315N use the primitive and vertex data generated by the vertex processor 1305 to generate frame buffers displayed on a display device. In one embodiment, fragment processors (multiple) 1315A to 1315N are optimized to execute fragment shader programs provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs provided in the Direct 3D API.

[0198] Additionally, the graphics processor 1310 includes one or more memory management units (MMUs) 1320A to 1320B, one or more caches 1325A to 1325B, and one or more circuit interconnects 1330A to 1330B. The one or more MMUs 1320A to 1320B provide virtual-to-physical address mappings for the graphics processor 1310, including vertex processors 1305 and / or (multiple) fragment processors 1315A to 1315N. Besides vertex or image / texture data stored in the one or more caches 1325A to 1325B, the virtual-to-physical address mappings may also reference vertex or image / texture data stored in memory. In one embodiment, the one or more MMUs 1320A to 1320B may interact with system interconnects including those within the same memory. Figure 23 The synchronization of one or more MMUs, including one or more MMUs associated with the one or more application processors 1205, image processor 1215, and / or video processor 1220, enables each processor 1205 to 1220 to participate in a shared or unified virtual memory system. According to an embodiment, the one or more circuit interconnects 1330A to 1330B enable the graphics processor 1310 to interact with other IP cores within the SoC via the SoC's internal bus or via a direct connection.

[0199] like Figure 24B As shown, the graphics processor 1340 includes Figure 24AThe graphics processor 1310 includes one or more MMUs 1320A to 1320B, caches 1325A to 1325B, and circuit interconnects 1330A to 1330B. The graphics processor 1340 includes one or more shader cores 1355A to 1355N (e.g., 1455A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N-1 and 1355N), which provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code to implement vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present may vary in embodiments and implementations. Additionally, the graphics processor 1340 includes an inter-core task manager 1345, which acts as a thread dispatcher for assigning execution threads to one or more shader cores 1355A to 1355N and a chunking unit 1358 for accelerating chunked operations for chunked rendering, in which rendering operations for a particular scene are subdivided in the image space, for example to take advantage of local spatial consistency within the scene or to optimize the use of internal caches.

[0200] Figures 25A to 25B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. Figure 25A A graphics core 1400 is shown, which can be included in... Figure 23 The graphics processor 1210 can be as follows Figure 24B The unified shader cores in the 1355A to 1355N. Figure 25B The 1430 is a highly parallel general-purpose graphics processing unit suitable for deployment on multi-chip modules.

[0201] like Figure 25AAs shown, the graphics core 1400 includes a shared instruction cache 1402, texture units 1418, and cache memory / shared memory 1420 common to the execution resources within the graphics core 1400. The graphics core 1400 may include multiple slices 1401A to 1401N or per core partition, and the graphics processor may include multiple instances of the graphics core 1400. Slices 1401A to 1401N may include supporting logic, including local instruction caches 1404A to 1404N, thread schedulers 1406A to 1406N, thread dispatchers 1408A to 1408N, and a set of registers 1410A. To perform logical operations, slices 1401A to 1401N may include a set of additional functional units (AFU1412A to 1412N), floating-point units (FPU 1414A to 1414N), integer arithmetic logic units (ALU1416 to 1416N), addressing calculation units (ACU 1413A to 1413N), double-precision floating-point units (DPFPU 1415A to 1415N), and matrix processing units (MPU 1417A to 1417N).

[0202] Some of these compute units operate with specific precision. For example, FPUs 1414A to 1414N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while DPFPUs 1415A to 1415N perform double-precision (64-bit) floating-point operations. ALUs 1416A to 1416N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision, and can be configured for mixed-precision operations. MPUs 1417A to 1417N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. MPUs 1417 to 1417N can perform a wide variety of matrix operations to accelerate machine learning application frameworks, including enabling accelerated Generalized Matrix-to-Matrix Multiplication (GEMM). AFUs 1412A to 1412N can perform additional logical operations not supported by floating-point or integer units, including trigonometric function operations (e.g., sine, cosine, etc.).

[0203] like Figure 25BAs shown, the General Purpose Processing Unit (GPGPU) 1430 can be configured to perform highly parallel computational operations by the graphics processing unit array. Additionally, the GPGPU 1430 can be directly linked to other instances of GPGPUs to create multi-GPU clusters, thereby improving the training speed, particularly for deep neural networks. The GPGPU 1430 includes a host interface 1432 for implementing connectivity with a host processor. In one embodiment, the host interface 1432 is a PCI Express interface. However, the host interface can also be a provider-specific communication interface or communication structure. The GPGPU 1430 receives commands from the host processor and uses a global scheduler 1434 to distribute the execution threads associated with those commands to a group of compute clusters 1436A to 1436H. Compute clusters 1436A to 1436H share a cache memory 1438. The cache memory 1438 can act as a higher-level cache of the cache memory within the compute clusters 1436A to 1436H.

[0204] The GPGPU 1430 includes memories 1434A to 1434B coupled to computing clusters 1436A to 1436H via a set of memory controllers 1442A to 1442B. In various embodiments, memories 1434A to 1434B may include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0205] In one embodiment, computing clusters 1436A to 1436H each include a set of graphics cores, such as Figure 25A The graphics core 1400 may include various types of integer logic units and floating-point logic units, which can perform computational operations suitable for machine learning within a certain precision range. For example, in one embodiment, at least a subset of the floating-point units in each of the computing clusters 1436A to 1436H may be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units may be configured to perform 64-bit floating-point operations.

[0206] Multiple instances of GPGPU 1430 can be configured to operate as a computing cluster. The computing mechanisms used by the computing cluster for synchronization and data exchange vary across embodiments. In one embodiment, multiple instances of GPGPU 1430 communicate via host interface 1432. In one embodiment, GPGPU 1430 includes an I / O hub 1439 that couples GPGPU 1430 to GPU links 1440 that implement direct connections to other instances of GPGPU. In one embodiment, GPU link 1440 is coupled to a dedicated GPU-to-GPU bridge that implements communication and synchronization between multiple instances of GPGPU 1430. In one embodiment, GPU link 1440 is coupled to a high-speed interconnect for transmitting and receiving data to and from other GPGPUs or parallel processors. In one embodiment, multiple instances of GPGPU 1430 reside in a separate data processing system and communicate via a network device accessible via host interface 1432. In one embodiment, in addition to or as an alternative to host interface 1432, GPU link 1440 can be configured to implement a connection to a host processor.

[0207] While the illustrated configuration of the GPGPU 1430 can be configured to train neural networks, one embodiment provides an alternative configuration of the GPGPU 1430 that can be deployed within a high-performance or low-power inference platform. In the inference configuration, the GPGPU 1430 includes fewer compute clusters from compute clusters 1436A to 1436H associated with the training configuration. Additionally, the memory technology associated with memories 1434A to 1434B can differ between the inference and training configurations, with higher-bandwidth memory technology dedicated to the training configuration. In one embodiment, the inference configuration of the GPGPU 1430 can support inference-specific instructions. For example, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which are typically used during the inference operations of the deployed neural network.

[0208] Additional notes and examples:

[0209] Example 1 may include a performance-enhanced computing system comprising an integrated graphics processor and logic for forming a determination, based on information from a connected display device, whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device, the information corresponding to whether the connected display device is driven by an integrated graphics processor or a discrete graphics processor.

[0210] Example 2 may include the system described in Example 1, wherein the determination is made during the startup sequence of the computing system, which includes the discrete graphics processor.

[0211] Example 3 may include the system described in Example 2, further including a multiplexer (MUX) electrically connected to the integrated graphics processor, the discrete graphics processor, and the connected display device, wherein the determination is to connect the discrete graphics processor to the connected display, and further wherein, after a startup sequence and based on the determination, logic is used to control the MUX to electrically connect the discrete graphics processor to the connected display device.

[0212] Example 4 may include the system of any one of Examples 1-3, wherein the logic includes a list of display devices driven by a discrete graphics processor, and further wherein the logic is configured to detect information from a connected display device, compare the information with the list, and, if the comparison indicates that the information is in the list, make the determination that the discrete graphics processor will be connected to the connected display device.

[0213] Example 5 may include the system described in Example 1, wherein the determination is to connect a discrete graphics processor to a connected display device, wherein logic is configured to override the determination when the discrete graphics processor is determined to be unavailable, and to electrically connect an integrated graphics processor to the connected display device.

[0214] Example 6 may include the system described in Example 1, wherein logic is used to determine, based on a user's selection, that another connected display device will be electrically connected to the discrete graphics processor.

[0215] Example 7 may include the system described in Example 1, and further include a substrate to which logic is coupled.

[0216] Example 8 may include a semiconductor package device including logic implemented in one or more of configurable logic or fixed-function hardware logic, the logic being configured to determine, based on information from a connected display device, whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device, the information corresponding to whether the connected display device is driven by an integrated graphics processor or a discrete graphics processor.

[0217] Example 9 may include the device described in Example 8, wherein the determination is made during the startup sequence of a computing system including an integrated graphics processor and a discrete graphics processor.

[0218] Example 10 may include the device described in Example 9, wherein the determination is to connect a discrete graphics processor to a connected display device, and after a startup sequence and based on the determination, logic is used to control a multiplexer electrically connected to the discrete graphics processor, the integrated graphics processor, and the connected display device to electrically connect the discrete graphics processor to the connected display device.

[0219] Example 11 may include the device of any one of Examples 8-10, wherein the logic includes a list of display devices driven by a discrete graphics processor, wherein the logic is configured to detect information from a connected display device, compare the information with the list, and, if the comparison indicates that the information is in the list, make the determination that the discrete graphics processor will be connected to the connected display device.

[0220] Example 12 may include the device described in Example 8, wherein the determination is to connect a discrete graphics processor to a connected display device, and wherein logic is configured to override the determination when the discrete graphics processor is determined to be unavailable, and to electrically connect an integrated graphics processor to the connected display device.

[0221] Example 13 may include the device described in Example 8, wherein logic is used to determine, based on a user's selection, that another connected display device will be electrically connected to the discrete graphics processor.

[0222] Example 14 may include the device described in Example 8, and further includes a substrate to which logic is coupled.

[0223] Example 15 may include a method of operating a semiconductor package device, comprising: forming a determination, based on information from a connected display device, whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device, the information corresponding to whether the connected display device will be driven by an integrated graphics processor or a discrete graphics processor.

[0224] Example 16 may include the method of Example 15, wherein the formation occurs during the startup sequence of a computing system including an integrated graphics processor and a discrete graphics processor.

[0225] Example 17 may include the method of Example 16, wherein the determination is to connect a discrete graphics processor to a connected display device, the method further including, after a startup sequence and based on the determination, electrically connecting the discrete graphics processor to the connected display device.

[0226] Example 18 may include the method of any one of Examples 15-17, further comprising detecting information from a connected display device, wherein forming the determination includes comparing the information of the connected display device with a list of devices for display to be driven by a discrete graphics processor, and if the information is in the list, making the determination that the discrete graphics processor will be connected to the connected display device.

[0227] Example 19 may include the method of Example 15, wherein the determination is to connect a discrete graphics processor to a connected display device, the method further comprising determining whether the discrete graphics processor is unavailable, and when the discrete graphics processor is determined to be unavailable, controlling the determination and electrically connecting an integrated graphics processor to the connected display device.

[0228] Example 20 may include the method described in Example 15, further including determining, based on the user's selection, that another connected display device will be electrically connected to the discrete graphics processor.

[0229] Example 21 may include at least one computer-readable storage medium comprising a set of instructions that, when executed by a computing device, cause the computing device to determine, based on information from a connected display device, whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device, wherein the information corresponds to whether the connected display device is driven by an integrated graphics processor or a discrete graphics processor.

[0230] Example 22 may include at least one computer-readable storage medium as described in Example 21, wherein the instructions, when executed, cause the computing device to form the determination during the startup sequence of the computing device, further wherein the computing device includes an integrated graphics processor and a discrete graphics processor.

[0231] Example 23 may include at least one computer-readable storage medium as described in Example 22, wherein the determination is to connect a discrete graphics processor to a connected display device, and the instructions, when executed, cause the computing device to electrically connect the discrete graphics processor to the connected display device after a startup sequence.

[0232] Example 24 may include at least one computer-readable storage medium as described in any one of Examples 21-23, wherein, when executed, the instructions cause a computing device to detect information from a connected display device, compare the information of the connected display device with a list including display devices to be driven by a discrete graphics processor, and if the information is in the list, make the determination that the discrete graphics processor will be connected to the connected display device.

[0233] Example 25 may include at least one computer-readable storage medium as described in Example 21, wherein the determination is to connect a discrete graphics processor to a connected display device, further wherein the instructions, when executed, cause the computing device to determine whether the discrete graphics processor is unavailable, and when it is determined that the discrete graphics processor is unavailable, override the determination and electrically connect an integrated graphics processor to the connected display device.

[0234] Example 26 may include at least one computer-readable storage medium as described in Example 21, wherein the instructions, when executed, cause a computing device to determine, based on a user's selection, that another connected display device will be electrically connected to a discrete graphics processor.

[0235] Example 27 may include a switching device comprising means for forming a determination, based on information from a connected display device, whether to connect a discrete graphics processor or an integrated graphics processor to the connected display device, the information corresponding to whether the connected display device will be driven by an integrated graphics processor or a discrete graphics processor.

[0236] Example 28 may include the device described in Example 27, wherein the means for forming makes the determination during a startup sequence of a computing system including an integrated graphics processor and discrete graphics processors.

[0237] Example 29 may include the device described in Example 28, wherein the determination is to connect a discrete graphics processor to a connected display device, the device further including means for electrically connecting the discrete graphics processor to the connected display device after a startup sequence and based on the determination.

[0238] Example 30 may include the device of any one of Examples 27-29, further including means for detecting information from a connected display device, wherein the means for forming includes means for comparing the information of the connected display device with a list of display devices to be driven by a discrete graphics processor, and means for making the determination that the discrete graphics processor will be connected to the connected display device if the information is in the list.

[0239] Example 31 may include the device described in Example 27, wherein the determination is to connect a discrete graphics processor to a connected display device, the device further including means for determining whether the discrete graphics processor is unavailable, and means for overriding the determination when the discrete graphics processor is determined to be unavailable and electrically connecting an integrated graphics processor to the connected display device.

[0240] Example 32 may include the device described in Example 27, and further include means for determining, based on a user’s selection, that another connected display device will be electrically connected to the discrete graphics processor.

[0241] Therefore, the technology described in this article can achieve a better VR experience, where users can read text more easily. In fact, this technology can improve the operation of HMD systems, allowing the entire scene to be presented in a clearer manner.

[0242] The term "coupled" is used herein to refer to any type of direct or indirect relationship between the components under discussion, and may be applied to electrical, mechanical, fluid, optical, electromagnetic, electromechanical, or other connections. Furthermore, the terms "first," "second," etc., are used herein for ease of discussion only and do not carry a specific time or chronological meaning, unless otherwise stated. Additionally, it should be understood that the indefinite article "a" or "an" carries the meaning of "one or more" or "at least one."

[0243] Those skilled in the art will understand from the foregoing description that the extensive techniques of the embodiments can be implemented in various forms. Therefore, although various embodiments have been described in conjunction with specific examples, the true scope of the embodiments is not limited thereto, as other modifications will become apparent to those skilled in the art upon studying the drawings, specification, and the following claims.

Claims

1. A performance enhanced computing system comprising: an integrated graphics processor; and logic to: form a determination whether a discrete graphics processor or the integrated graphics processor is to be connected to a connected display device based on information from the connected display device, the information corresponding to whether the connected display device is to be driven by the integrated graphics processor or the discrete graphics processor, wherein the determination is made during a boot sequence of the computing system, in response to the determination being that the discrete graphics processor is to be connected to the connected display device, defer enumeration of the connected display device until the discrete graphics processor is available after the boot sequence, and in response to the determination being that the integrated graphics processor is to be connected to the connected display device, enumerate the connected display device during the boot sequence. the computing system includes the discrete graphics processor.

2. The system of claim 1, wherein, further comprising a multiplexer (MUX) electrically connected to the integrated graphics processor, the discrete graphics processor, and the connected display device, 3. The system of claim 2, wherein, wherein the determination is to connect the discrete graphics processor to the connected display device, and further wherein, after the boot sequence and based on the determination, the logic is to control the MUX to electrically connect the discrete graphics processor to the connected display device. the logic includes a list of display devices to be driven by the discrete graphics processor, 4. The system of any one of claims 1-3, wherein, further wherein the logic is to detect the information from the connected display device, compare the information to the list, and make the determination that the discrete graphics processor is to be connected to the connected display device if the comparison indicates that the information is in the list. the determination is to connect the discrete graphics processor to the connected display device, 5. The system of claim 1, wherein, wherein the logic is to override the determination when it is determined that the discrete graphics processor is not available, and electrically connect the integrated graphics processor to the connected display device. the logic is to determine, based on a user's selection, that another connected display device is to be electrically connected to the discrete graphics processor.

6. The system of claim 1, wherein, further comprising a substrate, the logic being coupled to the substrate.

7. The system of claim 1, wherein, 8. A semiconductor package device comprising: logic implemented in one or more of configurable logic or fixed function hardware logic, the logic to: form a determination whether a discrete graphics processor or an integrated graphics processor is to be connected to a connected display device based on information from the connected display device, the information corresponding to whether the connected display device is to be driven by the integrated graphics processor or the discrete graphics processor, wherein the determination is made during a boot sequence of a computing system, in response to the determination being that the discrete graphics processor is to be connected to the connected display device, defer enumeration of the connected display device until the discrete graphics processor is available after the boot sequence, and in response to the determination being that the integrated graphics processor is to be connected to the connected display device, enumerate the connected display device during the boot sequence. in response to the determination being that the integrated graphics processor is to be connected to the connected display device, enumerating the connected display device during the boot sequence.

9. The apparatus of claim 8, wherein, the computing system includes the integrated graphics processor and the discrete graphics processor.

10. The apparatus of claim 9, wherein, the determination is that the discrete graphics processor is to be connected to the connected display device, and after the boot sequence and based on the determination, the logic is to control a multiplexer electrically connected to the discrete graphics processor, the integrated graphics processor, and the connected display device to electrically connect the discrete graphics processor to the connected display device.

11. The apparatus of any one of claims 8-10, wherein, the logic includes a list of display devices to be driven by the discrete graphics processor, and wherein the logic is to: detect the information from the connected display device, compare the information to the list, and make the determination that the discrete graphics processor is to be connected to the connected display device if the comparison indicates that the information is in the list.

12. The apparatus of claim 8, wherein, the determination is that the discrete graphics processor is to be connected to the connected display device, and wherein the logic is to override the determination when the discrete graphics processor is determined to be unavailable and electrically connect the integrated graphics processor to the connected display device.

13. The apparatus of claim 8, wherein, the logic is to determine that another connected display device is to be electrically connected to the discrete graphics processor based on a selection by a user.

14. The apparatus of claim 8, wherein, further comprising a substrate, the logic being coupled to the substrate.

15. A method of operating a semiconductor package device, comprising: forming a determination of whether a discrete graphics processor or an integrated graphics processor is to be connected to a connected display device based on information from the connected display device, the information corresponding to whether the connected display device is to be driven by the integrated graphics processor or the discrete graphics processor, wherein the determination is made during a boot sequence of a computing system; in response to the determination being that the discrete graphics processor is to be connected to the connected display device, deferring enumeration of the connected display device until after the boot sequence the discrete graphics processor is available, and in response to the determination being that the integrated graphics processor is to be connected to the connected display device, enumerating the connected display device during the boot sequence.

16. The method of claim 15, wherein, the computing system includes the integrated graphics processor and the discrete graphics processor.

17. The method of claim 16, wherein, the determination is that the discrete graphics processor is to be connected to the connected display device, the method further includes electrically connecting the discrete graphics processor to the connected display device after the boot sequence and based on the determination.

18. The method of any one of claims 15-17, wherein, further comprising: detecting the information from the connected display device, wherein forming the determination includes: comparing the information of the connected display device to a list, the list including display devices to be driven by the discrete graphics processor; and making the determination that the discrete graphics processor is to be connected to the connected display device if the information is in the list.

19. The method of claim 15, wherein, The determination is to connect the discrete graphics processor to the connected display device, The method further comprises: determining whether the discrete graphics processor is unavailable; and overriding the determination when it is determined that the discrete graphics processor is unavailable and electrically connecting the integrated graphics processor to the connected display device.

20. The method of claim 15, wherein, Further comprising determining, based on a user's selection, that another connected display device is to be electrically connected to the discrete graphics processor.

21. A switching device comprising: means for forming a determination, based on information from a connected display device, whether a discrete graphics processor or an integrated graphics processor is to be connected to the connected display device, the information corresponding to whether the connected display device is to be driven by the integrated graphics processor or the discrete graphics processor, wherein the determination is made during a boot sequence of a computing system; means for deferring enumeration of the connected display device until the discrete graphics processor is available after the boot sequence in response to the determination being that the discrete graphics processor is to be connected to the connected display device, and means for enumerating the connected display device during the boot sequence in response to the determination being that the integrated graphics processor is to be connected to the connected display device.

22. The apparatus of claim 21, wherein, The computing system comprises the integrated graphics processor and the discrete graphics processor.

23. The apparatus of claim 22, wherein, The determination is to connect the discrete graphics processor to the connected display device, The device further comprises means for electrically connecting the discrete graphics processor to the connected display device after the boot sequence and based on the determination.

24. The apparatus of any one of claims 21-23, wherein, Further comprising: means for detecting the information from the connected display device; and The means for forming comprises: means for comparing the information of the connected display device to a list, the list comprising display devices to be driven by the discrete graphics processor; and means for making the determination that the discrete graphics processor is to be connected to the connected display device if the information is in the list.

25. The apparatus of claim 21, wherein, The determination is to connect the discrete graphics processor to the connected display device, The device further comprises: means for determining whether the discrete graphics processor is unavailable; and means for overriding the determination when it is determined that the discrete graphics processor is unavailable and electrically connecting the integrated graphics processor to the connected display device.

26. The apparatus of claim 21, wherein, Further comprising means for determining, based on a user's selection, that another connected display device is to be electrically connected to the discrete graphics processor.

Citation Information

Patent Citations

  • Switchable hybrid graphics

    US10224003B1

  • Multiplexed graphics architecture for graphics power management

    US20080034238A1

  • Display controller, information processing device and display method

    US20120081374A1