Adaptive Foveal Encoder and Global Motion Prediction Factor

Through the adaptive foveal encoder and global motion prediction factor, the focus and motion information of the head-mounted device are used to optimize video encoding, which solves the problems of insufficient encoding efficiency and quality in the prior art, and achieves more efficient encoding and better user experience.

CN111164654BActive Publication Date: 2025-07-18INTEL CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201780095426.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-11-23
Publication Date
2025-07-18
Estimated Expiration
2037-11-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize user focus and motion information in the graphics system to optimize video encoding, resulting in insufficient encoding efficiency and quality.

Method used

Adaptive foveal encoder and global motion prediction factor are used to dynamically adjust video encoding parameters based on the focus and motion information of the head-mounted device, especially to improve encoding quality in the focus area, and optimize macroblock encoding through global motion prediction factor.

Benefits of technology

Improve the efficiency and quality of video encoding, especially under limited bandwidth conditions, enhance user experience, and reduce computing load and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111164654B_ABST
    Figure CN111164654B_ABST
Patent Text Reader

Abstract

Embodiments of an adaptive video encoder may include the following techniques: determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focus and information related to motion; and determining one or more video encoding parameters based on the information related to the head-mounted device. Other embodiments are disclosed and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The various embodiments generally relate to graphics systems. More specifically, the embodiments relate to an adaptive foveal encoder and a global motion predictor. Background Art

[0002] After an image is rendered by a graphics engine, the image can be encoded for display, transmission, and / or file storage. Fovea may refer to a small depression in the retina of the eye where visual acuity may be the highest. The center of the field of view can be focused on this area where cones may be particularly concentrated. In the context of certain graphics applications, fovea or the foveal region can correspond to the focused area in an image or a display. Brief Description of the Drawings

[0003] The various advantages of the embodiments will become apparent to those skilled in the art by reading the following specification and the appended claims and by referring to the following drawings, in which:

[0004] Figure 1 is a block diagram of an example of an electronic processing system according to an embodiment;

[0005] Figure 2 is a block diagram of an example of a sensing engine according to an embodiment;

[0006] Figure 3 is a block diagram of an example of a focusing engine according to an embodiment;

[0007] Figure 4 is a block diagram of an example of a motion engine according to an embodiment;

[0008] Figure 5 is a block diagram of another example of an electronic processing system according to an embodiment;

[0009] Figure 6 is a block diagram of an example of a semiconductor packaging device according to an embodiment;

[0010] Figures 7A to 7C is a flowchart of another example of a method of adaptive encoding according to an embodiment;

[0011] Figure 8 is a block diagram of another example of an electronic processing system according to an embodiment;

[0012] Figure 9 is a block diagram of an example of an adaptive encoder according to an embodiment;

[0013] Figure 10 is a flowchart of an example of a method of adaptive encoding according to an embodiment;

[0014] Figure 11 is a block diagram of an example of a motion vector calculator assisted by head movement according to an embodiment;

[0015] Figure 12 is a view of an example of a VR projection model according to an embodiment;

[0016] Figure 13 is a block diagram of another example of an adaptive encoder according to an embodiment;

[0017] Figure 14 is a flowchart of another example of a method of adaptive coding according to an embodiment;

[0018] Figures 15A to 15F is a view of an example of a set of regions according to an embodiment;

[0019] Figures 15G to 15H is a view of an example of foveal coding according to an embodiment;

[0020] Figure 16 is a block diagram of an example of a stereo virtual reality display according to an embodiment;

[0021] Figure 17 is a view of an example of a grid of macroblocks superimposed on a foveal region according to an embodiment;

[0022] Figure 18 is a block diagram of another example of an electronic processing system according to an embodiment;

[0023] Figure 19 is a flowchart of another example of a method of adaptive coding according to an embodiment;

[0024] Figure 20 is a block diagram of an example of a processing system according to an embodiment;

[0025] Figure 21 is a block diagram of an example of a processor according to an embodiment;

[0026] Figure 22 is a block diagram of an example of a graphics processor according to an embodiment;

[0027] Figure 23 is a block diagram of an example of a graphics processing engine of a graphics processor according to an embodiment;

[0028] Figure 24 is a block diagram of an example of the hardware logic of a graphics processor core according to an embodiment;

[0029] Figures 25A to 25B shows an example of thread execution logic according to an embodiment;

[0030] Figure 26is a block diagram showing an example of a graphics processor instruction format according to an embodiment;

[0031] Figure 27 is a block diagram of another example of a graphics processor according to an embodiment;

[0032] Figure 28A is a block diagram showing an example of a graphics processor command format according to an embodiment;

[0033] Figure 28B is a block diagram showing an example of a graphics processor command sequence according to an embodiment;

[0034] Figure 29 shows an example of an exemplary graphics software architecture for a data processing system according to an embodiment;

[0035] Figure 30A is a block diagram showing an example of an IP core development system according to an embodiment;

[0036] Figure 30B shows an example of a cross-sectional side view of an integrated circuit package assembly according to an embodiment;

[0037] Figure 31 is a block diagram showing an example of a system-on-chip integrated circuit according to an embodiment;

[0038] Figures 32A to 32B is a block diagram showing an exemplary graphics processor for use within a SoC according to an embodiment; and

[0039] Figures 33A to 33B shows additional exemplary graphics processor logic according to an embodiment. DETAILED DESCRIPTION

[0040] Turning now to Figure 1 , an embodiment of an electronic processing system 10 may include an application processor 11, a permanent storage medium 12 communicatively coupled to the application processor 11, and a graphics subsystem 13 communicatively coupled to the application processor 11. The system 10 may further include: a sensing engine 14 communicatively coupled to the graphics subsystem 13 to provide sensing information; a focusing engine 15 communicatively coupled to the sensing engine 14 and the graphics subsystem 13 to provide focusing information; a motion engine 16 communicatively coupled to the sensing engine 14, the focusing engine 15, and the graphics subsystem 13 to provide motion information; and an adaptive encoder 17 communicatively coupled to the motion engine 16, the focusing engine 15, and the sensing engine 14 to adjust one or more video encoding parameters of the graphics subsystem 13 based on one or more of the sensing information, the focusing information, and the motion information.

[0041] In some embodiments of system 10, the adaptive encoder 17 may further include an adaptive foveated encoder to encode an image based on the focusing information (e.g., as described in more detail below). Some embodiments of the adaptive encoder 17 may further include an adaptive motion encoder to determine global motion parameters for the encoder based on motion information (e.g., as described in more detail below). For example, the adaptive encoder 17 may be configured to determine information related to a headset, the information including at least one of information related to focusing and information related to motion, and determine one or more video coding parameters based on the information related to the headset (e.g., as described in more detail below). In some embodiments, the adaptive encoder 17 may also be configured to encode macroblocks of a video image based on one or more determined video coding parameters.

[0042] Each of the above-described application processor 11, permanent storage medium 12, graphics subsystem 13, sensing engine 14, focusing engine 15, motion engine 16, adaptive encoder 17, and other system components may be implemented in hardware, software, or any suitable combination thereof. For example, a hardware implementation may include configurable logic such as, for example, a programmable logic array (PLA), FPGA, complex programmable logic device (CPLD), or fixed-function logic hardware using circuit technologies such as, for example, ASIC, complementary metal oxide semiconductor (CMOS), or transistor-transistor logic (TTL) technology, or any combination thereof. Alternatively or additionally, these components may be implemented as a set of logic instructions stored in a machine or computer-readable storage medium (such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc.) to be executed by a processor or computing device in one or more modules. For example, computer program code for implementing the operation of the components can be written in any combination of one or more operating system-applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0043] Sensing Engine Example

[0044] Now turning to Figure 2 , the sensing engine 18 may obtain information from sensors, content, services, and / or other sources to provide sensing information. The sensed information may include, for example, image information, audio information, motion information, depth information, temperature information, biometric information, graphics processing unit (GPU) information, etc. Generally speaking, some embodiments may use the sensed information to adjust the video coding parameters of the graphics system.

[0045] For example, the sensing engine may include a sensor hub communicatively coupled to a two-dimensional (2D) camera, a three-dimensional (3D) camera, a depth camera, a gyroscope, an accelerometer, an inertial measurement unit (IMU), first- and second-order motion meters, a position service, a microphone, a proximity sensor, a thermometer, a biometric sensor, etc., and / or a combination of multiple sources that provide information to the focusing and / or motion engine. The sensor hub may be distributed across multiple devices. Information from the sensor hub may include or be combined with input data (e.g., touch data) from the user device.

[0046] For example, the user's device(s) may include one or more 2D, 3D, and / or depth cameras. The user's device(s) may also include a gyroscope, an accelerometer, an IMU, a positioning service, a thermometer, a biometric sensor, etc. For example, the user may wear a head-mounted display (HMD) that includes various cameras, motion sensors, and / or other sensors. Non-limiting examples of mixed reality HMDs include the Microsoft HoloLens. The user may also carry a smartphone (e.g., in the user's pocket) and / or may wear a wearable device (e.g., such as a smartwatch, an activity monitor, and / or a health tracker). The user's device(s) may also include a microphone that can be used to detect whether the user is speaking, on the phone, speaking to another person nearby, etc. The sensor hub may include some or all of the user's various devices capable of capturing information related to the user's actions or activities (e.g., including the input / output (I / O) interface of the user device that can capture keyboard / mouse / touch activities). The sensor hub may obtain information directly from the capture devices of the user device (e.g., wired or wirelessly), or the sensor hub is capable of integrating information from devices of a server or service (e.g., information from a health tracker can be uploaded to a cloud service and the sensor hub can download the information).

[0047] Examples of focusing engines

[0048] Now turning to Figure 3, the focusing engine 19 can obtain information from the sensing engine and / or the motion engine and other sources to provide focusing information. The focusing information can include, for example, focus, focusing area, eye position, eye movement, pupil size, pupil dilation, depth of field (DOF), content focus, content focusing object, content focusing area, etc. The focusing information can also include previous focusing information, determined future focusing information, and / or predicted focusing information (e.g., predicted focus, predicted focusing area, predicted eye position, predicted eye movement, predicted pupil size, predicted pupil dilation, predicted DOF, determined future content focus, determined future content focusing object, determined future content focusing area, predicted content focus, predicted content focusing object, predicted content focusing area, etc.).

[0049] Macroscopically, some embodiments can use the focusing information to adjust the video encoding parameters of the graphics system based on 1) assuming where the user is looking, 2) determining where the user is looking, 3) applying where the user is desired to look, and / or 4) predicting where the user will look in the future. In the focusing area where the user is viewing, certain focusing cues may be stronger. If the user looks straight ahead, they may see things with clear focus. In the case where the scene or object faces the periphery, the user may notice the movement but not the clear details of the focus.

[0050] For example, if there is limited sensing information or processing capacity of the graphics system (e.g., the attached HMD or the host cannot provide or utilize the information), the focusing information can be static and / or assumption-based (e.g., it can be assumed that the user is viewing the center of the screen with a fixed eye position, DOF, etc.). The focusing information can also change dynamically based on factors such as motion information (e.g., from a virtual reality (VR) headset), motion prediction information, content information (e.g., motion in the scene), etc. More preferably, a rich sensor set including eye tracking (e.g., sometimes also referred to as gaze tracking) can be utilized to provide a better user experience to identify the focusing area and provide the focusing information. For example, some embodiments can include an eye tracker or obtain eye information from an eye tracker to track the user's eyes. The eye information can include eye position, eye movement, pupil size / dilation, depth of field, etc. The eye tracker can capture an image of the user's eyes including the pupils. The user's focus and / or DOF can be determined, inferred, and / or estimated based on the eye position and pupil dilation. The user can perform a calibration process, which can help the eye tracker provide more accurate focusing and / or DOF information.

[0051] For example, when a user wears a VR headset, a camera can capture an image of the pupil, and the system can determine where the user is looking (e.g., the focus area, depth, and / or direction). The camera can capture pupil dilation information, and the system can infer where the user's focus area is based on this information. For example, the human eye has a certain depth of field (DOF) such that if a person focuses on something nearby, things in the distance may become blurry. Focus information can include the focus at focal length X and the DOF information of Δ(X), so the focus area can correspond to X+ / -Δ[X] located near the user's focus. The size of the DOF can vary with the distance X (e.g., having different Δs at different focal lengths). For example, the user's DOF can be calibrated and can vary in each direction (e.g., x, y, and z), such that the Δ[X] function is not necessarily spherical.

[0052] In some embodiments, the focus information can include content-based focus information. For example, in 3D, VR, augmented reality (AR), and / or mixed reality environments, depth and / or distance information can be provided from an application (e.g., where the user is in the virtual environment, where an object is, and / or how far an object is from the user, etc.). Content-based focus information may also include points, objects, or regions in the content that the application wants the user to focus on, such as the more interesting things happening that the application wants the user to pay attention to. The application may also be able to provide future content focus information because the application may know the motion information of the content and / or what objects / regions may be more interesting to the user in the next frame or next scene (e.g., an object about to enter the scene from the edge of the screen).

[0053] Motion Engine Example

[0054] Now turning to Figure 4 , the motion engine 20 can obtain information from the sensing engine and / or the focus engine and other sources to provide motion information. The motion information can include, for example, head position, head velocity, head acceleration, head movement direction, eye velocity, eye acceleration, eye movement direction, object position, object velocity, object acceleration, object movement direction, etc. The motion information can also include previous motion information, determined future motion information, and / or predicted motion information (e.g., predicted head velocity, predicted head acceleration, predicted head position, predicted head movement direction, predicted eye velocity, predicted eye acceleration, predicted eye movement direction, determined future content position, determined future content object velocity, determined future content object acceleration, predicted object position, predicted object velocity, predicted object acceleration, etc.).

[0055] At a macroscopic level, some embodiments may use motion information to adjust video encoding parameters of a graphics system based on 1) the user moving their head, 2) the user moving their eyes, 3) the user moving their body, 4) where the application wants the user to turn their head, eyes, and / or body, and / or 4) where it is predicted the user will turn their head, eyes, and / or body. Some motion information can be easily determined from sensed information. For example, head position, velocity, acceleration, direction of motion, etc. can be determined from an accelerometer. Eye motion information can be determined by tracking eye position information over time (e.g., if an eye tracker only provides eye position information).

[0056] Some motion information can be content-based. For example, in a game or on-the-fly 3D content, the application may know how fast and where an object is moving. The application can provide this information to the motion engine (e.g., via an API call). Future content-based object motion information for the next frame / scene can also be fed into the motion engine for decision-making. Some content-based motion information can be determined by performing image processing or machine vision processing on the content.

[0057] Some embodiments of a machine vision system can, for example, analyze features captured by a camera and / or perform feature / object recognition on an image captured by a camera. For example, machine vision and / or image processing can identify and / or recognize objects in a scene (e.g., the edge belongs to the front of a chair). The machine vision system can also be configured to perform face recognition, gaze tracking, facial expression recognition, and / or pose recognition, which includes body-level poses, arm / leg-level poses, hand-level poses, and / or finger-level poses. The machine vision system can be configured to classify the actions of a user. In some embodiments, a properly configured machine vision system can be capable of determining whether a user is present at a computer, whether typing on a keyboard, whether using a mouse, whether using a touchpad, whether using a touchscreen, whether using an HMD, whether using a VR system, whether sitting, whether standing, and / or whether taking other actions or activities.

[0058] For example, a motion engine can obtain camera data related to real objects in a scene and can use this information to identify the motion and direction of the real objects. The motion engine can obtain latency information from a graphics processor. The motion engine can then predict the next frame direction of the real object. The amount of latency can be based on one or more of the time used to render and / or encode the scene, the number of virtual objects in the scene, and the complexity of the scene, etc. For example, a sensing engine can include one or more cameras to capture a real scene. For example, one or more cameras can include one or more 2D cameras, 3D cameras, depth cameras, high-speed cameras, or other image capture devices. The real scene can include objects moving in the scene. The cameras can be coupled to an image processor to process data from the cameras to identify objects in the scene (e.g., including moving objects) and to identify the motion of the objects (e.g., including direction information). The motion engine can determine predicted motion information based on tracking the motion of the object and can predict the future position of the object based on the measured or estimated latency (e.g., from the time of capture to the time of rendering / encoding). According to some embodiments, various motion tracking and / or motion prediction techniques can be enhanced using optical flow and other real motion estimation techniques to determine the next position of the real object. For example, some embodiments can use extended common filtering and / or perspective processing (e.g., from autonomous driving applications) to predict the motion of the object.

[0059] Engine overlap example

[0060] Those skilled in the art will understand that aspects of the various engines described herein can overlap with other engines, and portions of each engine can be implemented or distributed throughout various parts of an electronic processing system. For example, a focusing engine can use motion information to provide a predicted future focusing area, and a motion engine can use focusing information to predict future motion. Eye motion information can come directly from the sensing engine, can be determined / predicted by the focusing engine, and / or can be determined / predicted by the motion engine. The examples herein should be considered illustrative and not limiting in terms of specific implementations.

[0061] Adaptive foveal encoder and global motion predictor example

[0062] Now turning to Figure 5, an embodiment of the electronic processing system 22 may include: a processor 23; a memory 24 communicatively coupled to the processor 23; and logic 25 communicatively coupled to the processor 23 for determining information related to the head-mounted device (including at least one of information related to focusing and information related to motion), and determining one or more video coding parameters based on the information related to the head-mounted device. In some embodiments, the logic 25 may further be configured to adjust quality parameters for encoding of macroblocks based on the information related to focusing. For example, the logic 25 may be configured to identify a focused area based on the information related to focusing, and adjust the quality parameters for encoding such that, compared to macroblocks outside the focused area, relatively higher quality is provided for macroblocks inside the focused area. Additionally or alternatively, in some embodiments, the logic 25 may further be configured to determine a global motion predictor for encoding of macroblocks based on the information related to motion. For example, the logic 25 may additionally or alternatively be configured to determine a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion. For any of the embodiments, the logic 25 may further be configured to encode macroblocks of a video image based on one or more determined video coding parameters. In some embodiments, any one of the processor 23, the memory 24, and the logic 25 may be fully or partially co-located on the same integrated circuit die.

[0063] Embodiments of each of the above processor 23, memory 24, and logic 25, as well as other system components, can be implemented in hardware, software, or any suitable combination thereof. For example, a hardware implementation may include configurable logic such as, for example, a programmable logic array (PLA), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), or fixed function logic hardware using circuit technologies such as, for example, an application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS), or transistor-transistor logic (TTL) technology, or any combination thereof.

[0064] Alternatively or additionally, all or part of these components may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), firmware, flash memory, etc.) to be executed by a processor or computing device. For example, the computer program code for implementing the operation of the components can be written in any combination of one or more operating system (OS)-applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages. For example, the memory 24, permanent storage medium, or other system memory may store a set of instructions that, when executed by the processor 23, cause the system 22 to implement one or more components, features, or aspects of the system 22 (e.g., logic 25, determining information related to the head-mounted device including at least one of information related to focusing and information related to motion, determining one or more video coding parameters based on the information related to the head-mounted device, etc.).

[0065] Turning now to Figure 6 , an embodiment of the semiconductor packaging device 27 may include one or more substrates 28 and logic 29 coupled to the one or more substrates 28, wherein the logic is implemented at least in part by one or more of configurable logic and fixed-function hardware logic, and the logic 29 is coupled to the one or more substrates 28 for determining information related to the head-mounted device including at least one of information related to focusing and information related to motion, and determining one or more video coding parameters based on the information related to the head-mounted device. In some embodiments, the logic 29 may further be configured to adjust quality parameters for the encoding of macroblocks based on the information related to focusing. For example, the logic 29 may be configured to identify a focused area based on the information related to focusing, and adjust the quality parameters for encoding to provide a relatively higher quality for the macroblocks inside the focused area compared to the macroblocks outside the focused area. Additionally or alternatively, in some embodiments, the logic 29 may further be configured to determine a global motion predictor for the encoding of macroblocks based on the information related to motion. For example, the logic 29 may additionally or alternatively be configured to determine a hierarchical motion estimation offset based on the current head position, previous head position, and center point from the information related to motion. For any of the embodiments, the logic 29 may further be configured to encode macroblocks of a video image based on one or more determined video coding parameters.

[0066] Embodiments of the logic 29, as well as other components of the apparatus 27, can be implemented in hardware, software, or any combination thereof (including at least a partial hardware implementation). For example, a hardware implementation may include configurable logic such as, for example, a PLA, FPGA, CPLD, or fixed function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL technologies, or any combination thereof. Additionally, portions of these components can be implemented in one or more modules as a set of logic instructions stored in a machine or computer-readable storage medium (such as, RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for implementing the operation of the components can be written in any combination of one or more OS-applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0067] Now turning to Figures 7A to 7C , embodiments of the adaptive coding method 30 may include: determining, at block 31, information related to the head-mounted device including at least one of information related to focus and information related to motion; and determining, at block 32, one or more video coding parameters based on the information related to the head-mounted device. In some embodiments, the method 30 may additionally or alternatively include adjusting, at block 33, a quality parameter for encoding of macroblocks based on the information related to focus. For example, the method 30 may include: identifying, at block 34, a focus region based on the information related to focus; and adjusting, at block 35, the quality parameter for encoding such that, compared to macroblocks outside the focus region, relatively higher quality is provided for macroblocks inside the focus region. Additionally or alternatively, some embodiments of the method 30 may include determining, at block 36, a global motion predictor for encoding of macroblocks based on the information related to motion. For example, the method 30 may include determining, at block 37, a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion. For any of the embodiments, the method 30 may further include encoding, at block 38, macroblocks of a video image based on one or more determined video coding parameters.

[0068] Embodiments of method 30 may be implemented in a system, apparatus, computer, device, etc. (such as, for example, those described herein). More specifically, a hardware implementation of method 30 may include configurable logic (such as, for example, PLA, FPGA, CPLD), or fixed-function logic hardware using circuit technologies (such as, for example, ASIC, CMOS or TTL technologies), or any combination thereof. Alternatively or additionally, method 30 may be implemented as a set of logical instructions stored in a machine or computer-readable storage medium (such as, RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for implementing the operations of components can be written in any combination of one or more OS-applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0069] For example, method 30 may be implemented on a computer-readable medium described in conjunction with Examples 19 to 24 below. Embodiments or portions of method 30 may be implemented in firmware, an application (e.g., via an application programming interface (API)), or driver software running on an operating system (OS).

[0070] Now turning to Figure 8 , embodiments of an electronic processing system 40 may include a host device 41 (such as, for example, a personal computer, a server, etc.) communicatively coupled to a user device 42 (such as, for example, a smart phone, a tablet, an HMD, etc.) in a wired or wireless manner. The host device 41 may include a VR renderer 43 (such as, for example, a high-performance graphics adapter) communicatively coupled to an adaptive encoder 44, which in turn may be communicatively coupled to a transceiver 45 (such as, for example, WIFI, WIGIG, Ethernet, etc.). For example, the adaptive encoder 44 may utilize Advanced Video Coding (AVC) and / or High Efficiency Video Coding (HEVC), and may be executed by a graphics processor or an external graphics adapter. The user device 42 may include a transceiver 46 (such as, for example, configured to exchange information with the transceiver 45 of the host device 41) communicatively coupled to a VR decoder 47, which in turn is communicatively coupled to a display 48. The user device 42 may further include one or more sensors 49.

[0071] The host device 41 may include one or more GPUs that implement all or a portion of the VR renderer 43 and the adaptive encoder 44. The host device 41 may render VR graphics / video image content, encode the content, and stream the VR content to the user device 42. The user device 42 may decode the VR content and present the graphics / video image on the display 48. The user device may also support other local functions such as asynchronous time warping (ATW), frame buffer rendering (FBR), barrel distortion correction, etc. According to some embodiments, the user device 42 may provide sensor-related information back to the host device 41 for beneficial use by the adaptive encoder 44 as described herein. For example, the user device 42 may send HMD position, 3 degrees of freedom (3DOF) information, 6 degrees of freedom (6DOF) information, and / or other sensor data back to the host device 41. For example, the adaptive encoder 44 may advantageously include one or both of an adaptive motion encoder and / or an adaptive foveal encoder to adjust one or more video encoding parameters based on sensor-related information received from the user device 42.

[0072] Adaptive Motion Encoder Example

[0073] Turning now to Figure 9 , an embodiment of the adaptive motion encoder device 50 may include: a motion engine 51 for providing motion information; and an adaptive motion encoder 52 communicatively coupled to the motion engine 51 to adjust one or more video encoding parameters based on the motion information. In some embodiments, the motion engine 51 may include a head tracker 53 to identify head position / motion (e.g., or obtain head position / motion information from an attached HMD). For example, the motion engine 51 may obtain motion-related information, including head position, from a head-mounted device (e.g., a VR HMD worn by a user). For example, the adaptive motion encoder 52 may be configured to determine global motion parameters based on the motion information. For example, one or more video encoding parameters may include motion vectors based on global motion predictors, and the adaptive motion encoder 52 may determine global motion predictors based on head position information from the motion engine 51. In some embodiments, the adaptive motion encoder 52 may adjust the motion vectors used to encode macroblocks based on the motion-related information. The adaptive motion encoder 52 may also be configured to encode macroblocks of a video image based on one or more determined video encoding parameters (e.g., including global motion predictor values determined according to head position / motion information).

[0074] Embodiments of each of the above-described motion engine 51, adaptive motion encoder 52, head tracker 53, and other components of device 50 may be implemented in hardware, software, or any combination thereof. For example, a part or all of device 50 may be implemented as part of a parallel GPU, further configured with the adaptive motion encoder described herein. Device 50 may also be adapted to work with a stereo HMD system. For example, a hardware implementation may include configurable logic such as, for example, PLA, FPGA, CPLD, or fixed function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL technology, or any combination thereof. Alternatively or additionally, these components may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium such as, for example, RAM, ROM, PROM, firmware, flash memory, etc., to be executed by a processor or computing device. For example, computer program code for implementing the operation of the components can be written in any combination of one or more operating system-applicable / suitable programming languages, including object-oriented programming languages such as, for example, PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0075] Now turning to Figure 10 , embodiments of the adaptive encoding method 60 may include: determining motion information at block 61, and adjusting one or more video encoding parameters based on the motion information at block 62. For example, at block 63, one or more video encoding parameters may include global motion parameters such as a global motion prediction factor value. Method 60 may further include determining head position information based on the motion information at block 64, and determining a global motion prediction factor value based on the head position information at block 65. In some embodiments, the method may further include: determining motion vector information for macroblocks of a video image based on the global motion prediction factor at block 66; and encoding the macroblocks of the video image based on one or more determined video encoding parameters at block 67.

[0076] Embodiments of method 60 may be implemented in a system, apparatus, GPU, PPU, or media processor. More specifically, a hardware implementation of method 60 may include configurable logic (such as, for example, PLA, FPGA, CPLD), or fixed function logic hardware using circuit technologies (such as, for example, ASIC, CMOS, or TTL technology), or any combination thereof. Alternatively or additionally, method 60 may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as, RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for implementing component operations can be written in any combination of one or more operating system applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0077] For example, embodiments or portions of method 60 may be implemented as firmware, an application (e.g., via an API), or driver software. Other embodiments or portions of method 60 may be implemented with specialized code (e.g., shaders) to be executed on a GPU. Other embodiments or portions of method 60 may be implemented in fixed function logic or specialized hardware (e.g., in a GPU or media processor).

[0078] Some embodiments may advantageously utilize the HMD position as a global motion prediction factor to improve the encoding efficiency of VR streaming. Non-limiting examples of applications for some embodiments may include AR, VR, computer-generated holography, VR gaming, automotive infotainment, and automotive driver assistance systems. In some embodiments, compared to an all-in-one system, a wireless VR streaming system (e.g., such as the system 40 in Figure 8 can improve the user experience. Specifically, a higher computational workload may be handled by the host device, while the battery-powered user device may only need to receive, decode, and display the image data (thus reducing weight and extending the battery life of the user device). Providing a wireless connection advantageously decouples the user from the host device and provides the user with a greater degree of freedom of movement.

[0079] For wireless VR streaming, video encoding may play an important role in generating high-quality content within the target bitrate. For example, a user may enjoy playing a VR game while wearing an HMD, and the scene may change according to the user's head movement. To handle large motions, hierarchical motion estimation (HME) can be used to predict the motion. Some other systems can utilize 32x motion estimation (32xME) (also known as ultra HME (UHME)) and 16x motion estimation (16xME) (also known as super HME (SHME)) to reduce the size of the original surface and the reference frame and estimate the global motion vector. The video encoding pipeline can provide the output of the 32xME calculation to the 16xME calculation, and the 16xME calculation can in turn be provided to the 4x motion estimation (4xME) calculation to predict the motion vector for encoding. Both the 32xME and 16xME calculations can be highly GPU compute-intensive. Some embodiments can estimate the global motion vector based on the head position information from the HMD to advantageously utilize much simpler and faster calculations in place of the 32xME and 16xME calculations. Advantageously, some embodiments can obtain the global motion before encoding through the change in HMD position and can utilize the HMD position information to calculate the global motion predictor to improve the encoding quality while also reducing the computational load.

[0080] In VR usage, the game scene may not change very rapidly, but the user's head position may change rapidly (e.g., in a VR game, the user may quickly turn their head). The VR software development kit (SDK) can obtain the matching head position and the rendered VR content. For example, 3DOF and / or 6DOF information is available to the VR SDK. Thus, some embodiments can use the head position to calculate the global motion predictor and guide the encoder to generate the best-matching motion vectors for all macroblocks of the video image to be encoded.

[0081] Now turning to Figure 11, the motion vector calculator 70 assisted by head movement may include an HME offset module 71 communicatively coupled to the 4xME module 72, and the 4xME module 72 may in turn be communicatively coupled to the macroblock encoding (MBEnc) module 73. The HME offset module 71 may be provided with the current head position information, the previous head position information, and the center point information, and may calculate the 4xME offset based on the provided information. For example, the HME offset module 71 may calculate the 4xHME offset based on the current and previous HMD positions and the center point 2D position (0, 0, w / 2, h / 2). The 4xHME offset may correspond to an offset for searching around this predictor rather than around the co-located (0, 0) predictor. In some embodiments, the HME offset calculation may be advantageously performed by a central processing unit (CPU) on the host device without processing on the original surface and the reference surface.

[0082] Now turning to Figure 12 , the illustration of the VR projection model 75 may include a user 76 wearing a VR HMD 77. The user 76 may change their head position from the previous head position (PrevHeadPos) to the current head position (CurHeadPos), such that the points in the VR scene move from point P(old) to point P(new) in the screen space. In addition to the X, Y, and Z dimensions, the projection geometry may also have an additional dimension, which is called W. This four-dimensional space may be called the projection space, and the coordinates in the projection space may be called homogeneous coordinates. The homogeneous coordinates P(x, y, z, w) may correspond to the points P(old) and P(new) in the projection space. The following equations explain an example of how to calculate the HME offset based on the HMD position by projecting the points in the 3D scene onto the 2D surface.

[0083] MatrixHeadPos4x4 = quatToMatrix(head.w, head.x, head.y, head.z) [Equation 1]

[0084] The function quatToMatrix(qw, qx, qy, qz) may convert the quaternion into a rotation matrix as follows:

[0085] A = [1.0 - 2.0 * qy * qy - 2.0 * qz * qz, 2.0 * qx * qy - 2.0 * qz * qw,... 2.0 * qx * qz + 2.0 * qy * qw, 0.0; 2.0 * qx * qy + 2.0 * qz * qw, 1.0 - 2.0 * qx * qx - 2.0 * qz * qz,... 2.0 * qy * qz - 2.0 * qx * qw, 0.0; 2.0 * qx * qz - 2.0 * qy * qw, 2.0 * qy * qz + 2.0 * qx * qw,... 1.0 - 2.0 * qx * qx - 2.0 * qy * qy, 0.0; 0.0, 0.0, 0.0, 1.0];

[0086] end

[0087] Point2D(x, y) = Point3D(x, y, z, w) * MatrixHeadPos4x4 [Equation 2]

[0088] where the head (w, x, y, z) is the current HMD position, Point3D(x, y, z, w) is the pixel in 3D world coordinates, and Point2D(x, y) is the pixel in the final 2D image.

[0089] As Figure 12 shown, triangle 78 corresponds to the previous head position, and triangle 79 corresponds to the latest head position. Two corresponding 2D points P(new) and P(old) can be determined based on Equation 1 and Equation 2 as follows:

[0090] Point2Dnew(x, y) = Point3D(x, y, z, w) * MatrixHeadPosNew4x4 [Equation 3]

[0091] Point2Dold(x, y) = Point3D(x, y, z, w) * MatrixHeadPosOld4x4 [Equation 4]

[0092] Another equation can be based on Equation 3 and Equation 4:

[0093] Point2Dnew(x, y) = Point2Dold(x, y) *

[0094] Invert(MatrixHeadPosOld4x4) * MatrixHeadPosNew4x4 [Equation 5]

[0095] Where Invert(MatrixHeadPosOld4x4) is used to calculate the inverse matrix of MatrixHeadPosOld4x4. To determine the global motion prediction factor based on the HMD position, the 4xHME offset can be calculated as follows:

[0096] MotionSearch_OFFSET(ΔX,ΔY) =

[0097] Point2Dold(x,y) – Point2Dnew(x,y) [Equation 6]

[0098] Some embodiments can calculate HME_OFFSET as follows:

[0099] MotionSearch_OFFSET(ΔX,ΔY) = Point2Dold(x,y) – Point2Dold(x,y)*Invert(MatrixHeadPosOld4x4)*MatrixHeadPosNew4x4 [Equation 7]

[0100] To simplify the calculation, some embodiments can select the center point of the previous 2D VR content of Point2Dold, i.e., (width / 2, length / 2). Based on Equation 7 (e.g., using the HMD position provided by the HMD), some embodiments can advantageously utilize the HMD position to predict the global motion prediction factor and as a guidance for motion search for the encoder.

[0101] For a wireless VR streaming system, some embodiments can improve the encoding quality with a smaller latency. For example, by replacing the 32XHME and 16XHME calculations in the AVC encoder, some embodiments can save several milliseconds because the number of search points can be significantly smaller. Compared with the encoding using 32xME and 16xME calculations, some embodiments can also provide an encoding quality improvement for various VR applications using adaptive motion encoding as described herein, and this quality improvement can be measured by a higher peak signal-to-noise ratio (PSNR).

[0102] Adaptive foveated encoder example

[0103] Now turn to Figure 13, embodiments of the adaptive foveated encoder device 80 may include: a focusing engine 81 for providing focusing information; and an adaptive foveated encoder 82 communicatively coupled to the focusing engine 81 to adjust one or more video coding parameters based on the focusing information. In some embodiments, the focusing engine 81 may include an eye tracker 83 to identify the foveal region (e.g., or obtain eye position / movement information from the HMD). For example, the focusing engine 81 may obtain information related to focusing, including the foveal region, from a head-mounted device (e.g., such as a VR HMD worn by the user). For example, the adaptive foveated encoder 82 may be configured to provide a higher coding quality for a first coding region inside the foveal region compared to the coding quality of a second coding region outside the foveal region. For example, one or more video coding parameters may include coding quality parameters such as a quantization parameter (QP), and the adaptive foveated encoder 82 may use a lower QP value for the first coding region compared to the QP value of the second coding region. In some embodiments, the adaptive foveated encoder 82 may adjust the QP value used for encoding macroblocks based on information related to focusing. For example, after identifying the foveal region based on information related to focusing, the adaptive foveated encoder 82 may adjust the QP value for encoding such that a relatively higher quality is provided for macroblocks inside the foveal region compared to macroblocks outside the foveal region. The adaptive foveated encoder 82 may also be configured to encode macroblocks of a video image based on one or more determined video coding parameters (e.g., including the focusing-based QP value).

[0104] Embodiments of each of the above-described focusing engine 81, adaptive foveal encoder 82, eye tracker 83, and other components of device 80 can be implemented in hardware, software, or any combination thereof. For example, a part or all of device 80 can be implemented as part of a parallel graphics processing unit (GPU), further configured with the adaptive foveal encoder described herein. Device 80 can also be adapted to work with a stereo HMD system. For example, a hardware implementation can include configurable logic such as, for example, PLA, FPGA, CPLD, or fixed function logic hardware using circuit technologies such as, for example, ASIC, CMOS, or TTL technology, or any combination thereof. Alternatively or additionally, these components can be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium such as, for example, RAM, ROM, PROM, firmware, flash memory, etc., to be executed by a processor or computing device. For example, the computer program code for implementing the operation of the components can be written in any combination of one or more operating system-applicable / suitable programming languages, including object-oriented programming languages such as, for example, PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0105] Now turning to Figure 14 , embodiments of the adaptive encoding method 90 can include: determining focus information at block 91, and adjusting one or more video encoding parameters based on the focus information at block 92. For example, at block 93, one or more video encoding parameters can include quality parameters such as QP values. Method 90 can further include: determining a focus region based on the focus information at block 94; and providing a higher quality encoding for a first encoding region inside the focus region compared to the encoding quality of a second encoding region outside the focus region at block 95. In some embodiments, the method can further include: providing a lower encoding quality to regions successively farther from the focus region at block 96. Method 90 can further include encoding macroblocks of a video image based on one or more adjusted video encoding parameters at block 97.

[0106] Embodiments of method 90 may be implemented in a system, apparatus, GPU, PPU, or media processor. More specifically, the hardware implementation of method 60 may include configurable logic (such as, for example, PLA, FPGA, CPLD), or fixed-function logic hardware using circuit technologies (such as, for example, ASIC, CMOS, or TTL technology), or any combination thereof. Alternatively or additionally, method 90 may be implemented in one or more modules as a set of logical instructions stored in a machine or computer-readable storage medium (such as, RAM, ROM, PROM, firmware, flash memory, etc.) to be executed by a processor or computing device. For example, computer program code for implementing the operations of the components can be written in any combination of one or more operating system-applicable / suitable programming languages, including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0107] For example, embodiments or portions of method 90 may be implemented as firmware, an application (e.g., via an API), or driver software. Other embodiments or portions of method 90 may be implemented with special-purpose code (e.g., shaders) to be executed on a GPU. Other embodiments or portions of method 90 may be implemented in fixed-function logic or dedicated hardware (e.g., in a GPU or media processor).

[0108] In some embodiments, a focus region may be provided from a focus engine that includes an eye tracker that provides eye position information. Some embodiments may then use the eye position information to increase the quality of where the user is looking and decrease the quality away from the focus region. According to some embodiments, reducing the bit rate, increasing the QP, and / or adjusting other encoding parameters may advantageously save memory, network, and / or computing bandwidth and may provide power savings.

[0109] Some embodiments may be implemented for both wired and wireless applications (e.g., wireless VR). Wireless applications may particularly benefit from selectively reducing the quality away from the focus region. For example, the QP value may increase away from the center (or focus region). Using a lower bit rate / quality can increase the data transmission speed. By primarily using a higher bit rate / quality in the focus region, some embodiments can effectively dedicate resources to the most important region for the user on the screen.

[0110] In one example of adjusting encoding quality, regions can be applied that are based on proximity to a focus region and a desired quality / QP / bitrate, etc. corresponding to that region. For example, regions can be designed based on studies of human visual preferences. For example, there may be multiple different regions, and each of the regions can differ for different quality factors. Any additional screen region outside the outermost region can be processed with the same settings as the outermost region or may be further degraded.

[0111] Now turning to Figures 15A to 15F In, an embodiment of a set of regions for foveal encoding can be represented by any one of two or more continuously enclosed, non - intersecting regions. Although the regions are shown together, each region can be defined and applied separately. The regions can have any suitable shape, such as circular (e.g., Figure 15A and 15C ), elliptical (e.g., Figure 15B , 15C and 15D), square or rectangular (e.g., Figure 15E ), or any shape (e.g., Figure 15F ). The innermost region is typically full encoding quality (e.g., 0% degraded quality / QP / etc.), but in some use cases (e.g., bandwidth - saving settings, see Figure 15F with 10% degradation) degradation can be applied. The innermost region can be surrounded by one or more continuous, non - intersecting regions, where more degradation is applied sequentially to each successive region (e.g., further away from the focus region). These regions can have a common center (e.g., Figure 15A and 15B ) or no common center (e.g., Figure 15D and 15F ). The orientation of the regions can be aligned with the display (e.g., Figure 15A , 15B , 15C and 15E) or not aligned with the display (e.g., Figure 15D and 15F ). The shape of each region can be the same (e.g., Figure 15A , 15B and 15D) or can be different from each other (e.g., Figure 15C , 15E and 15F).

[0112] Now turning to Figure 15G and Figure 15H , an embodiment of a set of regions 100 can be applied to an image region 102 based on a focus region. For example, if the focus region is approximately centered (e.g., if that is where the user is looking), then the regions 100 can also be approximately centered when applied to the image region (e.g., see Figure 15G)。If the focus area moves (e.g., based on gaze information, motion information, content, etc.), then region 100 can similarly move based on the new focus area (e.g., see Figure 15H ). For example, the innermost region can be encoded with a low QP value (e.g., 0% degradation), while the outermost region can be encoded with a higher QP value (e.g., 50% degradation). Additional consecutive regions can be provided with successively higher QP values (e.g., 10% degradation, 30% degradation, etc.) between the innermost and outermost regions.

[0113] In another example of adjusting the encoding quality, a formula can be applied based on the position of the target macroblock relative to the focus area. For example, the system can calculate the shortest distance from the target macroblock to the focus area boundary and degrade the quality proportionally to the calculated distance. Alternatively, a specific macroblock can be selected as the focus (e.g., the focus macroblock), and the distance from the target macroblock to the focus macroblock can be calculated. The system can use a linear formula, a non-linear formula (e.g., parabola, logarithm, etc.), or other suitable formula for proportional encoding quality degradation. The system can also maintain a set of ranges of quality degradation (e.g., adjacent blocks [no degradation]; 2 to 4 blocks apart [20% degradation]; 5 or more blocks apart [50% degradation]).

[0114] In some embodiments, an initial QP value can be selected, or the user can select from a set of predefined regions to determine how much degradation to perform on each region. For example, some people may be more sensitive to certain quality parameters and less sensitive to others. In some embodiments, there can be per-user calibration. Generally, an eye tracker involves some user calibration. Region calibration for various quality parameters can be performed simultaneously with the eye tracker calibration. The calibration can be based on, for example, the just noticeable difference (JND). During the calibration phase, the user may be asked to focus on a region and give a response when noticing a change in the surrounding quality or detail. Or when there is actually a certain difference between the encoded QP values of the images, the user can be asked whether two images look the same on the periphery (e.g., to determine how much change the user perceives). In some embodiments, the calibration or other settings can be user-adjustable (e.g., a setting / slider for more compression / less compression), or included as part of various power / performance settings.

[0115] The human eye can be sensitive to movement on the periphery. Users may not immediately recognize an object, but they may notice the movement. Subsequently, the movement can guide the user's gaze to the movement. Preferably, there is no sharp drop from one region to the next. In the focused region, for example, the encoding quality can be high quality / full quality. In the next region, if the resolution is degraded by 50%, the change may be too obvious. A gradual degradation may be preferred so that the boundaries are less distinct (e.g., 0% to 25% to 50%, etc.). For example, some displays may have a higher pixel density (e.g., 4K displays), and flicker can cause motion sickness or other adverse user experiences. In some embodiments, the gradual degradation can also reduce the perceivable flicker from one region to the next.

[0116] Advantageously, some embodiments can provide adaptive foveal encoding quality. For example, the central region or the focused region can have high-quality encoding, while the edge regions can have lower-quality encoding. Some embodiments can identify the focused region and provide a higher encoding quality at the focused region and a lower encoding quality outside the focused region.

[0117] Some embodiments can provide multiple bands of decreasing encoding quality further away from the focused region, and can encode the outermost band at low quality. For example, the focused region is encoded with a higher quality, and as moving out from the focused region, each additional band or threshold has an encoding quality adjusted to be lower. During encoding, parameters can be set to indicate the encoding quality for each band (e.g., different QP values for each band).

[0118] In some embodiments, a media pipeline including a media engine can perform encoding. Both the focused region (e.g., the eye position) and the macroblock position can be known at the time of encoding. For each macroblock, there may be a parameter indicating how far it is from the focused region. Alternatively, the media engine can calculate how far a macroblock is from the focused region and reduce the encoding quality based on that calculation (e.g., a higher QP value can be used for macroblocks that are further away). After the rendering operation is completed, some embodiments can be implemented in the media pipeline. Computational threads can be dispatched to post-process the raster image. For example, a shader unit can run a compute shader to determine the distance of each macroblock from the focused region, to determine which region each macroblock belongs to and set the corresponding parameters.

[0119] Some embodiments can advantageously provide foveated encoding with visual quality adaptation for VR streaming. Non-limiting applications for some embodiments can include AR, VR, merged reality, mixed reality, computer-generated holography, VR gaming, automotive infotainment, automotive driver assistance systems, holographic lens type applications. Some wired VR systems may have problems in terms of mobility (e.g., because the HMD may be tethered to the base system) and / or cost (e.g., due to the complex graphics operations involved). Certain wireless VR systems may have problems in terms of network bandwidth (e.g., due to the amount of data that needs to be transmitted) and / or image quality (e.g., because the bitrate may be reduced to support the available bandwidth). To improve the visual quality of VR streaming, some wireless VR systems may use high-bandwidth wireless technologies (e.g., 60 GHz wireless technology can provide 4 Gbps data bandwidth). However, such high-bandwidth systems are more costly and may require many proprietary wireless adapters that are not natively supported by user devices. Some embodiments can advantageously provide adaptive foveated encoding based on focus information to improve the quality in the focused area while reducing the bandwidth requirements outside the focused area.

[0120] Visual quality can be very important for various VR experiences. Some embodiments can support a resolution of at least 2560x1440 for a binocular VR system. To have better visual quality, some embodiments can provide a bitrate of at least 80 Mbps for a 2560x1440 display. Perceived latency can be very important for the user experience. A higher bitrate may introduce additional latency. To have low-latency network transmission, some embodiments can utilize a network transport protocol based on the User Datagram Protocol (UDP) to provide very low transmission overhead (e.g., compared to the Transmission Control Protocol (TCP)). Using the UDP mechanism, the size of the Group of Pictures (GOP) may be limited. Some embodiments can utilize an intra-frame-only encoding strategy. One challenge in VR content streaming is how to obtain good visual quality based on a limited bitrate budget (e.g., 80 Mbps@60 FPS with a resolution of 2560x1440) and a limited GOP size. Some embodiments can advantageously provide foveated encoding with visual quality adaptation to improve visual quality in very limited network bandwidth without sacrificing the user experience.

[0121] Some other coding techniques can process the entire frame equally and assign the same QP to all macroblocks. Some embodiments can advantageously provide multiple quality levels based on the foveated vision model. The multiple quality levels can be based on the human eye, which is much more sensitive to the foveated region. Some embodiments can maintain the quality of the foveated region while reducing the quality of the far peripheral regions. By doing so, some embodiments can guarantee the visual quality within the target bitrate. Some embodiments can also adaptively change the GOP size based on the quality. In particular, some embodiments can achieve a high visual quality under limited network bandwidth (e.g., <100Mbps WIFI network), which can reduce the cost and / or complexity of the hardware configuration required to support VR applications. Some embodiments can be easily adapted to follow existing codec standards without any special changes on the receiving side. Some embodiments can be applied to VR applications on cloud WIGIG-based wireless streaming, as well as wired streaming applications (e.g., Ethernet, USB, etc.).

[0122] Now turning to Figure 16 , the binocular system 104 can include a left-eye display area 105 and a right-eye display area 106. Without being limited to the theory of operation, the human visual system can include foveated vision and peripheral vision. The area on the retina corresponding to the fixation center (referred to as the fovea) can have a higher cone cell density than any other area on the retina. In the fovea, retinal ganglion cells may have a smaller receptive field, while in the periphery, they may have a much larger receptive field. On an HMD display, the human eye visual system can follow the foveated vision system, where the user is more sensitive to the quality of the foveated region and less sensitive to the quality of the far peripheral regions. Some VR scenes may be very complex and may be difficult to compress at a low bitrate. Some embodiments can advantageously use a higher bitrate for the foveated region. To simplify the human visual system, some embodiments can divide the quality of the entire VR content into three levels. The first-priority quality regions 105a, 106a can include the foveated region and the near peripheral regions (e.g., +30 degrees, -30 degrees). Since the user may be most sensitive to this region, some embodiments can ensure the quality of this region 105a, 106a. The second-priority quality regions 105b, 106b can include the intermediate peripheral regions. Compared with the first-priority quality regions 105a, 106a, the quality of the second-priority regions will be worse. The third-priority quality regions 105c, 106c can include the far peripheral regions. The regions 105c, 106c can have a relatively worst quality for the entire VR content because the user may be insensitive to this region 105c, 106c.

[0123] Now turning to Figure 17, the image region 110 can be divided into macroblocks 111 for encoding. Some embodiments may include more or fewer macroblocks. In some embodiments, the sizes of the macroblocks may not be uniform. In some embodiments, for example, the macroblocks positioned towards the periphery may be larger than those located near the focus region. Some embodiments may identify a focus region 112 and may provide higher quality encoding for the macroblocks in the focus region 112. Some embodiments may include only the macroblocks that are entirely within the focus region 112 for the highest quality encoding, while some embodiments may additionally include any macroblocks that intersect the focus region 112 for the highest quality encoding. Some embodiments may identify a second region 113 outside the focus region 112 and adjust the encoding quality for the macroblocks that are in the second region 113 but outside the focus region 112 to a lower quality encoding (e.g., a lower quality encoding compared to the highest quality encoding). Some embodiments may include only the macroblocks that are entirely within the second region 113 for the lower quality encoding, while some embodiments may additionally include any macroblocks that intersect the second region 113 for the lower quality encoding. Some embodiments may adjust the encoding quality of the macroblocks outside the focus region 112 and outside the second region 113 to a relatively lowest encoding quality.

[0124] Now turning to Figure 18 , embodiments of the electronic processing system 120 may include a host device 121 (e.g., a personal computer, a server, etc.) communicatively coupled to a user device 122 (e.g., a smart phone, a tablet, an HMD, etc.) in a wired or wireless manner. The host device 121 may include a VR renderer 123 (e.g., a high-performance graphics adapter) communicatively coupled to an adaptive encoder 124, which in turn may be communicatively coupled to a transceiver 125 (e.g., WIFI, WIGIG, Ethernet, etc.). For example, the adaptive encoder 124 may utilize Advanced Video Coding (AVC) and / or High Efficiency Video Coding (HEVC), and may be executed by a graphics processor or an external graphics adapter. The user device 122 may include a transceiver 126 (e.g., configured to exchange information with the transceiver 125 of the host device 121) communicatively coupled to a VR decoder 127, which in turn is communicatively coupled to a display 128. The user device 122 may further include one or more sensors 129. The adaptive encoder 124 may include VR foveated encoder software 124a communicatively coupled to a hardware encoder 124b.

[0125] The host device 121 may include one or more GPUs that implement all or part of the VR renderer 123 and the adaptive encoder 124. The host device 121 may render VR graphics / video image content, encode the content, and stream the VR content to the user device 122. The user device 122 may decode the VR content and present the graphics / video image on the display 128. The user device may also support other local functions such as asynchronous time warping (ATW), frame buffer rendering (FBR), barrel distortion correction, etc. According to some embodiments, the user device 122 may provide sensor-related information back to the host device 121 to be advantageously utilized by the adaptive encoder 124, as described herein. For example, the user device 122 may send back the HMD position, 3DOF information, 6DOF information, and / or other sensor data to the host device 121. For example, the adaptive encoder 124 may also include an adaptive motion encoder as described herein to adjust one or more video encoding parameters based on the sensor-related information received from the user device 122.

[0126] Regardless of the scenario, some embodiments of the system 120 may advantageously ensure visual quality by providing quality-adaptive foveated encoding to improve visual quality. For example, the hardware encoder 124b on the transmitter side may report the frame-level QP for each frame. Based on the reported previous frame-level QP, the VR foveated encoder software 124a may have a general guidance for that level of quality. For example, if the QP is about 41, the quality may be poor, which may affect the user experience. If the QP is about 25 or below, the quality is acceptable and no change may be needed. The VR foveated encoder software 124a may then generate a new set of QP differences / changes (ΔQP) for both the first priority region and the second priority region. Additionally or alternatively, the VR foveated encoder software 124a may adjust the GOP size. In some embodiments, the actual QP values for the first priority region and the second priority region may be determined as follows:

[0127] QP of the first / second region = ΔQP of the first / second region + frame-level QP [Equation 8]

[0128] If ΔQP is less than 0, the quality may be adjusted to be better than other regions.

[0129] In a wireless environment, network packets may sometimes be dropped. Some embodiments may use intra-frame encoding only for a wireless environment. If a VR scene is too complex to be compressed with high quality, some embodiments may increase the GOP size to provide good quality, since inter-frame references can be utilized to have a better compression ratio. Some embodiments may keep the GOP size less than 4 to avoid artifacts or long motion-to-photon effects and improve the VR experience. Advantageously, some embodiments may affect only the encoder side without any special changes on the receiver side.

[0130] Now turning to Figure 19 , the method 130 of adaptive encoding may include initializing a VR encoder at block 131 and preparing a new encoding workload at block 132. When the VR encoder SW starts to prepare a new encoding workload, it may check the encoding statistics from the previously encoded frame (e.g., the frame-level QP of the previous frame). If at block 133, the foveal region QP (frame-level QP + ΔQP) ≤ threshold QP, the VR encoder SW may apply the same settings as the previous frame at block 134 (e.g., at the beginning, ΔQP may be defaultly set to -1). The method 130 may then submit the workload to the GPU at block 135. The method 130 may then determine at block 136 whether all encoding is completed, and if not, may prepare the next encoding workload at block 132.

[0131] If at block 133, the foveal QP is greater than the threshold QP, then at block 137, the method 130 may determine whether there is room to adjust ΔQP. For example, if the frame-level QP is less than the maximum QP (e.g., QP = 51), there may still be room to change ΔQP. Otherwise, there may be no room to adjust ΔQP within the target bitrate. If there is still room to adjust ΔQP, ΔQP may be shrunk a lot to ensure better quality for the fovea. If ΔQP is not applied at block 137, the method 130 may apply ΔQP at block 138 to increase the foveal quality at block 138 and generate a new set of encoding parameters at block 139. The method 130 may then submit the workload to the GPU at block 135. The method 130 may then determine at block 136 whether all encoding is completed, and if not, may prepare the next encoding workload at block 132.

[0132] If ΔQP has been applied at block 137, the VR encoder can attempt to use more P-frames instead of only I-frames (e.g., new GOP = previous GOP + 1 && new GOP < 4) at block 140 and generate a new set of encoding parameters at block 139. In some embodiments, the new GOP can be constrained to a maximum threshold (e.g., GOP less than 4) to avoid longer motion-to-photon delays due to packet loss. If the VR encoder finds that the quality of the foveal region has met the target and can be adjusted by ΔQP, the VR encoder can attempt to reduce the size of the GOP to recheck the quality. Method 130 can then submit the workload to the GPU at block 135. Method 130 can then determine at block 136 whether all encoding is completed, and if not, can prepare the next encoding workload at block 132.

[0133] In some embodiments, a conservative setting of ΔQP can ensure that the far periphery can also have good quality. With such a setting, compared to applying the same QP to each macroblock, at the foveal region of 1024x1024@60fps at a 6Mbps bitrate, foveal encoding can have a quality improvement of about 1.5dB to 3dB. If a more aggressive ΔQP is applied, there can be an even better quality improvement in the foveal region.

[0134] System Overview

[0135] Figure 20 is a block diagram of a processing system 150 according to an embodiment. In various embodiments, system 150 includes one or more processors 152 and one or more graphics processors 158, and can be a single-processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 152 or processor cores 157. In one embodiment, system 150 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile devices, handheld devices, or embedded devices.

[0136] In one embodiment, system 150 may include or may be incorporated within: a server-based gaming platform, a gaming console including a gaming and media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In some embodiments, system 150 is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. Processing system 150 may also include a wearable device, coupled to or integrated within the wearable device, such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In some embodiments, processing system 150 is a television or a set-top box device having one or more processors 152 and a graphical interface generated by one or more graphics processors 158.

[0137] In some embodiments, each of the one or more processors 152 includes one or more processor cores 157 for processor instructions that, when executed, perform operations for system and user software. In some embodiments, each of the one or more processor cores 157 is configured to process a particular instruction set 159. In some embodiments, instruction set 159 may facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). The multiple processor cores 157 may each process a different instruction set 159, and the different instruction sets 159 may include instructions for facilitating the emulation of other instruction sets. Processor core 157 may also include other processing devices, such as a digital signal processor (DSP).

[0138] In some embodiments, processor 152 includes cache memory 154. Depending on the architecture, processor 152 may have a single internal cache or multiple levels of internal cache. In some embodiments, the cache memory is shared among the various components of processor 152. In some embodiments, processor 152 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and the external cache may be shared among the processor cores 157 using known cache coherence techniques. A register file 156 is additionally included in processor 152, and register file 156 may include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). Some registers may be general purpose registers, while other registers may be dedicated to the design of processor 152.

[0139] In some embodiments, one or more processors 152 are coupled to one or more interface buses 160 to transfer communication signals, such as address, data, or control signals, between the processors 152 and other components in the system 150. In one embodiment, the interface bus 160 may be a processor bus, such as a certain version of the Direct Media Interface (DMI) bus. However, the processor bus is not limited to the DMI bus and may include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In one embodiment, the processor(s) 152 include an integrated memory controller 166 and a Platform Controller Hub 180. The memory controller 166 facilitates communication between the memory device and other components of the system 150, while the Platform Controller Hub (PCH) 180 provides connections to I / O devices via a local I / O bus.

[0140] The memory device 170 may be a Dynamic Random Access Memory (DRAM) device, a Static Random Access Memory (SRAM) device, a flash memory device, a Phase Change Memory device, or some other memory device having suitable performance to act as process memory. In one embodiment, the memory device 170 may operate as system memory for the system 150 to store data 172 and instructions 171 for use when one or more processors 152 execute an application or process. The memory controller 166 is also coupled to an optional external graphics processor 162, which may communicate with one or more graphics processors 158 in the processor 152 to perform graphics operations and media operations. In some embodiments, a display device 161 may be connected to the processor(s) 152. The display device 161 may be one or more of the following: an internal display device, such as in a mobile electronic device or a laptop device; or an external display device attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, the display device 161 may be a Head-Mounted Display (HMD), such as a stereoscopic display device for use in Virtual Reality (VR) applications or Augmented Reality (AR) applications.

[0141] In some embodiments, the Platform Controller Hub 180 enables peripheral devices to be connected to the memory device 170 and the processor 152 via a high-speed I / O bus. The I / O peripheral devices include, but are not limited to, an audio controller 196, a network controller 184, a firmware interface 178, a wireless transceiver 176, a touch sensor 175, and a data storage device 174 (e.g., a hard disk drive, a flash memory, etc.). The data storage device 174 may be connected via a storage interface (e.g., SATA) or via a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). The touch sensor 175 may include a touch screen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 176 may be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. The firmware interface 178 enables communication with system firmware and may be, for example, a Unified Extensible Firmware Interface (UEFI). The network controller 184 enables a network connection to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 160. In one embodiment, the audio controller 196 is a multi-channel high-definition audio controller. In one embodiment, the system 150 includes an optional legacy I / O controller 190 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. The Platform Controller Hub 180 may also be connected to one or more Universal Serial Bus (USB) controllers 192 to connect input devices such as a keyboard and mouse 193 combination, a camera 194, or other USB input devices.

[0142] It will be appreciated that the illustrated system 150 is exemplary and not restrictive, as other types of data processing systems configured in different ways may also be used. For example, instances of the memory controller 166 and the Platform Controller Hub 180 may be integrated into a discrete external graphics processor such as the external graphics processor 162. In one embodiment, the Platform Controller Hub 180 and / or the memory controller 166 may be external to one or more of the processors 152. For example, the system 150 may include an external memory controller 166 and a Platform Controller Hub 180 that may be configured as a memory controller hub and a peripheral controller hub within a system chipset that communicates with the (one or more) processors 152.

[0143] Figure 21 is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A - 202N, an integrated memory controller 214, and an integrated graphics processor 208. Figure 21Those elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in a manner similar to any described elsewhere herein, but are not limited thereto. The processor 200 can include additional cores, up to and including additional cores 202N represented by the dashed boxes. Each of the processor cores 202A - 202N includes one or more internal cache units 204A - 204N. In some embodiments, each processor core also has access to one or more shared cache units 206.

[0144] The internal cache units 204A - 204N and the shared cache unit 206 represent a cache memory hierarchy within the processor 200. The cache memory hierarchy can include at least one level of instruction and data cache within each processor core and one or more levels of shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In some embodiments, cache coherence logic maintains coherence between the cache units 206 and 204A - 204N.

[0145] In some embodiments, the processor 200 can also include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI buses or PCI Express buses. The system agent core 210 provides management functions for the various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).

[0146] In some embodiments, one or more of the processor cores 202A - 202N include support for simultaneous multithreading. In such embodiments, the system agent core 210 includes components for coordinating and operating the cores 202A - 202N during multithreading. The system agent core 210 can additionally include a power control unit (PCU), which includes logic and components for regulating the power states of the processor cores 202A - 202N and the graphics processor 208.

[0147] In some embodiments, the processor 200 additionally includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to a set of shared cache units 206 and to a system agent core 210 that includes one or more integrated memory controllers 214. In some embodiments, the system agent core 210 also includes a display controller 211 for driving the graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 can also be a separate module coupled to the graphics processor via at least one interconnect, or can be integrated within the graphics processor 208.

[0148] In some embodiments, a ring-based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units can be used, such as point-to-point interconnects, switched interconnects, or other techniques, including techniques well known in the art. In some embodiments, the graphics processor 208 is coupled to the ring interconnect 212 via an I / O link 213.

[0149] Exemplary I / O link 213 represents at least one of a plurality of various I / O interconnects, including an on-package I / O interconnect that facilitates communication between the various processor components and a high-performance embedded memory module 218, such as an eDRAM module. In some embodiments, each of the processor cores 202A - 202N and the graphics processor 208 can use the embedded memory module 218 as a shared last-level cache.

[0150] In some embodiments, the processor cores 202A - 202N are homogeneous cores that execute the same instruction set architecture. In another embodiment, the processor cores 202A - 202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A - 202N execute a first instruction set, while at least one of the other cores executes a subset of the first instruction set or a different instruction set. In one embodiment, the processor cores 202A - 202N are heterogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled to one or more power cores with lower power consumption. Additionally, the processor 200 can be implemented on one or more chips, or can be implemented as a SoC integrated circuit with the components shown, among other components.

[0151] Figure 22is a block diagram of a graphics processor 300, which can be a discrete graphics processing unit or can be a graphics processor integrated with multiple processing cores. In some embodiments, the graphics processor communicates via a memory-mapped I / O interface to registers on the graphics processor and using commands placed in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.

[0152] In some embodiments, the graphics processor 300 further includes a display controller 302 for driving display output data to a display device 320. The display controller 302 includes hardware for the composition of one or more overlay planes for the display and multi-layered video or user interface elements. The display device 320 can be an internal or external display device. In one embodiment, the display device 320 is a head-mounted display device, such as, a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding media into one or more media encoding formats, decoding media from one or more media encoding formats, or transcoding media between one or more media encoding formats, the one or more media encoding formats including but not limited to: Moving Picture Experts Group (MPEG) formats (such as, MPEG-2), Advanced Video Coding (AVC) formats (such as, H.264 / MPEG-4 AVC and Society of Motion Picture and Television Engineers (SMPTE) 421M / VC-1), and Joint Photographic Experts Group (JPEG) formats (such as, JPEG, and Motion JPEG (MJPEG) formats).

[0153] In some embodiments, the graphics processor 300 includes a Block Image Transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterizer operations, including for example, bit-boundary block transfers. However, in one embodiment, one or more components of the Graphics Processing Engine (GPE) 310 are used to perform 2D graphics operations. In some embodiments, the GPE 310 is a computational engine for performing graphics operations, the graphics operations including three-dimensional (3D) graphics operations and media operations.

[0154] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed-function elements that perform various tasks into the elements of the 3D / media subsystem 315 and / or within the generated execution threads. Although the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 that is dedicated to performing media operations, such as video post-processing and image enhancement.

[0155] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units for performing one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of, or on behalf of, the video codec engine 306. In some embodiments, the media pipeline 316 additionally includes a thread generation unit to generate threads for execution on the 3D / media subsystem 315. The generated threads perform calculations for media operations on one or more graphics execution units included in the 3D / media subsystem 315.

[0156] In some embodiments, the 3D / media subsystem 315 includes logic for executing the threads generated by the 3D pipeline 312 and the media pipeline 316. In one embodiment, the pipeline sends thread execution requests to the 3D / media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching various requests for available thread execution resources. The execution resources include an array of graphics execution units for processing 3D threads and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes a shared memory for sharing data between threads and for storing output data, which includes registers and addressable memory.

[0157] Graphics Processing Engine

[0158] Figure 23 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is Figure 22 a certain version of the GPE 310 shown in Figure 23 Those elements having the same reference numerals (or names) as the elements in any other figure herein can operate or function in a manner similar to any described elsewhere herein, but are not limited thereto. For example, shown Figure 223D pipeline 312 and media pipeline 316. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.

[0159] In some embodiments, the GPE 410 is coupled to or includes a command stream converter 403 that provides a command stream to the 3D pipeline 312 and / or the media pipeline 316. In some embodiments, the command stream converter 403 is coupled to a memory, which may be a system memory, or one or more of an internal cache memory and a shared cache memory. In some embodiments, the command stream converter 403 receives commands from the memory and sends these commands to the 3D pipeline 312 and / or the media pipeline 316. These commands are instructions fetched from a ring buffer that stores commands for the 3D pipeline 312 and the media pipeline 316. In one embodiment, the ring buffer may additionally include a batch command buffer that stores a batch of multiple commands. Commands for the 3D pipeline 312 may also include references to data stored in the memory, such as but not limited to vertex data and geometry data for the 3D pipeline 312 and / or image data and memory objects for the media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process commands and data by performing operations via logic within their respective pipelines or by dispatching one or more execution threads to the graphics core array 414. In one embodiment, the graphics core array 414 includes a block of one or more graphics cores (e.g., (multiple) graphics cores 415A, (multiple) graphics cores 415B), each block including one or more graphics cores. Each graphics core includes a set of graphics execution resources, the set of graphics execution resources including: general execution logic and graphics-specific execution logic for performing graphics operations and computational operations; and fixed-function texture processing logic and / or machine learning and artificial intelligence acceleration logic.

[0160] In embodiments, the 3D pipeline 312 includes fixed-function and programmable logic for processing one or more shader programs by processing instructions and dispatching execution threads to the graphics core array 414, the one or more shader programs such as, vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs. The graphics core array 414 provides a unified block of execution resources for use in processing these shader programs. The multi-functional execution logic (e.g., execution units) within the (multiple) graphics cores 415A - 415B of the graphics core array 414 includes support for various 3D API shader languages and can execute multiple synchronized execution threads associated with multiple shaders.

[0161] In some embodiments, the graphics core array 414 further includes execution logic for performing media functions such as video and / or image processing. In one embodiment, the execution units additionally include general-purpose logic that is programmable to perform parallel general-purpose computing operations in addition to performing graphics processing operations. The general-purpose logic can execute processing operations in parallel with or in combination with the (multiple) processor cores 157 of Figure 20 or the general-purpose logic within the cores 202A-202N in Figure 21 .

[0162] Output data generated by threads executing on the graphics core array 414 can output the data to memory in a unified return buffer (URB) 418. The URB 418 can store data for multiple threads. In some embodiments, the URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, the URB 418 can additionally be used for synchronization between threads on the graphics core array and fixed-function logic within the shared function logic 420.

[0163] In some embodiments, the graphics core array 414 is scalable such that the array includes a variable number of graphics cores, each having a variable number of execution units based on the target power and performance levels of the GPE 410. In one embodiment, the execution resources are dynamically scalable such that execution resources can be enabled or disabled as needed.

[0164] The graphics core array 414 is coupled to shared function logic 420, which includes a plurality of resources that are shared among the graphics cores in the graphics core array. The shared functions within the shared function logic 420 are hardware logic units that provide specialized complementary functions to the graphics core array 414. In embodiments, the shared function logic 420 includes, but is not limited to, sampler logic 421, math logic 422, and inter-thread communication (ITC) logic 423. Additionally, some embodiments implement one or more caches 425 within the shared function logic 420.

[0165] Implementing shared functionality in cases where the demand for a given specialized function is not sufficient to be included in the graphics core array 414. Instead, a single instantiation of that specialized function is implemented as a separate entity within the shared function logic 420 and is shared among the execution resources within the graphics core array 414. The exact set of functions that are shared among and included within the graphics core array 414 varies from embodiment to embodiment. In some embodiments, specific shared functions that are widely used by the graphics core array 414 within the shared function logic 420 may be included within the shared function logic 416 within the graphics core array 414. In various embodiments, the shared function logic 416 within the graphics core array 414 may include some or all of the logic within the shared function logic 420. In one embodiment, all of the logic elements within the shared function logic 420 may be replicated within the shared function logic 416 of the graphics core array 414. In one embodiment, the shared function logic 420 is excluded in favor of the shared function logic 416 within the graphics core array 414.

[0166] Figure 24 is a block diagram of the hardware logic of a graphics processing unit core 500 according to some embodiments described herein. Figure 24 Those elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in a manner similar to any described elsewhere herein, but are not limited thereto. In some embodiments, the illustrated graphics processing unit core 500 is included within Figure 23 the graphics core array 414. The graphics processing unit core 500 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. An example of a graphics processing unit core 500 is a graphics core slice, and based on the target power envelope and performance envelope, a graphics processor as described herein can include multiple graphics core slices. Each graphics core 500 may include a fixed function block 530 that is coupled to a plurality of sub-cores 501A - 501F (also referred to as sub-slices), the plurality of sub-cores 501A - 501F including blocks of modular general and fixed function logic.

[0167] In some embodiments, the fixed function block 530 includes a geometry / fixed function pipeline 536 that can be shared by all of the sub-cores within the graphics processor 500, for example, in a lower performance and / or lower power graphics processor implementation. In various embodiments, the geometry / fixed function pipeline 536 includes a 3D fixed function pipeline (e.g., Figure 22 and Figure 23 the 3D pipeline 312 in Figure 23Unified return buffer 418.

[0168] In one embodiment, the fixed function block 530 further includes a graphics SoC interface 537, a graphics microcontroller 538, and a media pipeline 539. The graphics SoC interface 537 provides an interface between the graphics core 500 and other processor cores within a system-on-chip integrated circuit. The graphics microcontroller 538 is a programmable sub-processor configurable to manage various functions of the graphics processor 500, including thread dispatching, scheduling, and preemption. The media pipeline 539 (e.g., Figure 22 and Figure 23 media pipeline 316) includes logic for facilitating decoding, encoding, pre-processing, and / or post-processing of multimedia data including image data and video data. The media pipeline 539 implements media operations via requests to the computing or sampling logic within the sub-cores 501A - 501F.

[0169] In one embodiment, the SoC interface 537 enables the graphics core 500 to communicate with general-purpose application processor cores (e.g., CPUs) and / or other components within the SoC, including memory hierarchy elements such as shared last-level cache memories, system RAM, and / or embedded on-chip or on-package DRAM. The SoC interface 537 may also enable communication with fixed function devices within the SoC such as camera imaging pipelines, and enable the use and / or implementation of global memory atomicity, which may be shared between the graphics core 500 and a CPU within the SoC. The SoC interface 537 may also implement power management control for the graphics core 500, and enable an interface between the clock domain of the graphics core 500 and other clock domains within the SoC. In one embodiment, the SoC interface 537 enables receipt of command buffers from a command stream converter and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. When media operations are to be performed, these commands and instructions may be dispatched to the media pipeline 539, or when graphics processing operations are to be performed, these commands and instructions may be dispatched to the geometry and fixed function pipelines (e.g., geometry and fixed function pipeline 536, geometry and fixed function pipeline 514).

[0170] The graphics microcontroller 538 can be configured to perform various scheduling tasks and management tasks for the graphics core 500. In one embodiment, the graphics microcontroller 538 can schedule graphics and / or compute workloads for the execution unit (EU) arrays 502A - 502F within the sub - cores 501A - 501F and the respective graphics parallel engines within 504A - 504F. In this scheduling model, host software executing on the CPU core of the SoC that includes the graphics core 500 can submit a workload via one of the multiple graphics processor doorbells, which invokes a scheduling operation for the appropriate graphics engine. The scheduling operations include: determining which workload to run next, submitting the workload to the command stream converter, pre - empting an existing workload running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is complete. In one embodiment, the graphics microcontroller 538 can also facilitate the low - power or idle state of the graphics core 500, thereby providing the ability to save and restore the registers within the graphics core 500 across low - power state transitions independent of the operating system and / or the graphics driver software on the system.

[0171] The graphics core 500 can have more or fewer sub - cores 501A - 501F than shown, up to N modular sub - cores. For each set of N sub - cores, the graphics core 500 can also include shared functional logic 510, shared and / or cache memory 512, a geometry / fixed - function pipeline 514, and additional fixed - function logic 516 for accelerating various graphics and compute processing operations. The shared functional logic 510 can include logic units associated with shared functional logic 420 (e.g., sampler logic, math logic, and / or inter - thread communication logic) that can be shared by each N sub - cores within the graphics core 500. Figure 23 The shared and / or cache memory 512 can be the last - level cache for the set 501A - 501F of N sub - cores within the graphics core 500 and can also act as shared memory accessible by multiple sub - cores. The geometry / fixed - function pipeline 514, rather than the geometry / fixed - function pipeline 536, can be included within the fixed - function block 530, and the geometry / fixed - function pipeline 514 can include the same or similar logic units.

[0172] In one embodiment, the graphics core 500 includes additional fixed function logic 516, which may include various fixed function acceleration logics for use by the graphics core 500. In one embodiment, the additional fixed function logic 516 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are two geometry pipelines: the full geometry pipeline within the geometry / fixed function pipelines 516, 536; and a culling pipeline, which is an additional geometry pipeline that may be included within the additional fixed function logic 516. In one embodiment, the culling pipeline is a streamlined version of the full geometry pipeline. The full pipeline and the culling pipeline may execute different instances of the same application, each instance having a separate context. Position-only shading can hide the long culling runs of discarded triangles, thus enabling earlier completion of shading in some instances. For example and in one embodiment, the culling pipeline logic within the additional fixed function logic 516 can execute the position shader in parallel with the main application and generally generate key results faster than the full pipeline, because the culling pipeline only takes the position attributes of the vertices and shades the position attributes of the vertices, without performing rasterization and rendering of pixels to the frame buffer. The culling pipeline can use the generated key results to calculate the visibility information of all triangles, regardless of whether those triangles are culled. The full pipeline (which may be referred to as a replay pipeline in this instance) can consume this visibility information to skip the culled triangles, thus only shading the visible triangles that are ultimately passed to the rasterization stage.

[0173] In one embodiment, the additional fixed function logic 516 may further include machine learning acceleration logic, such as fixed function matrix multiplication logic, for an implementation that includes optimizations for machine learning training or inference.

[0174] Each graphics sub - core 501A - 501F includes a set of execution resources that can be used to perform graphics operations, media operations, and compute operations in response to requests made by the graphics pipeline, media pipeline, or shader program. The graphics sub - cores 501A - 501F include: multiple EU arrays 502A - 502F, 504A - 504F; thread dispatch and inter - thread communication (TD / IC) logic 503A - 503F; 3D (e.g., texture) samplers 505A - 505F; media samplers 506A - 506F; shader processors 507A - 507F; and shared local memory (SLM) 508A - 508F. Each of the EU arrays 502A - 502F, 504A - 504F includes multiple execution units, which are general - purpose graphics processing units capable of performing floating - point and integer / fixed - point logical operations to serve graphics operations, media operations, or compute operations (including graphics programs, media programs, or compute shader programs). The TD / IC logic 503A - 503F performs local thread dispatch and thread control operations for the execution units within the sub - core and facilitates communication between the threads executing on the execution units of the sub - core. The 3D samplers 505A - 505F can read texture or other 3D graphics - related data into memory. The 3D samplers can read texture data in different ways based on the configured sample state and the texture format associated with a given texture. The media samplers 506A - 506F can perform similar read operations based on the type and format associated with the media data. In one embodiment, each graphics sub - core 501A - 501F can alternatively include a unified 3D and media sampler. Threads executing on the execution units within each of the sub - cores 501A - 501F can utilize the shared local memory 508A - 508F within each sub - core to enable the threads executing within a thread group to use a common pool of on - chip memory for execution.

[0175] Execution Unit

[0176] Figures 25A - 25B Illustrates thread execution logic 600 according to an embodiment described herein, the thread execution logic 600 including an array of processing elements employed in a graphics processor core. Figures 25A - 25B Those elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in a manner similar to any described elsewhere herein, but are not limited thereto. Figure 25A Illustrates an overview of thread execution logic 600, which may include a variant of the hardware logic shown in Figure 24 each of the sub - cores 501A - 501F. Figure 25B Illustrates exemplary internal details of an execution unit.

[0177] As inFigure 25A As shown, in some embodiments, thread execution logic 600 includes a shader processor 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array including a plurality of execution units 608A - 608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 608A, 608B, 608C, 608D, up to 608N - 1 and 608N) based on the computational requirements of the workload. In one embodiment, the included components are interconnected via an interconnect structure that links to each of the components. In some embodiments, thread execution logic 600 includes one or more connections to memory (such as system memory or cache memory) through instruction cache 606, data port 614, sampler 610, and one or more of execution units 608A - 608N. In some embodiments, each execution unit (e.g., 608A) is an independent programmable general - purpose computing unit capable of executing multiple synchronous hardware threads and processing multiple data elements in parallel for each thread. In embodiments, the array of execution units 608A - 608N is scalable to include any number of individual execution units.

[0178] In some embodiments, execution units 608A - 608N are primarily used to execute shader programs. Shader processor 602 can process various shader programs and can dispatch execution threads associated with the shader programs via thread dispatcher 604. In one embodiment, the thread dispatcher includes logic for arbitrating requests for threads initiated from the graphics pipeline and the media pipeline and instantiating the requested threads on one or more of execution units 608A - 608N. For example, the geometry pipeline can dispatch a vertex shader, a tessellation shader, or a geometry shader to the thread execution logic for processing. In some embodiments, thread dispatcher 604 can also process runtime thread generation requests from executed shader programs.

[0179] In some embodiments, execution units 608A - 608N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct3D and OpenGL) to be executed with minimal translation. These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). Each of the execution units 608A - 608N is capable of multi - issue single - instruction multiple - data (SIMD) execution, and multi - threaded operation enables an efficient execution environment in the face of higher - latency memory accesses. Each hardware thread within each execution unit has a dedicated high - bandwidth register file and associated independent thread state. The pipeline, which is capable of integer operations, single - precision floating - point operations, and double - precision floating - point operations, capable of SIMD branching, capable of logical operations, capable of transcendental operations, and capable of other miscellaneous operations, is multi - issue per clock. When waiting for data from one of the shared functions in memory or a shared function, the dependency logic within execution units 608A - 608N puts the waiting threads to sleep until the requested data has returned. While the waiting threads are sleeping, the hardware resources can be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution units can perform operations for pixel shaders, fragment shaders, or another type of shader program including a different vertex shader.

[0180] Each of the execution units 608A - 608N operates on an array of data elements. The number of data elements is the "execution size", or the number of channels for an instruction. Execution channels are the logical units for data - element access, masking, and execution of flow control within an instruction. The number of channels can be independent of the number of physical arithmetic - logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In some embodiments, execution units 608A - 608N support integer and floating - point data types.

[0181] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution unit will process each element based on the data size of the element. For example, when operating on a 256-bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 64-bit packed data elements (quadword (QW) size data elements), eight separate 32-bit packed data elements (doubleword (DW) size data elements), sixteen separate 16-bit packed data elements (word (W) size data elements), or thirty-two separate 8-bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible.

[0182] In one embodiment, one or more execution units can be combined into fused execution units 609A - 609N that have thread control logic (607A - 607N) common to the fused EUs. Multiple EUs can be fused into an EU group. Each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary according to the embodiment. Additionally, various SIMD widths can be executed per EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 609A - 609N includes at least two execution units. For example, fused execution unit 609A includes a first EU 608A, a second EU 608B, and thread control logic 607A common to the first EU 608A and the second EU 608B. Thread control logic 607A controls the threads executed on the fused graphics execution unit 609A, allowing each EU within the fused execution units 609A - 609N to execute using a common instruction pointer register.

[0183] One or more internal instruction caches (e.g., 606) are included in the thread execution logic 600 to cache thread instructions for the execution units. In some embodiments, one or more data caches (e.g., 612) are included to cache thread data during thread execution. In some embodiments, a sampler 610 is included to provide texture sampling for 3D operations and media sampling for media operations. In some embodiments, the sampler 610 includes specialized texture or media sampling functions to process texture data or media data during the sampling process before providing the sampled data to the execution units.

[0184] During execution, the graphics pipeline and the media pipeline send thread initiation requests to the thread execution logic 600 via the thread generation and dispatch logic. Once a set of geometric objects has been processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 602 is invoked to further compute the output information and cause the results to be written to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes the values of vertex attributes, and the values of the vertex attributes are interpolated across the rasterized objects. In some embodiments, the pixel processor logic within the shader processor 602 then executes a pixel shader program or a fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 602 dispatches threads to execution units (e.g., 608A) via the thread dispatcher 604. In some embodiments, the shader processor 602 uses the texture sampling logic in the sampler 610 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometric data compute the pixel color data for each geometric fragment or discard one or more pixels without further processing.

[0185] In some embodiments, the data port 614 provides a memory access mechanism for the thread execution logic 600 to output the processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, the data port 614 includes or is coupled to one or more cache memories (e.g., the data cache 612) to cache the data for memory access via the data port.

[0186] As Figure 25B shown, the graphics execution unit 608 may include an instruction fetch unit 637, a general register file array (GRF) 624, an architectural register file array (ARF) 626, a thread arbiter 622, a send unit 630, a branch unit 632, a set of SIMD floating point units (FPUs) 634, and in one embodiment, a set of dedicated integer SIMD ALUs 635. The GRF 624 and the ARF 626 include a set of general register files and architectural register files associated with each synchronous hardware thread that can be active in the graphics execution unit 608. In one embodiment, the per-thread architectural state is maintained in the ARF 626, while the data used during thread execution is stored in the GRF 624. The execution state of each thread, including the instruction pointer for each thread, can be held in a thread-specific register in the ARF 626.

[0187] In one embodiment, the graphics execution unit 608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronous threads and the number of registers per execution unit, where the execution unit resources are divided across the logic for executing multiple synchronous threads.

[0188] In one embodiment, the graphics execution unit 608 can co-issue multiple instructions, which can each be different instructions. The thread arbiter 622 of the graphics execution unit thread 608 can dispatch the instructions to one of the following for execution: the send unit 630, the branch unit 642, or the (multiple) SIMD FPU 634. Each execution thread can access 128 general-purpose registers within the GRF 624, where each register can store 32 bytes that can be accessed as a SIMD 8-element vector with 32-bit data elements. In one embodiment, each execution unit thread has access to 4 kilobytes within the GRF 624, but the embodiments are not limited thereto, and more or fewer register resources can be provided in other embodiments. In one embodiment, up to seven threads execute simultaneously, but the number of threads per execution unit can also vary according to the embodiment. In an embodiment where seven threads can access 4 kilobytes, the GRF 624 can store a total of 28 kilobytes. Flexible addressing modes can permit addressing multiple registers together to create a wider virtual register or represent a strided rectangular block data structure.

[0189] In one embodiment, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the messaging send unit 630. In one embodiment, branch instructions are dispatched to a dedicated branch unit 632 to facilitate SIMD divergence and eventual convergence.

[0190] In one embodiment, the graphics execution unit 608 includes one or more SIMD floating-point units (FPUs) 634 for performing floating-point operations. In one embodiment, the (multiple) FPUs 634 also support integer computations. In one embodiment, the (multiple) FPUs 634 can SIMD execute up to a number M of 32-bit floating-point (or integer) operations, or SIMD execute up to 2M of 16-bit integer or 16-bit floating-point operations. In one embodiment, at least one of the (multiple) FPUs provides extended math capabilities that support high-throughput transcendental math functions and double-precision 64-bit floating point. In some embodiments, a set 635 of 8-bit integer SIMD ALUs also exists and can be specifically optimized to perform operations associated with machine learning computations.

[0191] In one embodiment, an array of multiple instances of the graphics execution unit 608 may be instantiated in graphics sub-core groupings (e.g., sub-slices). For scalability, the product architect may choose the exact number of execution units per sub-core grouping. In one embodiment, the execution unit 608 may execute instructions across multiple execution channels. In a further embodiment, each thread executed on the graphics execution unit 608 is executed on a different channel.

[0192] Figure 26 FIG. 4 is a block diagram showing a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set with instructions of multiple formats. The solid box diagrams generally represent components that are included in the execution unit instructions, while the dashed lines include optional or components that are only included in a subset of the instructions. In some embodiments, the described and shown instruction format 700 is a macro-instruction, as they are the instructions supplied to the execution unit, as opposed to micro-instructions that result from instruction decoding once the instruction has been processed.

[0193] In some embodiments, the graphics processor execution unit natively supports instructions in the 128-bit instruction format 710. Based on the selected instructions, instruction options, and number of operands, the 64-bit compact instruction format 730 may be used for some instructions. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The native instructions available in the 64-bit format 730 vary by embodiment. In some embodiments, a set of index values in the index field 713 is used to partially compress the instruction. The execution unit hardware references a set of compression tables based on the index values and uses the compressed table output to reconstruct the native instruction in the 128-bit instruction format 710.

[0194] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel representing a texture element or a picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control of certain execution options, such as channel selection (e.g., assertion) and data channel order (e.g., blend). For instructions in the 128-bit instruction format 710, the execution size field 716 limits the number of data channels that will be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.

[0195] Some execution unit instructions have up to three operands, including two source operands src0 720, src1 722, and one destination operand 718. In some embodiments, the execution unit supports dual-destination instructions, where one of the dual destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction may be an immediate (e.g., hard-coded) value passed with the instruction.

[0196] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that, for example, specifies whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are provided directly by bits in the instruction.

[0197] In some embodiments, the 128-bit instruction format 710 includes an access / addressing mode field 726 that specifies the addressing mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes including a 16-byte alignment access mode and a 1-byte alignment access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte-aligned addressing for all source and destination operands.

[0198] In one embodiment, the addressing mode portion of the access / addressing mode field 726 determines whether the instruction is to use direct addressing or indirect addressing. When using direct register addressing mode, the register addresses of one or more operands are provided directly by bits in the instruction. When using indirect register addressing mode, the register addresses of one or more operands may be calculated based on the address register value and the address immediate field in the instruction.

[0199] In some embodiments, instructions are grouped based on the 712-bit opcode field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode grouping shown is merely an example. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares the five most significant bits (MSBs), where the move (mov) instruction takes the form 0000xxxxb and the logic instruction takes the form 0001xxxxb. The flow control instruction group 744 (e.g., call, jump) includes instructions in the form 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) in the form 0011xxxxb (e.g., 0x30). The parallel math instruction group 748 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form 0100xxxxb (e.g., 0x40). The parallel math group 748 performs arithmetic operations in parallel across data channels. The vector math group 750 includes arithmetic instructions (e.g., dp4) in the form 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations.

[0200] Graphics Pipeline

[0201] Figure 27 is a block diagram of another embodiment of the graphics processor 800. Figure 27 Those elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to that described elsewhere herein, but are not limited thereto.

[0202] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a render output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued through the ring interconnect 802 to the graphics processor 800. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command stream translator 803, which supplies instructions to the various components of the geometry pipeline 820 or the media pipeline 830.

[0203] In some embodiments, the command stream converter 803 directs the operation of the vertex fetcher 805, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 803. In some embodiments, the vertex fetcher 805 provides the vertex data to the vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, the vertex fetcher 805 and the vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A - 852B via a thread dispatcher 831.

[0204] In some embodiments, the execution units 852A - 852B are an array of vector processors having instruction sets for performing graphics operations and media operations. In some embodiments, the execution units 852A - 852B have attached L1 caches 851 dedicated to each array or shared between the arrays. The cache can be configured as a data cache, an instruction cache, or a single cache partitioned to contain data and instructions in different partitions.

[0205] In some embodiments, the geometry pipeline 820 includes a tessellation component for performing hardware - accelerated tessellation of 3D objects. In some embodiments, the programmable hull shader 811 configures the tessellation operation. The programmable domain shader 817 provides backend evaluation of the tessellation output. The tessellator 813 operates under the direction of the hull shader 811 and includes dedicated logic for generating a set of detailed geometric objects based on a coarse geometric model that is provided as input to the geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation component (e.g., the hull shader 811, the tessellator 813, and the domain shader 817) can be bypassed.

[0206] In some embodiments, a complete geometric object can be processed by the geometry shader 819 via one or more threads dispatched to the execution units 852A - 852B, or can proceed directly to the clipper 829. In some embodiments, the geometry shader operates on entire geometric objects rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 can be programmed by a geometry shader program to perform geometric tessellation in the case where the tessellation unit is disabled.

[0207] Before rasterization, the clipper 829 processes vertex data. The clipper 829 can be a fixed-function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer and depth test component 873 in the render output pipeline 870 dispatches pixel shaders to convert geometric objects to a per-pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, an application can bypass the rasterizer and depth test component 873 and access the un-rasterized vertex data via the out-flow unit 823.

[0208] The graphics processor 800 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed among the major components of the processor. In some embodiments, the execution units 852A - 852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via data ports 856 to perform memory accesses and communicate with the render output pipeline components of the processor. In some embodiments, the sampler 854, caches 851, 858, and execution units 852A - 852B each have separate memory access paths. In one embodiment, the texture cache 858 can also be configured as a sampler cache.

[0209] In some embodiments, the render output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects to associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. Associated render cache 878 and depth cache 879 are also available in some embodiments. The pixel operation component 877 performs pixel-based operations on data, but in some instances, pixel operations associated with 2D operations (e.g., bit-block blitting with blending) are performed by the 2D engine 841 or, at display time, by the display controller 843 using an overlay display plane instead. In some embodiments, a shared L3 cache 875 is available for all graphics components, allowing data to be shared without using the main system memory.

[0210] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from the command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front end 834 processes the commands before sending the media commands to the media engine 837. In some embodiments, the media engine 837 includes a thread generation function for generating threads for dispatch to the thread execution logic 850 via the thread dispatcher 831.

[0211] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and is coupled to the graphics processor via the ring interconnect 802, or some other interconnect bus or fabric. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 includes dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system integrated display device (such as in a laptop computer) or an external display device attached via a display device connector.

[0212] In some embodiments, the geometry pipeline 820 and the media pipeline 830 can be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). In some embodiments, the driver software of the graphics processor converts API calls dedicated to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for all of the open graphics library (OpenGL), open computing language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. In some embodiments, support can also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, combinations of these libraries can be supported. Support can also be provided for the open source computer vision library (OpenCV). Future APIs with a compatible 3D pipeline will also be supported if a mapping from the pipeline of the future API to the pipeline of the graphics processor can be made.

[0213] Graphics Pipeline Programming

[0214] Figure 28A is a block diagram showing a graphics processor command format 900 according to some embodiments. Figure 28B is a block diagram showing a graphics processor command sequence 910 according to an embodiment. Figure 28AThe solid boxes therein show the components generally included in a graphics command, while the dashed lines include optional components or those only included in a subset of the graphics commands. Figure 28A An exemplary graphics processor command format 900 includes data fields for identifying a client 902 of the command, a command operation code (opcode) 904, and data 906. A sub-opcode 905 and a command size 908 are also included in some commands.

[0215] In some embodiments, the client 902 specifies the client unit of the graphics device that processes the command data. In some embodiments, the graphics processor command parser examines the client field of each command to adjust further processing of the command and route the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 904 and the sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses the information within the data field 906 to execute the command. For some commands, an explicit command size 908 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands in the command based on the command opcode. In some embodiments, commands are aligned via multiples of a double word.

[0216] Figure 28B The flowchart therein shows an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of the graphics processor uses a certain version of the shown command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described only for illustrative purposes, as embodiments are not limited to these specific commands or this command sequence. Also, commands may be issued as a batch of commands in a command sequence such that the graphics processor will process the command sequence in at least a partially concurrent manner.

[0217] In some embodiments, the graphics processor command sequence 910 may begin with a pipeline dump flush command 912 to cause any active graphics pipeline to complete the current outstanding commands of the pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate concurrently. A pipeline dump flush is performed to cause the active graphics pipeline to complete any outstanding commands. In response to the pipeline dump flush, the command parser for the graphics processor will pause command processing until the active rendering engine has completed the outstanding operations and the associated read caches are invalidated. Optionally, any data marked as "dirty" in the render cache may be dumped to memory. In some embodiments, the pipeline dump flush command 912 may be used for pipeline synchronization or prior to placing the graphics processor in a low power state.

[0218] In some embodiments, a pipeline select command 913 is used when the command sequence requires the graphics processor to explicitly switch between pipelines. In some embodiments, only one pipeline select command 913 is required in the execution context prior to issuing pipeline commands, unless the context will issue commands for both pipelines. In some embodiments, a pipeline dump flush command 912 is required immediately prior to a pipeline switch via the pipeline select command 913.

[0219] In some embodiments, a pipeline control command 914 configures the graphics pipeline for operation and programs the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state of the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline prior to processing a batch of commands.

[0220] In some embodiments, a return buffer status command 916 is used to configure a set of return buffers for the corresponding pipeline to write data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during operation. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer status 916 includes selecting the size and number of return buffers to be used for a set of pipeline operations.

[0221] The remaining commands in the command sequence differ based on the active pipeline for operation. Based on the pipeline determination 920, the command sequence is customized for the 3D pipeline 922 starting with the 3D pipeline state 930, or the media pipeline 924 starting at the media pipeline state 940.

[0222] Commands for configuring the 3D pipeline state 930 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured prior to processing 3D primitive commands. The values of these commands are determined at least in part based on the particular 3D API in use. In some embodiments, the 3D pipeline state 930 commands are also capable of selectively disabling or bypassing certain pipeline elements if they will not be used.

[0223] In some embodiments, the 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate multiple vertex data structures. The vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 commands are used to perform vertex operations on the 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.

[0224] In some embodiments, the 3D pipeline 922 is triggered via an execute 934 command or event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a "go" or "kick" command in a command sequence. In one embodiment, a pipeline synchronization command is used to trigger command execution in order to flush a command sequence through the graphics pipeline. The 3D pipeline will perform geometric processing on the 3D primitives. Once the operations are complete, the resulting geometric objects are rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands for controlling pixel coloring and pixel backend operations may also be included.

[0225] In some embodiments, when performing media operations, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the particular uses and ways of programming the media pipeline 924 depend on the media or compute operations to be performed. During media decoding, specific media decoding operations may be transferred to the media pipeline. In some embodiments, the media pipeline may also be bypassed, and the resources provided by one or more general-purpose processing cores may be used to perform media decoding wholly or in part. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0226] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. A set of commands for configuring the media pipeline state 940 is dispatched or placed into a command queue, before the media object commands 942. In some embodiments, the commands 940 for the media pipeline state include data for configuring media pipeline elements that will be used to process media objects. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the commands 940 for the media pipeline state also support the use of one or more pointers to "indirect" state elements that point to a bulk of state settings.

[0227] In some embodiments, the media object commands 942 supply pointers to media objects to be processed by the media pipeline. The media objects include memory buffers that contain video data to be processed. In some embodiments, all media pipeline states must be valid before the media object commands 942 are issued. Once the pipeline state is configured and the media object commands 942 are queued, the media pipeline 924 is triggered via an execute command 944 or an equivalent execution event (e.g., register write). Subsequently, the output from the media pipeline 924 can be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.

[0228] Graphics Software Architecture

[0229] Figure 29 An exemplary graphics software architecture for a data processing system 1000 is shown, according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.

[0230] In some embodiments, the 3D graphics application 1010 includes one or more shader programs, which include shader instructions 1012. The shader language instructions may be in a high-level shader language, such as High-Level Shader Language (HLSL), or OpenGL Shading Language (GLSL). The application also includes executable instructions 1014 in machine language suitable for execution by the general-purpose processor cores 1034. The application also includes graphics objects 1016 defined by vertex data.

[0231] In some embodiments, the operating system 1020 is an operating system from Microsoft Corporation, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a variant of the Linux kernel. The operating system 1020 may support a graphics API 1022, such as the Direct3D API, the OpenGL API, or the Vulkan API. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation may be just-in-time (JIT) compilation, or the application executable shader may be pre-compiled. In some embodiments, during the compilation of the 3D graphics application 1010, high-level shaders are compiled into low-level shaders. In some embodiments, the shader instructions 1012 are provided in an intermediate form, such as a version of the standard portable intermediate representation (SPIR) used by the Vulkan API.

[0232] In some embodiments, the user-mode graphics driver 1026 includes a back-end shader compiler 1027 that is used to convert the shader instructions 1012 into a hardware-specific representation. When the OpenGL API is in use, the shader instructions 1012 in the high-level GLSL language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses the operating system kernel-mode functionality 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.

[0233] IP Core Implementation

[0234] One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions representing various logic within the processor. When read by the machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit that may be stored on a tangible, machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model may be supplied to various consumers or manufacturing facilities that load the hardware model on a manufacturing machine for fabricating the integrated circuit. The integrated circuit may be fabricated such that the circuit performs operations described in association with any of the embodiments described herein.

[0235] Figure 30AFIG. 0 is a block diagram showing an IP core development system 1100 according to an embodiment, which can be used to fabricate an integrated circuit to perform operations. The IP core development system 1100 can be used to generate a modular, reusable design that can be incorporated into a larger design or used to build an entire integrated circuit (e.g., an SOC integrated circuit). A design facility 1130 can generate a software simulation 1110 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 can include functional simulation, behavioral simulation, and / or timing simulation. Subsequently, a register transfer level (RTL) design 1115 can be created or synthesized from the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic executed using the modeled digital signals). In addition to the RTL design 1115, lower-level designs at the logic level or transistor level can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can vary.

[0236] The RTL design 1115 or an equivalent can be further synthesized by the design facility into a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted via a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The manufacturing facility 1165 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0237] Figure 30BFIG. 0 shows a cross-sectional side view of an integrated circuit package component 1170 in accordance with some embodiments described herein. The integrated circuit package component 1170 shows an implementation of one or more processor or accelerator devices as described herein. The package component 1170 includes a plurality of hardware logic units 1172, 1174 coupled to a substrate 1180. The logic 1172, 1174 may be at least partially implemented in configurable logic or fixed function logic hardware and may include any one or more portions of a processor core, a graphics processor, or other accelerator device as described herein, including (multiple) processor cores, (multiple) graphics processors, or other accelerator devices. Each logic unit 1172, 1174 may be implemented within a semiconductor die and is coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between the logic 1172, 1174 and the substrate 1180 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, the interconnect structure 1173 may be configured to route electrical signals such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of the logic 1172, 1174. In some embodiments, the substrate 1180 is an epoxy-based laminated substrate. In other embodiments, the package substrate 1180 may include other suitable types of substrates. The package component 1170 may be connected to other electrical devices via package interconnects 1183. The package interconnects 1183 may be coupled to the surface of the substrate 1180 to route electrical signals to other electrical devices such as a motherboard, other chip sets, or multi-chip modules.

[0238] In some embodiments, the logic units 1172, 1174 are electrically coupled to a bridge 1182 that is configured to route electrical signals between the logic 1172 and the logic 1174. The bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. The bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Circuitry features may be formed on the bridge substrate to provide a chip-to-chip connection between the logic 1172 and the logic 1174.

[0239] Although two logic units 1172, 1174 and a bridge 1182 are shown, embodiments described herein may include more or fewer logic units on one or more dies. The one or more dies may be connected by zero or more bridges, as the bridge 1182 may be excluded when the logic is included on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Additionally, in other possible configurations, including three-dimensional configurations, multiple logic units, dies, and bridges may be connected together.

[0240] Exemplary System-on-Chip Integrated Circuit

[0241] Figures 31 - 3 3 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, other logic and circuitry can be included, including additional graphics processor(s) / core(s), peripheral interface controllers, or general-purpose processor cores.

[0242] Figure 31 is a block diagram showing an exemplary system-on-chip integrated circuit 1200 that can be fabricated using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, any of which can be modular IP cores from the same design facility or multiple different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an 2 S / I 2 C controller 1240. Additionally, the integrated circuit can include a display device 1245 that is coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display interface 1255. Storage can be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface can be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Some integrated circuits additionally include an embedded security engine 1270.

[0243] Figures 32A - 32B is a block diagram showing an exemplary graphics processor for use within a SoC according to embodiments described herein. Figure 32A shows an exemplary graphics processor 1310 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. Figure 32B shows an additional exemplary graphics processor 1340 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment. Figure 32A The graphics processor 1310 of is an example of a low-power graphics processor core. Figure 32B The graphics processor 1340 of is an example of a higher-performance graphics processor core. Each of the graphics processors 1310, 1340 can be Figure 31 a variant of the graphics processor 1210 of.

[0244] As Figure 32AAs shown, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A - 1315N (e.g., 1315A, 1315B, 1315C, 1315D, up to 1315N - 1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic such that the vertex processor 1305 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1315A - 1315N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. The vertex processor 1305 executes the vertex processing stage of the 3D graphics pipeline and generates primitive data and vertex data. The one or more fragment processors 1315A - 1315N use the primitive data and vertex data generated by the vertex processor 1305 to produce a frame buffer that is displayed on a display device. In one embodiment, the one or more fragment processors 1315A - 1315N are optimized to execute fragment shader programs as provided in the OpenGL API, and these fragment shader programs can be used to perform operations similar to pixel shader programs as provided in the Direct 3D API.

[0245] The graphics processor 1310 additionally includes one or more memory management units (MMUs) 1320A - 1320B, one or more caches 1325A - 1325B, and one or more circuit interconnects 1330A - 1330B. The one or more MMUs 1320A - 1320B provide virtual - to - physical address mapping for the graphics processor 1310 (including for the vertex processor 1305 and / or the one or more fragment processors 1315A - 1315N), and this virtual - to - physical address mapping can reference vertex data or image / texture data stored in memory in addition to vertex data or image / texture data stored in the one or more caches 1325A - 1325B. In one embodiment, the one or more MMUs 1320A - 1320B can be synchronized with other MMUs within the system such that each processor 1205 - 1220 can participate in a shared or unified virtual memory system, and the other MMUs within the system include one or more MMUs associated with Figure 31 one or more application processors 1205, an image processor 1215, and / or a video processor 1220. According to an embodiment, the one or more circuit interconnects 1330A - 1330B enable the graphics processor 1310 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0246] As Figure 32B shown, the graphics processor 1340 includes Figure 32AOne or more MMUs 1320A - 1320B, caches 1325A - 1325B, and circuit interconnects 1330A - 1330B of the graphics processor 1310. The graphics processor 1340 includes one or more shader cores 1355A - 1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N - 1 and 1355N), which provide a unified shader core architecture in which a single core or type of core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary depending on the embodiment and implementation. Additionally, the graphics processor 1340 includes an inter - core task manager 1345, which acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1355A - 1355N and a tiling unit 1358 for accelerating tiling operations for tile - based rendering, in which rendering operations for a scene are subdivided in image space, e.g., to take advantage of local spatial coherence within the scene or to optimize the use of internal caches.

[0247] Figures 33A - 33B Additional exemplary graphics processor logic in accordance with embodiments described herein is shown. Figure 33A A graphics core 1400 is shown, which may be included in Figure 31 the graphics processor 1210 and may be a unified shader core 1355A - 1355N as in Figure 32B . Figure 33B A highly parallel general - purpose graphics processing unit 1430 suitable for deployment on a multi - chip module is shown.

[0248] As Figure 33AAs shown, graphics core 1400 includes shared instruction cache 1402, texture units 1418, and cache memory / shared memory 1420 that are common to the execution resources within graphics core 1400. Graphics core 1400 may include multiple slices 1401A - 1401N or partitions for each core, and the graphics processor may include multiple instances of graphics core 1400. Slices 1401A - 1401N may include support logic that includes local instruction caches 1404A - 1404N, thread schedulers 1406A - 1406N, thread dispatchers 1408A - 1408N, and a set of registers 1410A. To perform logical operations, slices 1401A - 1401N may include a set of additional functional units (AFU 1412A - 1412N), floating - point units (FPU1414A - 1414N), integer arithmetic logic units (ALU 1416A - 1416N), address calculation units (ACU 1413A - 1413N), double - precision floating - point units (DPFPU 1415A - 1415N), and matrix processing units (MPU 1417A - 1417N).

[0249] Some of these computational units operate at specific precisions. For example, FPU 1414A - 1414N may perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, while DPFPU 1415A - 1415N performs double - precision (64 - bit) floating - point operations. ALU 1416A - 1416N can perform variable - precision integer operations at 8 - bit, 16 - bit, and 32 - bit precisions and can be configured for mixed - precision operations. MPU 1417A - 1417N can also be configured for mixed - precision matrix operations, including half - precision floating - point operations and 8 - bit integer operations. MPU 1417A - 1417N may perform various matrix operations to accelerate machine - learning application frameworks, including enabling support for accelerated general matrix - matrix multiplication (GEMM). AFU 1412A - 1412N may perform additional logical operations not supported by the floating - point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0250] As Figure 33BAs shown, a general-purpose processing unit (GPGPU) 1430 can be configured to enable highly parallel computing operations to be performed by an array of graphics processing units. Additionally, GPGPU 1430 can be directly linked to other instances of GPGPUs to create a multi-GPU cluster, thereby improving the training speed, especially for deep neural networks. GPGPU 1430 includes a host interface 1432 for enabling connection to a host processor. In one embodiment, host interface 1432 is a PCI Express interface. However, the host interface can also be a vendor-specific communication interface or communication fabric. GPGPU 1430 receives commands from the host processor and distributes execution threads associated with those commands to a set of compute clusters 1436A - 1436H using a global scheduler 1434. Compute clusters 1436A - 1436H share a cache memory 1438. Cache memory 1438 can act as a higher-level cache for the cache memories within compute clusters 1436A - 1436H.

[0251] GPGPU 1430 includes memories 1434A - 1434B coupled to compute clusters 1436A - 1436H via a set of memory controllers 1442A - 1442B. In embodiments, memories 1434A - 1434B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.

[0252] In one embodiment, each of compute clusters 1436A - 1436H includes a set of graphics multi-cores, such as Figure 33A graphics cores 1400, which can include multiple types of integer logic units and floating-point logic units that can perform computing operations, including those suitable for machine learning computations, within a certain precision range. For example, and in one embodiment, at least a subset of the floating-point units in each of compute clusters 1436A - 1436H can be configured to perform 16-bit or 32-bit floating-point operations, while a different subset of the floating-point units can be configured to perform 64-bit floating-point operations.

[0253] Multiple instances of GPGPU 1430 can be configured to operate as a compute cluster. The communication mechanisms used by the compute cluster for synchronization and data exchange vary across embodiments. In one embodiment, multiple instances of GPGPU 1430 communicate via host interface 1432. In one embodiment, GPGPU 1430 includes an I / O hub 1439 that couples GPGPU 1430 to a GPU link 1440 that enables direct connections to other instances of the GPGPU. In one embodiment, GPU link 1440 is coupled to a dedicated GPU-GPU bridge that implements communication and synchronization between multiple instances of GPGPU 1430. In one embodiment, GPU link 1440 is coupled to a high-speed interconnect to transfer and receive data to and from other GPGPUs or parallel processors. In one embodiment, multiple instances of GPGPU 1430 are located in separate data processing systems and communicate via a network device that can be accessed via host interface 1432. In one embodiment, in addition to or in place of host interface 1432, GPU link 1440 can be configured to enable a connection to a host processor.

[0254] Although the illustrated configurations of GPGPU 1430 can be configured to train neural networks, one embodiment provides an alternative configuration of GPGPU 1430 that can be configured for deployment within a high-performance or low-power inference platform. In the inference configuration, GPGPU 1430 includes fewer compute clusters among compute clusters 1436A - 1436H relative to the training configuration. Additionally, the memory technology associated with memories 1434A - 1434B can differ between the inference configuration and the training configuration, with higher bandwidth memory technology dedicated to the training configuration. In one embodiment, the inference configuration of GPGPU 1430 can support inference-specific instructions. For example, the inference configuration can provide support for one or more 8-bit integer dot product instructions that are typically used during inference operations of a deployed neural network.

[0255] The above-described embodiments in combination with Figures 20 to 33B can be configured to include one or more features or aspects of an adaptive encoder as described herein, including those described in the following additional description and examples.

[0256] Additional Notes and Examples:

[0257] Example 1 may include an electronic processing system, the electronic processing system including: a processor; a memory communicatively coupled to the processor; and logic communicatively coupled to the processor for determining information related to a head-mounted device and determining one or more video coding parameters based on the information related to the head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion.

[0258] Example 2 may include the system of Example 1, wherein the logic is further configured to: adjust a quality parameter for encoding of a macroblock based on the information related to focusing.

[0259] Example 3 may include the system of Example 2, wherein the logic is further configured to: identify a focus area based on the information related to focusing; and adjust a quality parameter for encoding such that, compared to macroblocks outside the focus area, relatively higher quality is provided for macroblocks inside the focus area.

[0260] Example 4 may include the system of Example 1, wherein the logic is further configured to: determine a global motion prediction factor for encoding of a macroblock based on the information related to motion.

[0261] Example 5 may include the system of Example 4, wherein the logic is further configured to: determine a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion.

[0262] Example 6 may include the system of any one of Examples 1 to 5, wherein the logic is further configured to: encode macroblocks of a video image based on one or more determined video coding parameters.

[0263] Example 7 may include a semiconductor packaging device, including: one or more substrates; and logic coupled to the one or more substrates, wherein the logic is at least partially implemented in one or more of configurable logic and fixed-function hardware logic, and the logic coupled to the one or more substrates is configured to: determine information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; and determine one or more video coding parameters based on the information related to the head-mounted device.

[0264] Example 8 may include the device of Example 7, wherein the logic is further configured to: adjust a quality parameter for encoding of a macroblock based on the information related to focusing.

[0265] Example 9 may include the device of Example 8, wherein the logic is further configured to: identify a focus area based on the information related to focusing; and adjust a quality parameter for encoding such that, compared to macroblocks outside the focus area, relatively higher quality is provided for macroblocks inside the focus area.

[0266] Example 10 may include the apparatus of Example 7, wherein the logic is further configured to: determine a global motion predictor for encoding of a macroblock based on motion-related information.

[0267] Example 11 may include the apparatus of Example 10, wherein the logic is further configured to: determine a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from motion-related information.

[0268] Example 12 may include the apparatus of any one of Examples 7 to 11, wherein the logic is further configured to: encode a macroblock of a video image based on one or more determined video coding parameters.

[0269] Example 13 may include a method of adaptive coding, comprising: determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; and determining one or more video coding parameters based on the information related to the head-mounted device.

[0270] Example 14 may include the method of Example 13, further comprising: adjusting a quality parameter for encoding of a macroblock based on the information related to focusing.

[0271] Example 15 may include the method of Example 14, further comprising: identifying a focus region based on the information related to focusing; and adjusting the quality parameter for encoding such that a relatively higher quality is provided for macroblocks inside the focus region compared to macroblocks outside the focus region.

[0272] Example 16 may include the method of Example 13, further comprising: determining a global motion predictor for encoding of a macroblock based on the information related to motion.

[0273] Example 17 may include the method of Example 16, further comprising: determining a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion.

[0274] Example 18 may include the method of any one of Examples 13 to 17, further comprising: encoding a macroblock of a video image based on one or more determined video coding parameters.

[0275] Example 19 may include at least one computer-readable medium including an instruction set that, when executed by a computing device, causes the computing device to: determine information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; and determine one or more video coding parameters based on the information related to the head-mounted device.

[0276] Example 20 may include at least one computer-readable medium of Example 19, including a further set of instructions that, when executed by a computing device, cause the computing device to: adjust a quality parameter for encoding of a macroblock based on information related to focus.

[0277] Example 21 may include at least one computer-readable medium of Example 20, including a further set of instructions that, when executed by a computing device, cause the computing device to: identify a focus area based on information related to focus; and adjust a quality parameter for encoding to provide a relatively higher quality for macroblocks inside the focus area compared to macroblocks outside the focus area.

[0278] Example 22 may include at least one computer-readable medium of Example 19, including a further set of instructions that, when executed by a computing device, cause the computing device to: determine a global motion prediction factor for encoding of a macroblock based on information related to motion.

[0279] Example 23 may include at least one computer-readable medium of Example 22, including a further set of instructions that, when executed by a computing device, cause the computing device to: determine a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion.

[0280] Example 24 may include at least one computer-readable medium of any one of Examples 19 to 23, including a further set of instructions that, when executed by a computing device, cause the computing device to: encode macroblocks of a video image based on one or more determined video coding parameters.

[0281] Example 25 may include an adaptive encoder device, including: means for determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focus and information related to motion; and means for determining one or more video coding parameters based on the information related to the head-mounted device.

[0282] Example 26 may include the device of Example 25, further including: means for adjusting a quality parameter for encoding of a macroblock based on information related to focus.

[0283] Example 27 may include the device of Example 26, further including: means for identifying a focus area based on information related to focus; and means for adjusting a quality parameter for encoding to provide a relatively higher quality for macroblocks inside the focus area compared to macroblocks outside the focus area.

[0284] Example 28 may include the apparatus of Example 25, further comprising: means for determining a global motion predictor for encoding of a macroblock based on the motion-related information.

[0285] Example 29 may include the apparatus of Example 28, further comprising: means for determining a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the motion-related information.

[0286] Example 30 may include the apparatus of any one of Examples 25 to 29, further comprising: means for encoding a macroblock of a video image based on one or more determined video coding parameters.

[0287] Embodiments are applicable for use with all types of semiconductor integrated circuit (“IC”) chips. Examples of such IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLAs), memory chips, network chips, system-on-chip (SoC), SSD / NAND controller ASICs, and the like. Additionally, in some of the figures, signal conductors are represented by lines. Some of the lines may be different to indicate more constitutive signal paths, may have digital labels to indicate the number of constitutive signal paths, and / or may have arrows at one or more ends to indicate a primary information flow direction. However, this should not be construed in a limiting manner. Instead, such added details may be used in conjunction with one or more exemplary embodiments to facilitate easier understanding of the circuitry. Any represented signal line, whether or not having additional information, may in fact include one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, such as digital or analog lines implemented using differential pairs, fiber optic lines, and / or single-ended lines.

[0288] Example sizes / models / values / ranges may have been given, but the embodiments are not limited thereto. As manufacturing technology (e.g., lithography) matures over time, it is expected that devices of smaller sizes can be manufactured. Additionally, for simplicity of illustration and discussion and in order not to obscure certain aspects of the embodiments, well-known power / ground connections to IC chips and other components may or may not be shown in the figures. Further, various configurations may be shown in block diagram form to avoid obscuring the embodiments, and in view of the fact that the specific details of the implementation with respect to these block diagram configurations largely depend on the platform on which the embodiments are implemented, i.e., these specific details should be within the purview of those skilled in the art. In instances where specific details (e.g., circuitry) are set forth to describe exemplary embodiments, it is apparent that those skilled in the art can implement the embodiments without or with variations to these specific details. The description is thus to be regarded as illustrative rather than restrictive.

[0289] The term "coupled" may be used herein to denote any type of direct or indirect relationship between the components being discussed, and may apply to electrical, mechanical, fluidic, optical, electromagnetic, electromechanical, or other connections. Additionally, the terms "first", "second", etc. may be used herein only for convenience of discussion and do not carry a particular temporal or chronological significance, unless otherwise stated.

[0290] As used in this application and the claims, a list of items joined by the term "one or more of" may mean any combination of the listed items. For example, both the phrase "one or more of A, B, and C" and the phrase "one or more of A, B, or C" may mean A; B; C; A and B; A and C; B and C; or A, B, and C.

[0291] Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Thus, while the embodiments have been described in connection with specific examples of the embodiments, the true scope of the embodiments should not be so limited, since other modifications will become apparent to those skilled in the art after they have studied the drawings, the specification, and the appended claims.

Claims

1. An electronic processing system, comprising: a processor; a memory communicatively coupled to the processor; and logic communicatively coupled to the processor, the logic implemented in hardware, software, or a combination of hardware and software, the logic for: determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; determining one or more video coding parameters based on the information related to the head-mounted device; adjusting a quality parameter for coding of a macroblock based on the information related to focusing; identifying a focusing area based on the information related to focusing; and adjusting the quality parameter for the coding such that, compared to macroblocks outside the focusing area, relatively higher quality is provided for macroblocks inside the focusing area.

2. The system according to claim 1, characterized in that, The logic is further for: determining a global motion prediction factor for coding of a macroblock based on the information related to motion.

3. The system according to claim 2, wherein The logic is further for: determining a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion.

4. The system according to any one of claims 1 to 3, characterized in that, The logic is further for: encoding macroblocks of a video image based on one or more determined video coding parameters.

5. A semiconductor packaging device, comprising: a substrate; and logic coupled to the substrate, wherein the logic is implemented at least partially in one or more of configurable logic and fixed-function hardware logic, the logic coupled to the substrate for: determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; determining one or more video coding parameters based on the information related to the head-mounted device; adjusting a quality parameter for coding of a macroblock based on the information related to focusing; identifying a focusing area based on the information related to focusing; and adjusting the quality parameter for the coding such that, compared to macroblocks outside the focusing area, relatively higher quality is provided for macroblocks inside the focusing area.

6. The device according to claim 5, characterized in that, The logic is further for: determining a global motion prediction factor for coding of a macroblock based on the information related to motion.

7. The device according to claim 6, characterized in that, The logic is further for: determining a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the information related to motion.

8. The device according to any one of claims 5 to 7, characterized in that, The logic is further for: encoding macroblocks of a video image based on one or more determined video coding parameters.

9. An adaptive coding method, comprising: determining information related to a head-mounted device, the information related to the head-mounted device including at least one of information related to focusing and information related to motion; determining one or more video coding parameters based on the information related to the head-mounted device; adjusting a quality parameter for coding of a macroblock based on the information related to focusing; identifying a focusing area based on the information related to focusing; and Adjust the quality parameter for the encoding to provide a relatively higher quality for macroblocks inside the focus region compared to macroblocks outside the focus region.

10. The method according to claim 9, further comprising: Determine a global motion prediction factor for encoding of a macroblock based on the motion-related information.

11. The method according to claim 10, further comprising: Determine a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the motion-related information.

12. The method according to any one of claims 9 to 11, further comprising: Encode macroblocks of a video image based on one or more determined video coding parameters.

13. An adaptive video encoder device, comprising: Means for determining information related to a head-mounted device, the information related to the head-mounted device including at least one of focus-related information and motion-related information; Means for determining one or more video coding parameters based on the information related to the head-mounted device; Means for adjusting a quality parameter for encoding of a macroblock based on the focus-related information; Means for identifying a focus region based on the focus-related information; And Means for adjusting the quality parameter for the encoding to provide a relatively higher quality for macroblocks inside the focus region compared to macroblocks outside the focus region.

14. The device according to claim 13, further comprising: Means for determining a global motion prediction factor for encoding of a macroblock based on the motion-related information.

15. The device according to claim 14, further comprising: Means for determining a hierarchical motion estimation offset based on a current head position, a previous head position, and a center point from the motion-related information.

16. The device according to any one of claims 13 to 15, further comprising: Means for encoding macroblocks of a video image based on one or more determined video coding parameters.

Citation Information

Patent Citations

  • Focus adjusting virtual reality headset

    US20170160518A1