Viewpoint prediction for video streaming and adaptive resolution for point clouds in computing environment
By employing viewpoint prediction and point cloud adaptive resolution technology, the problems of expensive point cloud data rendering computation and high data rate are solved, enabling efficient 6DoF video streaming, reducing resource waste and bandwidth pressure, and ensuring high-fidelity content transmission.
Patent Information
- Application Number
- CN202511220161.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-31
- Filing Date
- 2019-06-28
- Publication Date
- 2025-12-09
AI Technical Summary
Rendering point cloud data is computationally expensive and has a high data rate. Existing technologies cannot efficiently transmit relevant data from 6DoF video, resulting in resource waste and bandwidth pressure.
Employing viewpoint prediction and point cloud adaptive resolution technology, it streams only relevant data based on user perspective knowledge, predicts viewer movement, and performs viewport-related 6DoF video streaming, reducing unnecessary data transmission.
It achieves resource savings, reduces bandwidth requirements, ensures high-fidelity content transmission, and improves computing efficiency and network transmission effectiveness.
Smart Images

Figure CN121099019A_ABST
Abstract
Description
[0001] Related applications This application relates to U.S. Patent Application No. 16 / 050,153, filed July 31, 2018, by Jill Boyce, entitled “REDUCEDRENDERING OF SIX-DEGREE OF FREEDOM VIDEO,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments described herein generally relate to computers. More specifically, embodiments for facilitating viewpoint prediction and adaptive resolution of point clouds for video streaming in a computing environment are described. Background Technology
[0003] Graphics processing has evolved from two-dimensional (2D) views to three-dimensional (3D) views. This ongoing evolution of graphics processing foreshadows the further development of 360-degree video, as well as three-degree-of-freedom (3DoF) and six-degree-of-freedom (6DoF) video, for the presentation of immersive video (IV), augmented reality (AR), and virtual reality (VR) experiences.
[0004] Six-DoF video is an emerging use case for immersive video, providing viewers with an immersive media experience where they control the viewpoint of the scene. Simpler 3DoF video (such as 360-degree or panoramic video) allows viewers to change their orientation around the X, Y, and Z axes from a fixed position, described as yaw, pitch, and rotation, while 6DoF video enables viewers to change their position by translating along the X, Y, and Z axes.
[0005] It is foreseeable that point clouds can be used to represent 6DoF video. However, rendering point cloud data is computationally expensive, making it difficult to render point cloud video containing a large number of points at high frame rates. Furthermore, point cloud data is very large, requiring significant capacity for storage or transmission.
[0006] Conventional techniques are known to transmit too much data, including irrelevant and unnecessary data, because such systems are known to operate without any knowledge of the relevance of the data being streamed or from the user's perspective. Attached Figure Description
[0007] The embodiments are illustrated in the figures as examples rather than as limitations, and similar reference numerals in the figures refer to similar elements.
[0008] Figure 1 This is a block diagram of a processing system according to one embodiment.
[0009] Figure 2It is a block diagram of an embodiment of a processor having one or more processor cores, an integrated memory controller, and an integrated graphics processor.
[0010] Figure 3 It is a block diagram of a graphics processing unit, which can be a discrete graphics processing unit or a graphics processing unit integrated with multiple processing cores.
[0011] Figure 4 This is a block diagram of a graphics processing engine for a graphics processor according to some embodiments.
[0012] Figure 5 This is a block diagram of the hardware logic of a graphics processor core according to some embodiments.
[0013] Figures 6A-6B The illustration depicts thread execution logic included in an array of processing elements employed in a graphics processor core, according to some embodiments.
[0014] Figure 7 This is a block diagram illustrating a graphics processor instruction format according to some embodiments.
[0015] Figure 8 This is a block diagram of another embodiment of a graphics processor.
[0016] Figure 9A This is a block diagram illustrating a graphics processor command format according to an embodiment.
[0017] Figure 9B This is a block diagram illustrating a sequence of graphics processor commands according to an embodiment.
[0018] Figure 10 The illustration shows a sample graphical software architecture of a data processing system according to some embodiments.
[0019] Figure 11A This is a block diagram illustrating an IP core development system according to an embodiment, which can be used to manufacture integrated circuits for performing operations.
[0020] Figure 11B The illustration shows a cross-sectional side view of an integrated circuit package assembly according to some embodiments.
[0021] Figure 12 This is a block diagram illustrating an exemplary system-on-a-chip integrated circuit that can be fabricated using one or more IP cores according to an embodiment.
[0022] Figures 13A-13B This is a block diagram illustrating an exemplary graphics processor used within a system-on-a-chip (SoC) according to embodiments described herein.
[0023] Figures 14A-14BAdditional exemplary graphics processor logic according to embodiments described herein is illustrated.
[0024] Figure 15A The illustrations depict various forms of immersive video.
[0025] Figure 15B The illustration shows the image projection and texture plane used for immersive video.
[0026] Figure 16 The diagram illustrates a client-server system through which immersive video content can be generated and encoded by server infrastructure for delivery to one or more client devices.
[0027] Figures 17A-17B The diagram illustrates a system used for encoding and decoding 3DoF+ content.
[0028] Figures 18A-18B The diagram illustrates a system for encoding and decoding 6DoF content using textured geometry data.
[0029] Figures 19A-19B The diagram illustrates a system for encoding and decoding 6DoF content using point cloud data.
[0030] Figure 20 The illustration shows a computing device for a managed prediction and correlation mechanism according to one embodiment.
[0031] Figure 21 The illustration shows a prediction and correlation mechanism according to one embodiment.
[0032] Figure 22 The illustration shows a transaction sequence for adaptive resolution of a point cloud according to one embodiment.
[0033] Figure 23A The diagram illustrates the transaction sequence used to perform client-side viewport prediction.
[0034] Figure 23B The diagram illustrates the transaction sequence used to perform server-side viewport prediction.
[0035] Figure 23C The diagram illustrates the transaction sequence used for unpacking and viewport selection.
[0036] Figure 24A The illustration shows a transaction sequence of a video-based method for an immersive media user viewport according to one embodiment.
[0037] Figure 24B The illustration shows an encoding system for a video-based method for an immersive media user viewport according to one embodiment.
[0038] Figure 24CThe illustration shows a decoding system for a video-based method for an immersive media user viewport according to one embodiment.
[0039] Figure 25 The illustration shows a transaction sequence of a point cloud-based method for an immersive media user viewport according to one embodiment.
[0040] Figure 26 The illustration depicts a system for viewport-dependent immersive media streaming according to one embodiment. Detailed Implementation
[0041] Numerous specific details are set forth in the following description. However, as described herein, embodiments can be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0042] The embodiments provide a novel technique for adaptive resolution of point clouds, viewpoint prediction that ensures only relevant data is streamed, prediction of viewer motion for subsequent frames, and viewport-dependent 6DoF video streaming. For example, in one embodiment, the novel technique provides knowledge of the user's perspective to effectively stream only relevant data, which allows for resource savings, such as reducing user bandwidth, and sending content with the highest possible fidelity based on throughput limitations.
[0043] As is expected, terms such as “request,” “query,” “job,” “work,” “work item,” and “workload” can be used interchangeably throughout this document. Similarly, “application” or “agent” can refer to or include applications through application programming interfaces (APIs) such as free rendering APIs, open graphics libraries, etc. Open Computing Language The term "dispatch" refers to computer programs, software applications, games, workstation applications, etc., provided by [various entities]. "Dispatch" can be interchangeably referred to as "unit of work" or "drawing," and similarly, "application" can be interchangeably referred to as "workflow" or simply "agent." For example, a workload such as a 3D game workload can contain and dispatch any number and type of "frames," where each frame can represent an image (e.g., a sailboat, a face). Additionally, each frame can contain and dispatch any number and type of unit of work, where each unit of work can represent a portion of the image (e.g., a sailboat, a face) represented by its corresponding frame (e.g., the mast of a sailboat, the forehead of a face). However, for consistency, throughout this document, each item may be referred to by a single term (e.g., "dispatch," "agent," etc.).
[0044] In some embodiments, terms like "display" and "display surface" can be used interchangeably to refer to the visible portion of a display device, while the remainder of the display device may be embedded within a computing device (such as a smartphone, wearable device, etc.). It is anticipated and should be noted that the embodiments are not limited to any specific computing device, software application, hardware component, display device, display screen or surface, protocol, standard, etc. For example, the embodiments can be applied to and used with any number and type of real-time applications on any number and type of computers (such as desktop computers, laptop computers, tablet computers, smartphones, head-mounted displays, and other wearable devices, etc.). Furthermore, for example, scenarios for rendering effective performance using this novel technique can vary from simple scenarios (e.g., desktop compositing) to complex scenarios (e.g., 3D games, augmented reality applications, etc.).
[0045] System Overview Figure 1 This is a block diagram of a processing system 100 according to one embodiment. In various embodiments, system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single-processor desktop system, a multiprocessor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, system 100 is a processing platform integrated within a system-on-a-chip (SoC) integrated circuit for use in mobile, handheld, or embedded devices.
[0046] In one embodiment, system 100 may include, or be incorporated into, a server-based gaming platform, a game console (including game and media consoles, mobile game consoles, handheld game consoles, or online game consoles). In some embodiments, system 100 is a mobile phone, smartphone, tablet computing device, or mobile internet device. Processing system 100 may also include, coupled to, or integrated with a wearable device (such as a smartwatch, smart glasses, augmented reality, or virtual reality device). In some embodiments, processing system 100 is a television or set-top box device having one or more processors 102 and a graphical interface generated by one or more graphics processors 108.
[0047] In some embodiments, one or more processors 102 each include one or more processor cores 107 to process instructions that, when executed, perform operations on the system and user software. In some embodiments, each of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, the instruction set 109 may facilitate Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation via Very Long Instruction Word (VLIW). Multiple processor cores 107 may each process a different instruction set 109, which may include instructions that facilitate the emulation of other instruction sets. Processor cores 107 may also include other processing means, such as digital signal processors (DSPs).
[0048] In some embodiments, processor 102 includes cache memory 104. Depending on the architecture, processor 102 may have a single internal cache or multiple levels of internal cache. In some embodiments, the cache memory is shared among various components of processor 102. In some embodiments, processor 102 also uses external caches (e.g., level 3 (L3) cache or last level cache (LLC)) (not shown), which may be shared among processor cores 107 using known cache coherence techniques. Register file 106 is also included in processor 102, which may include different types of registers (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers) for storing different types of data. Some registers may be general-purpose registers, while others may be design-specific to processor 102.
[0049] In some embodiments, processor(s) 102 is coupled to one or more interface buses 110 to transmit communication signals, such as address, data, or control signals, between processor(s) 102 and other components in system(s) 100. In one embodiment, interface bus 110 may be a processor bus, such as a version of a Direct Media Interface (DMI) bus. However, the processor bus is not limited to a DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In one embodiment, processor(s) 102 includes an integrated memory controller 116 and a platform controller hub (PCH) 130. The memory controller 116 facilitates communication between memory devices and other components of system(s) 100, while the platform controller hub (PCH) 130 provides connectivity to I / O devices via local I / O buses.
[0050] Memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase-change memory device, or some other memory device with suitable performance for acting as process memory. In one embodiment, memory device 120 can operate as system memory of system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. Memory controller 116 is also coupled to an optional external graphics processor 112, which can communicate with one or more graphics processors 108 in processors 102 to perform graphics and media operations. In some embodiments, display device 111 can be connected to processor(s) 102. Display device 111 can be one or more of an internal display device (such as in a mobile electronic device or laptop device) or an external display device attached via a display interface (e.g., a display port, etc.). In one embodiment, display device 111 can be a head-mounted display (HMD), such as a stereoscopic display device for use in virtual reality (VR) or augmented reality (AR) applications.
[0051] In some embodiments, the platform controller hub 130 enables peripherals to connect to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripherals include, but are not limited to, an audio controller 146, a network controller 134, a firmware interface 128, a wireless transceiver 126, a touch sensor 125, and a data storage device 124 (e.g., a hard disk drive, flash memory, etc.). The data storage device 124 can be connected via a storage interface (e.g., SATA) or via a peripheral bus (such as a peripheral component interconnect bus (e.g., PCI, PCI Express)). The touch sensor 125 can include a touchscreen sensor, a pressure sensor, or a fingerprint sensor. The wireless transceiver 126 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or LTE transceiver. The firmware interface 128 enables communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). The network controller 134 enables network connectivity to a wired network. In some embodiments, a high-performance network controller (not shown) is coupled to the interface bus 110. In one embodiment, the audio controller 146 is a multi-channel high-definition audio controller. In one embodiment, the system 100 includes an optional legacy I / O controller 140 for coupling legacy (e.g., a Personal System 2 (PS / 2)) devices to the system. The platform controller hub 130 can also be connected to one or more Universal Serial Bus (USB) controllers 142 to connect input devices, such as a keyboard and mouse combination 143, a camera 144, or other USB input devices.
[0052] It will be appreciated that the system 100 shown is exemplary and not limiting, and other types of data processing systems with different configurations may also be used. For example, instances of the memory controller 116 and platform controller hub 130 may be integrated into a discrete external graphics processor, such as external graphics processor 112. In one embodiment, the platform controller hub 130 and / or memory controller 160 may be external to one or more processors 102. For example, system 100 may include an external memory controller 116 and platform controller hub 130, which may be configured as a memory controller hub and peripheral controller hub within a system chipset that communicate with one or more processors 102.
[0053] Figure 2 This is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A-202N, an integrated memory controller 214, and an integrated graphics processor 208. The elements have the same reference numerals (or names) as any other figures herein. Figure 2 Those components can operate or function in any manner similar to, but not limited to, those described elsewhere herein. Processor 200 may include additional cores, up to and including additional cores 202N, indicated by dashed boxes. Each processor core 202A-202N includes one or more internal cache units 204A-204N. In some embodiments, each processor core may also use one or more shared cache units 206.
[0054] Internal cache units 204A-204N and shared cache unit 206 represent cache memory hierarchies within processor 200. A cache memory hierarchy may contain at least one level of instruction and data cache within each processor core, and one or more levels of shared intermediate cache (such as L2, L3, L4, or other levels of cache), wherein the highest-level cache preceding external memory is classified as LLC. In some embodiments, cache coherence logic maintains coherence between the various cache units 206 and 204A-204N.
[0055] In some embodiments, the processor 200 may further include a set of one or more bus controller units 216 and a system agent core 210. The one or more bus controller units 216 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. The system agent core 210 provides management functionality for various processor components. In some embodiments, the system agent core 210 includes one or more integrated memory controllers 214 to manage access to various external memory devices (not shown).
[0056] In some embodiments, one or more of the processor cores 202A-202N include support for simultaneous multithreading. In such embodiments, the system agent core 210 includes components for coordinating and operating the cores 202A-202N during multithreaded processing. The system agent core 210 may further include a power control unit (PCU) containing components and logic for regulating the power states of the graphics processor 208 and the processor cores 202A-202N.
[0057] In some embodiments, processor 200 further includes a graphics processor 208 that performs graphics processing operations. In some embodiments, graphics processor 208 is coupled to the shared cache unit 206 and system proxy core 210, which includes one or more integrated memory controllers 214. In some embodiments, system proxy core 210 also includes a display controller 211 to drive graphics processor output to one or more coupled displays. In some embodiments, display controller 211 may also be a separate module coupled to graphics processor via at least one interconnect, or it may be integrated within graphics processor 208.
[0058] In some embodiments, ring-based interconnect units 212 are used to couple the internal components of processor 200. However, alternative interconnect units, such as point-to-point interconnects, switched interconnects, or other technologies, including those well known in the art, may be used. In some embodiments, graphics processor 208 is coupled to ring interconnect 212 via I / O link 213.
[0059] The example I / O link 213 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 218 (such as eDRAM modules). In some embodiments, each of the processor cores 202A-202N and the graphics processor 208 uses the embedded memory module 218 as a shared final-level cache.
[0060] In some embodiments, processor cores 202A-202N are homogeneous cores executing the same instruction set architecture. In another embodiment, processor cores 202A-202N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more of processor cores 202A-202N execute a first instruction set, while at least one other core executes a subset of the first instruction set or a different instruction set. In one embodiment, processor cores 202A-202N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively higher power consumption are coupled to one or more power cores with lower power consumption. Furthermore, processor 200 can be implemented on one or more chips, or implemented as a SoC integrated circuit having the illustrated components and other components.
[0061] Figure 3 This is a block diagram of a graphics processor 300, which may be a discrete graphics processing unit or a graphics processor integrated with multiple processing cores. In some embodiments, the graphics processor communicates via a memory-mapped I / O interface to registers on the graphics processor and using commands placed in processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing memory. The memory interface 314 may be an interface to local memory, one or more internal caches, one or more shared external caches, and / or system memory.
[0062] In some embodiments, the graphics processor 300 further includes a display controller 302 to drive display output data to the display device 320. The display controller 302 includes hardware for compositing multiple layers of user interface elements or video and one or more overlay planes of the display. The display device 320 may be an internal or external display device. In one embodiment, the display device 320 is a head-mounted display device, such as a virtual reality (VR) display device or an augmented reality (AR) display device. In some embodiments, the graphics processor 300 includes a video codec engine 306 to encode media to one or more media encoding formats, decode media from one or more media encoding formats, or perform code conversion between one or more media encoding formats, including but not limited to Moving Picture Experts Group (MPEG) formats (such as MPEG-2), Advanced Video Decoding (AVC) formats (such as H.264 / MPEG-4 AVC), and SMPTE 421M / VC-1 and Joint Picture Experts Group (JPEG) formats (such as JPEG) and Motion JPEG (MJPEG) formats.
[0063] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 to perform two-dimensional (2D) rasterizer operations, such as bit boundary block transfer. However, in one embodiment, 2D graphics operations are performed using one or more components of a graphics processing engine (GPE) 310. In some embodiments, the GPE 310 is a computational engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0064] In some embodiments, GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering 3D images and scenes using processing functions that act on 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed-function elements that perform various tasks within the elements and / or generate execution threads to the 3D / media subsystem 315. While the 3D pipeline 312 can be used to perform media operations, embodiments of GPE 310 also include a media pipeline 316 specifically for performing media operations such as video post-processing and image enhancement.
[0065] In some embodiments, the media pipeline 316 includes fixed-function or programmable logic units to perform one or more dedicated media operations, such as video decoding acceleration, video deinterleaving, and video encoding acceleration, in place of or on behalf of the video codec engine 306. In some embodiments, the media pipeline 316 further includes a thread generation unit to generate threads for execution on the 3D / media subsystem 315. The generated threads perform computations of media operations on one or more graphics execution units contained in the 3D / media subsystem 315.
[0066] In some embodiments, the 3D / media subsystem 315 includes logic for executing threads generated by the 3D pipeline 312 and the media pipeline 316. In one embodiment, the pipelines send thread execution requests to the 3D media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching various requests to available thread execution resources. Execution resources include an array of graphics execution units that process 3D and media threads. In some embodiments, the 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory, including registers and addressable memory for sharing data between threads and storing output data.
[0067] Graphics processing engine Figure 4 This is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, the graphics processing engine (GPE) 410 is... Figure 3 The version of GPE 310 shown is illustrated. Elements having the same reference numerals (or names) as any other figures in this document are also included. Figure 4 The components can operate or function in any manner similar to, but not limited to, those described elsewhere herein. For example, illustrations show... Figure 3The 3D pipeline 312 and media pipeline 316 are included. The media pipeline 316 is optional in some embodiments of the GPE 410 and may not be explicitly included within the GPE 410. For example, and in at least one embodiment, a separate media and / or image processor is coupled to the GPE 410.
[0068] In some embodiments, GPE 410 is coupled to or includes command streamer 403, which provides a command stream to 3D pipeline 312 and / or media pipeline 316. In some embodiments, command streamer 403 is coupled to memory, which may be system memory, or one or more of internal cache memory and shared cache memory. In some embodiments, command streamer 403 receives commands from memory and sends the commands to 3D pipeline 312 and / or media pipeline 316. Commands are instructions fetched from a ring buffer that stores commands for 3D pipeline 312 and media pipeline 316. In one embodiment, the ring buffer may further include a command buffer that stores multiple commands in batches. Commands for 3D pipeline 312 may also include references to data stored in memory, such as, but not limited to, vertex and geometry data for 3D pipeline 312 and / or image data and memory objects for media pipeline 316. The 3D pipeline 312 and the media pipeline 316 process commands and data by performing operations via logic within their respective pipelines or by dispatching one or more execution threads to the graphics core array 414. In one embodiment, the graphics core array 414 comprises one or more blocks of graphics cores (e.g., one or more graphics cores 415A, one or more graphics cores 415B), each block containing one or more graphics cores. Each graphics core contains a set of graphics execution resources, which includes general and graphics-specific execution logic for performing graphics and computational operations, as well as fixed-function texture processing and / or machine learning and artificial intelligence acceleration logic.
[0069] In various embodiments, the 3D pipeline 312 includes fixed functions and programmable logic to process one or more shader programs, such as vertex shaders, geometry shaders, pixel shaders, fragment shaders, compute shaders, or other shader programs, by processing instructions and dispatching execution threads to the graphics core array 414. The graphics core array 414 provides a unified block of execution resources for processing these shader programs. The multipurpose execution logic (e.g., execution units) of the graphics core(s)(s)(s)415A-414B within the graphics core array 414 includes support for various 3D API shader languages and is capable of executing multiple concurrent threads associated with multiple threads.
[0070] In some embodiments, the graphics core array 414 further includes execution logic for performing media functions (such as video and / or image processing). In one embodiment, the execution unit further includes general-purpose logic programmable to perform parallel general-purpose computational operations in addition to graphics processing operations. The general-purpose logic can be coupled with… Figure 2 The core 202A-202N or Figure 1 The general logic within one or more processor cores 107 performs processing operations in parallel or cooperatively.
[0071] Output data generated by threads executing on the graphics core array 414 can be output to memory in a unified return buffer (URB) 418. URB 418 can store data for multithreading. In some embodiments, URB 418 can be used to send data between different threads executing on the graphics core array 414. In some embodiments, URB 418 can also be used for synchronization between fixed-function logic within shared-function logic 420 and threads on the graphics core array.
[0072] In some embodiments, the graphics core array 414 is scalable, such that the array contains a variable number of graphics cores, each having a variable number of execution units based on the target power and performance level of the GPE 410. In one embodiment, the execution resources are dynamically scalable, such that the execution resources can be enabled or disabled as needed.
[0073] The graphics core array 414 is coupled to shared functional logic 420, which includes multiple resources shared among the graphics cores contained within the array. The shared functions within the shared functional logic 420 are hardware logic units that provide dedicated supplemental functionality to the graphics core array 414. In various embodiments, the shared functional logic 420 includes, but is not limited to, sampler 421, math 422, and inter-thread communication (ITC) 423 logic. Furthermore, some embodiments implement one or more caches 425 within the shared functional logic 420.
[0074] To implement shared functionality, the requirement for a given dedicated function is insufficient to be included within the graphics core array 414. Instead, a single instantiation of that dedicated function is implemented as a separate entity within shared function logic 420 and shared among execution resources within the graphics core array 414. The exact set of functions shared between and included within the graphics core array 414 varies across embodiments. In some embodiments, a specific shared function within shared function logic 420, which is widely used by the graphics core array 414, may be included within shared function logic 416 within the graphics core array 414. In various embodiments, shared function logic 416 within the graphics core array 414 can include some or all of the logic within shared function logic 420. In one embodiment, all logic elements within shared function logic 420 may be replicated within shared function logic 416 of the graphics core array 414. In one embodiment, shared function logic 420 is excluded to favor shared function logic 416 within the graphics core array 414.
[0075] Figure 5 This is a block diagram of the hardware logic of a graphics processor core 500 according to some embodiments described herein. Elements have the same reference numerals (or names) as those in any other figures herein. Figure 5 The components can operate or function in any manner similar to, but not limited to, those described elsewhere herein. In some embodiments, the illustrated graphics processor core 500 is included... Figure 4 Within the graphics core array 414. A graphics processor core 500 (sometimes referred to as a core slice) can be one or more graphics cores within a modular graphics processor. A graphics processor core 500 is an example of a graphics core slice, and a graphics processor as described herein can contain multiple graphics core slices based on a target power and performance envelope. Each graphics core 500 can contain a fixed-function block 530 coupled to multiple sub-cores 501A-501F (also referred to as sub-slices), said sub-cores containing modular blocks of general-purpose and fixed-function logic.
[0076] In some embodiments, the fixed-function block 530 includes a geometry / fixed-function pipeline 536, which can be shared by all sub-cores of the graphics processor 500, for example, in a lower-performance and / or lower-power graphics processor implementation. In various embodiments, the geometry / fixed-function pipeline 536 includes a 3D fixed-function pipeline (e.g., such as...). Figure 3 and Figure 4 The 3D pipeline (312) includes a video front-end unit, a thread generator and a thread dispatcher, and a unified return buffer manager, which manages the unified return buffer, such as... Figure 4 The unified return buffer 418.
[0077] In one embodiment, fixed function block 530 further includes a graphics SoC interface 537, a graphics microcontroller 538, and a media pipeline 539. The graphics SoC interface 537 provides an interface between the graphics core 500 and other processor cores within the system-on-a-chip integrated circuit. The graphics microcontroller 538 is a programmable subprocessor configurable to manage various functions of the graphics processor 500, including thread dispatch, scheduling, and preemption. The media pipeline 539 (e.g., ...) Figure 3 and Figure 4 The media pipeline 316 contains logic that facilitates the decoding, encoding, preprocessing, and / or post-processing of multimedia data (including image and video data). The media pipeline 539 performs media operations via requests from computation or sampling logic within subcores 501-501F.
[0078] In one embodiment, SoC interface 537 enables graphics core 500 to communicate with a general-purpose application processor core (e.g., CPU) and / or other components within the SoC, including memory-level elements such as shared last-level cache memory, system RAM, and / or embedded on-chip or packaged DRAM. SoC interface 537 also enables communication with fixed-function devices within the SoC (e.g., camera imaging pipelines) and allows the use and / or implementation of global memory atoms that can be shared between graphics core 500 and the CPU within the SoC. SoC interface 537 also enables power management control of graphics core 500 and interfaces between the clock domain of graphics core 500 and other clock domains within the SoC. In one embodiment, SoC interface 537 enables the receiving of command buffers from a command streamer and a global thread dispatcher configured to provide commands and instructions to each of one or more graphics cores within the graphics processor. When a media operation is to be performed, commands and instructions can be dispatched to the media pipeline 539, or when a graphics processing operation is to be performed, commands and instructions can be dispatched to the geometry and fixed-function pipelines (e.g., geometry and fixed-function pipeline 536, geometry and fixed-function pipeline 514).
[0079] The graphics microcontroller 538 can be configured to perform various scheduling and management tasks for the graphics core 500. In one embodiment, the graphics microcontroller 538 can perform graphics and / or compute workload scheduling on various graphics parallel engines within the execution unit (EU) arrays 502A-502F and 504A-504F within sub-cores 501A-501F. In this scheduling model, host software executing on the CPU core of the SoC containing the graphics core 500 can submit a workload to one of multiple graphics processor doorbells, which invokes scheduling operations on the appropriate graphics engine. The scheduling operations include: determining which workload should run next, submitting the workload to the command streamer, preempting existing workloads running on the engine, monitoring the progress of the workload, and notifying the host software when the workload is completed. In one embodiment, the graphics microcontroller 538 can also facilitate low-power or idle states of the graphics core 500, providing the graphics core 500 with the ability to save and restore registers within the graphics core 500 across low-power state transitions, independent of the operating system and / or graphics driver software on the system.
[0080] The graphics core 500 may have more or fewer subcores than the illustrated subcores 501A-501F, up to N modular subcores. For each group of N subcores, the graphics core 500 may also include shared function logic 510, shared and / or cache memory 512, geometry / fixed function pipeline 514, and additional fixed function logic 516 for accelerating various graphics and computational processing operations. Figure 4 The shared functional logic 420 is associated with logic units (e.g., samplers, mathematical and / or inter-thread communication logic), which can be shared by the various N sub-cores within the graphics core 500. The shared and / or cache memory 512 can be the final-level cache for the group of N sub-cores 501A-501F within the graphics core 500, and can also be used as shared memory accessible by multiple sub-cores. The fixed-function block 530 can contain a geometry / fixed-function pipeline 514 instead of the geometry / fixed-function pipeline 536, and can contain the same or similar logic units.
[0081] In one embodiment, the graphics core 500 includes additional fixed-function logic 516, which can contain various fixed-function acceleration logics for use by the graphics core 500. In one embodiment, the additional fixed-function logic 516 includes an additional geometry pipeline for use in position-only shading. In position-only shading, there are two geometry pipelines: the full geometry pipeline within geometry / fixed-function pipelines 516 and 536, and a cull pipeline, which can be included within the additional fixed-function logic 516. In one embodiment, the cull pipeline is a trimmed version of the full geometry pipeline. The full pipeline and the cull pipeline can execute different instances of the same application, each with a separate context. Position-only shading can hide long cull runs of discarded triangles, allowing shading to complete earlier in some instances. For example, and in one embodiment, the pick pipeline logic within the additional fixed-function logic 516 can execute the position shader in parallel with the main application and generally generates key results faster than the full pipeline because the pick pipeline extracts and colors only the positional attributes of vertices without performing rendering and rasterization of pixels to the frame buffer. The pick pipeline can use the generated key results to compute visibility information for all triangles, regardless of whether those triangles were picked. The full pipeline (which may be referred to as the replay pipeline in this instance) can consume visibility information to skip picked triangles in order to color only the visible triangles that ultimately pass to the rasterization stage.
[0082] In one embodiment, the additional fixed-function logic 516 may also include machine learning acceleration logic (such as fixed-function matrix multiplication logic) for use in implementing optimizations for machine learning training or inference.
[0083] Each graphics subcore 501A-501F contains a set of execution resources that can be used to perform graphics, media, and computational operations in response to requests from the graphics pipeline, media pipeline, or shader program. The graphics subcores 501A-501F include multiple EU arrays 502A-502F, 504A-504F, thread dispatch and inter-thread communication (TD / IC) logic 503A-503F, 3D (e.g., texture) samplers 505A-505F, media samplers 506A-506F, shader processors 507A-507F, and shared local memory (SLM) 508A-508F. The EU arrays 502A-502F and 504A-504F each contain multiple execution units, which are general-purpose graphics processing units capable of performing floating-point and integer / fixed-point logic operations within services of graphics, media, or computational operations (including graphics, media, or computational shader programs). The TD / IC logic 503A-503F performs local thread dispatch and thread control operations for execution units within a subcore and facilitates communication between threads executing on execution units within the subcore. The 3D sampler 505A-505F can read textures or other 3D graphics-related data into memory. The 3D sampler can read texture data differently based on the configured sample state and the texture format associated with a given texture. The media sampler 506A-506F can perform similar read operations based on the type and format associated with the media data. In one embodiment, each graphics subcore 501A-501F can optionally include a unified 3D and media sampler. Threads executing on execution units within each of the subcores 501A-501F can utilize the shared local memory 508A-508F within each subcore, enabling thread execution within a thread group to utilize a common pool of on-chip memory.
[0084] Execution unit; Figures 6A-6B The illustration shows thread execution logic 600 comprising an array of processing elements employed in a graphics processor core according to an embodiment described herein. Elements have the same reference numerals (or names) as any other figures herein. Figures 6A-6B The components can operate or function in any manner similar to, but not limited to, those described elsewhere in this document. Figure 6A The diagram illustrates an overview of thread execution logic 600, which can include... Figure 5 The hardware logic of each sub-core 501A-501F is illustrated as a variant. Figure 6B The illustration shows exemplary internal details of the execution unit.
[0085] Figures 6A-6BThe illustration shows thread execution logic 600 comprising an array of processing elements employed in a graphics processor core according to an embodiment described herein. Elements have the same reference numerals (or names) as any other figures herein. Figures 6A-6B The components can operate or function in any manner similar to, but not limited to, those described elsewhere in this document. Figure 6A The diagram illustrates an overview of thread execution logic 600, which can include... Figure 5 The hardware logic of each sub-core 501A-501F is illustrated as a variant. Figure 6B The illustration shows exemplary internal details of the execution unit.
[0086] like Figure 6A As illustrated, in some embodiments, thread execution logic 600 includes a shader processor 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array comprising multiple execution units 608A-608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any one of execution units 608A, 608B, 608C, 608D through 608N-1 and 608N) based on workload computational requirements. In one embodiment, the included components are interconnected via an interconnect architecture linking each component. In some embodiments, thread execution logic 600 includes one or more connections to memory (such as system memory or cache memory) via one or more of the instruction cache 606, data port 614, sampler 610, and execution units 608A-608N. In some embodiments, each execution unit (e.g., 608A) is an independent programmable general-purpose computing unit capable of executing multiple concurrent hardware threads, processing multiple data elements in parallel for each thread. In various embodiments, the array of execution units 608A-608N can be scaled to include any number of individual execution units.
[0087] In some embodiments, execution units 608A-608N are primarily used to execute shader programs. Shader processor 602 can handle various shader programs and dispatch execution threads associated with the shader programs via thread dispatcher 604. In one embodiment, the thread dispatcher includes logic for arbitrating thread initiation requests from the graphics and media pipeline and instantiating the requested thread on one or more execution units in execution units 608A-608N. For example, a geometry pipeline can dispatch vertex, tessellation, or geometry shaders to thread execution logic for processing. In some embodiments, thread dispatcher 604 can also handle runtime thread generation requests from executing shader programs.
[0088] In some embodiments, execution units 608A-608N support instruction sets that include native support for many standard 3D graphics shader instructions, enabling the execution of shader programs from graphics libraries (such as Direct3D and OpenGL) with minimal translation. Execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general-purpose processing (e.g., computation and media shaders). Each execution unit 608A-608N is capable of multiple-shot single-instruction multiple-data (SIMD) execution, and multithreaded operation provides an efficient execution environment in the face of higher-latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. For pipelines capable of integer, single-precision, and double-precision floating-point operations, SIMD branching capabilities, logical operations, transcendental operations, and other hybrid operations, multiple shots are executed per clock cycle. While waiting for data from memory or one of the shared functions, dependency logic within the execution unit 608A-608N causes the waiting thread to sleep until the requested data has been returned. While waiting threads are sleeping, hardware resources can be dedicated to processing other threads. For example, during the delay associated with vertex shader operations, the execution unit can perform operations on pixel shaders, fragment shaders, or other types of shader programs, including different vertex shaders.
[0089] Each execution unit in the 608A-608N operates on an array of data elements. The number of data elements is the "execution size," or the number of channels used for instructions. An execution channel is a logical unit used for flow control, masking, and data element access within an instruction. The number of channels may be independent of the number of floating-point units (FPUs) or physical arithmetic logic units (ALUs) used for a specific graphics processor. In some embodiments, the execution units 608A-608N support both integer and floating-point data types.
[0090] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as compact data types, and the execution unit processes each element based on its data size. For example, when operating on a 256-bit wide vector, all 256 bits of the vector are stored in registers, and the execution unit operates on the vector as four individual 64-bit compact data elements (quad-word (QW) size data elements), eight individual 32-bit compact data elements (double-word (DW) size data elements), sixteen individual 16-bit compact data elements (word (W) size data elements), or 32 individual 8-bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible.
[0091] In one embodiment, one or more execution units can be combined into fused execution units 609A-609N having thread control logic (607A-607N) common to the fused EUs. Multiple EUs can be fused into EU groups. Each EU in a fused EU group can be configured to execute a separate SIMD hardware thread. According to embodiments, the number of EUs in a fused EU group can vary. Furthermore, various SIMD widths can be executed per EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 609A-609N contains at least two execution units. For example, fused execution unit 609A includes a first EU 608A, a second EU 608B, and thread control logic 607A, which is common to both the first EU 608A and the second EU 608B. Thread control logic 607A controls the threads executing on the fused graphics execution unit 609A, allowing each EU within the fused execution units 609A-609N to execute using a common instruction pointer register.
[0092] One or more internal instruction caches (e.g., 606) are included in thread execution logic 600 to cache thread instructions for the execution unit. In some embodiments, one or more data caches (e.g., 612) are included to cache thread data during thread execution. In some embodiments, a sampler 610 is included to provide texture sampling for 3D operations and media sampling for media operations. In some embodiments, sampler 610 includes dedicated texture or media sampling functionality to process texture or media data during a sampling process prior to providing sampled data to the execution unit.
[0093] During execution, the graphics and media pipeline sends thread initiation requests to thread execution logic 600 via thread creation and dispatch logic. Once a set of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 602 is invoked to further compute output information and write the results to an output interface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some embodiments, the pixel shader or fragment shader computes values of various vertex attributes to be interpolated across the rasterized objects. In some embodiments, the pixel processor logic within shader processor 602 then executes the pixel or fragment shader program provided by the Application Programming Interface (API). To execute the shader program, shader processor 602 dispatches threads to execution units (e.g., 608A) via thread dispatcher 604. In some embodiments, shader processor 602 uses texture sampling logic in sampler 610 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and input geometry data compute pixel color data for each geometric fragment, or discard one or more pixels from further processing.
[0094] In some embodiments, data port 614 provides a memory access mechanism for thread execution logic 600 to output processed data to memory for further processing on the graphics processor output pipeline. In some embodiments, data port 614 includes or is coupled to one or more cache memories (e.g., data cache 612) to cache data for memory access via the data port.
[0095] like Figure 6B As illustrated, the graphics execution unit 608 may include an instruction fetch unit 637, a general-purpose register file array (GRF) 624, an architecture register file array (ARF) 626, a thread arbiter 622, a send unit 630, a branch unit 632, a set of SIMD floating-point units (FPUs) 634, and in one embodiment, a set of dedicated integer SIMD ALUs 635. The GRF 624 and ARF 626 contain the general-purpose register file and architecture register file associated with each concurrent hardware thread that is valid in the graphics execution unit 608. In one embodiment, the architecture state of each thread is maintained in the ARF 626, while data used during thread execution is stored in the GRF 624. The execution state of each thread, including the instruction pointer for each thread, can be held in thread-specific registers in the ARF 626.
[0096] In one embodiment, the graphics execution unit 608 has an architecture that is a combination of simultaneous multithreading (SMT) and fine-grained interleaved multithreading (IMT). This architecture features a modular configuration that can be fine-tuned at design time based on a target number of simultaneous threads and the number of registers per execution unit, wherein execution unit resources are partitioned across logic used to execute multiple simultaneous threads.
[0097] In one embodiment, the graphics execution unit 608 can issue multiple instructions, which can be distinct instructions. The thread arbiter 622 of the graphics execution unit thread 608 can dispatch instructions to one of the sending unit 630, the branching unit 642, or one or more SIMD FPUs 634 for execution. Each execution thread can access 128 general-purpose registers within the GRF 624, each register capable of storing 32 bytes, accessible as a 32-bit SIMD 8-element vector. In one embodiment, each execution unit thread can access 4 kilobytes within the GRF 624, although embodiments are not limited to this, and more or fewer register resources may be provided in other embodiments. In one embodiment, up to seven threads can execute concurrently, although the number of threads per execution unit can vary depending on the embodiment. In an embodiment where seven threads can access 4 kilobytes, the GRF 624 can store a total of 28 kilobytes. Flexible addressing modes allow registers to be addressed together to efficiently construct wider registers, or represent rectangular block data structures spanning across.
[0098] In one embodiment, memory operations, sampler operations, and other long-latency system communications are dispatched via a "send" instruction executed by message sending unit 630. In one embodiment, branch instructions are dispatched to dedicated branch unit 632 to facilitate SIMD divergence and eventual convergence.
[0099] In one embodiment, the graphics execution unit 608 includes one or more SIMD floating-point units (FPUs) 634 to perform floating-point operations. In one embodiment, the FPU(s) 634 also supports integer computation. In one embodiment, the FPU(s) 634 can perform up to M 32-bit floating-point (or integer) operations in SIMD, or up to 2M 16-bit integer or 16-bit floating-point operations in SIMD. In one embodiment, at least one of the FPUs provides extended mathematical capabilities to support high-throughput transcendental mathematical functions and double-precision 64-bit floating-point operations. In some embodiments, a set of 8-bit integer SIMD ALUs 635 are also present and can be specifically optimized to perform operations associated with machine learning computations.
[0100] In one embodiment, an array of multiple instances of the graphics execution unit 608 can be instantiated in a graphics subcore group (e.g., a sub-slice). For scalability, the product architect can select the exact number of execution units in each subcore group. In one embodiment, the execution unit 608 can execute instructions across multiple execution channels. In another embodiment, each thread executing on the graphics execution unit 608 executes on a different channel.
[0101] Figure 7 This is a block diagram illustrating a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution unit supports an instruction set having multiple instruction formats. Solid lines illustrate components generally included in the execution unit instructions, while dashed lines contain optional components or those included only in subsets of the instructions. In some embodiments, the illustrated and described instruction format 700 is macro instructions, as they are instructions provided to the execution unit, in contrast to micro-operations generated by instruction decoding once the instructions are processed.
[0102] In some embodiments, the graphics processor execution unit natively supports instructions in 128-bit instruction format 710. A 64-bit compact instruction format 730 can be used for some instructions based on the selected instruction, instruction options, and number of operands. The native 128-bit instruction format 710 provides access to all instruction options, while some options and operations are constrained to the 64-bit format 730. The native instructions available in the 64-bit format 730 vary depending on the embodiment. In some embodiments, a set of index values in the index field 713 is used to compact the instructions. The execution unit hardware references a set of compact tables based on the index values and uses the compact table output to reconstruct the native instructions in 128-bit instruction format 710.
[0103] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a simultaneous add operation across each color channel representing a texture element or picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control over certain execution options, such as channel selection (e.g., prediction) and data channel order (e.g., scrambling). For instructions in the 128-bit instruction format 710, the execution size field 716 limits the number of data channels that will be executed in parallel. In some embodiments, the execution size field 716 is not available for use with the 64-bit compact instruction format 730.
[0104] Some execution unit instructions have up to three operands, including two source operands, src0720, src1722, and a destination 718. In some embodiments, the execution unit supports dual-destination instructions, where one of the destinations is implied. Data manipulation instructions can have a third source operand (e.g., SRC2724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction can be an immediate (e.g., hard-coded) value passed by the instruction.
[0105] In some embodiments, the 128-bit instruction format 710 includes, for example, an access / address mode field 726 that specifies whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits in the instruction.
[0106] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726, which specifies the address mode and / or access mode for the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes that include 16-byte aligned access modes and 1-byte aligned access modes, wherein the byte alignment of the access mode determines the access alignment of the instruction operands. For example, in a first mode, the instruction may use byte-aligned addressing for both source and destination operands, while in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands.
[0107] In one embodiment, the addressing mode portion of the access / address mode field 726 determines whether the instruction uses direct or indirect addressing. When using direct register addressing mode, bits in the instruction directly provide the register addresses of one or more operands. When using indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address immediate field and address register value in the instruction.
[0108] In some embodiments, instructions are grouped based on the opcode (712) bit field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the opcode type. The precise opcode grouping shown is merely an example. In some embodiments, the move and logic opcode group 742 contains data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares five most significant bits (MSBs), where move (mov) instructions are in the form of 0000xxxxb, and logic instructions are in the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jump (jmp)) contains instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 contains a mixture of instructions, including synchronous instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). Parallel math instruction group 748 contains component-wise arithmetic instructions (e.g., addition, multiplication (mul)) in the form of 0100xxxxb (e.g., 0x40). Parallel math group 748 performs arithmetic operations in parallel across data channels. Vector math group 750 contains arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). Vector math group performs arithmetic such as calculating the dot product on vector operands.
[0109] Graphics Pipeline Figure 8 This is a block diagram of another embodiment of a graphics processor 800. Elements have the same reference numerals (or names) as any other figures herein. Figure 8 The components can operate or function in any manner similar to, but not limited to, those described elsewhere in this document.
[0110] In some embodiments, the graphics processor 800 includes a geometry pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a rendering output pipeline 870. In some embodiments, the graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or by commands issued to the graphics processor 800 via a ring interconnect 802. In some embodiments, the ring interconnect 802 couples the graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 802 are interpreted by a command streamer 803, which provides instructions to the various components of the media pipeline 830 or the geometry pipeline 820.
[0111] In some embodiments, command streamer 803 directs the operation of vertex extractor 805, which reads vertex data from memory and executes vertex processing commands provided by command streamer 803. In some embodiments, vertex extractor 805 provides vertex data to vertex shader 807, which performs coordinated spatial transformation and lighting operations on each vertex. In some embodiments, vertex extractor 805 and vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A-852B via thread dispatcher 831.
[0112] In some embodiments, execution units 852A-852B are vector processor arrays having an instruction set for performing graphics and media operations. In some embodiments, execution units 852A-852B have attached L1 caches 851, which are specific to each array and shared between arrays. The caches can be configured as data caches, instruction caches, or a single cache, which is partitioned into different partitions containing data and instructions.
[0113] In some embodiments, the geometry pipeline 820 includes tessellation components to perform hardware-accelerated tessellation of 3D objects. In some embodiments, a programmable shell shader 811 configures the tessellation operation. A programmable domain shader 817 provides back-end evaluation of the tessellation output. A tessellation 813 operates in the direction of the shell shader 811 and contains dedicated logic to generate a set of detailed geometric objects based on a coarse geometric model provided as input to the geometry pipeline 820. In some embodiments, if tessellation is not used, the tessellation components (e.g., shell shader 811, tessellation 813, and domain shader 817) can be bypassed.
[0114] In some embodiments, the complete geometry object can be processed by the geometry shader 819 via one or more threads dispatched to execution units 852A-852B, or it can proceed directly to the trimmer 829. In some embodiments, the geometry shader operates on the entire geometry object, rather than on vertices or vertex patches as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, if the tessellation unit is disabled, the geometry shader 819 is programmable by the geometry shader program to perform geometric tessellation.
[0115] Prior to rasterization, trimmer 829 processes vertex data. Trimmer 829 can be a programmable trimmer with trimming and geometry shader functionality or a fixed-function trimmer. In some embodiments, the rasterizer and depth testing component 873 in the render output pipeline 870 dispatch pixel shaders to convert geometry objects into per-pixel representations. In some embodiments, pixel shader logic is contained within thread execution logic 850. In some embodiments, the application can bypass the rasterizer and depth testing component 873 and access unrasterized vertex data via outgoing unit 823.
[0116] The graphics processor 800 has an interconnect bus, interconnect architecture, or some other interconnect mechanism that allows data and messages to be passed between the main components of the processor. In some embodiments, execution units 852A-852B and associated logic units (e.g., L1 cache 851, sampler 854, texture cache 858, etc.) are interconnected via data port 856 to perform memory accesses and communicate with the processor's rendering output pipeline components. In some embodiments, sampler 854, caches 851, 858, and execution units 852A-852B each have a separate memory access path. In one embodiment, texture cache 858 can also be configured as a sampler cache.
[0117] In some embodiments, the rendering output pipeline 870 includes a rasterizer and a depth testing component 873 that converts vertex-based objects into associated pixel-based representations. In some embodiments, the rasterizer logic includes a window / mask unit to perform fixed-function triangle or line rasterization. In some embodiments, associated rendering cache 878 and depth cache 879 are also available. Pixel manipulation component 877 performs pixel-based operations on the data; however, in some instances, pixel operations associated with 2D operations (e.g., bit-block image transfers with blending) are performed by the 2D engine 841, or alternatively by the display controller 843 at display time using an overlay display plane. In some embodiments, a shared L3 cache 875 is available to all graphics components, allowing data to be shared without using main system memory.
[0118] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front-end 834. In some embodiments, the video front-end 834 receives pipeline commands from a command streamer 803. In some embodiments, the media pipeline 830 includes a separate command streamer. In some embodiments, the video front-end 834 processes media commands before sending them to the media engine 837. In some embodiments, the media engine 837 includes thread generation functionality to generate threads for dispatch to thread execution logic 850 via a thread dispatcher 831.
[0119] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and coupled to the graphics processor via a ring interconnect 802 or some other interconnect bus or architecture. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 includes dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system-integrated display device, as in a laptop computer, or an external display device attached via a display device connector.
[0120] In some embodiments, the geometry pipeline 820 and media pipeline 830 may be configured to perform operations based on multiple graphics and media programming interfaces (APIs), and are not specific to any single application programming interface (API). In some embodiments, driver software for the graphics processor translates API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for OpenGL, OpenCL, and / or Vulkan graphics and computing APIs, all from the Khronos Group. In some embodiments, support may also be provided for the Direct3D library from Microsoft Corporation. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the OpenCV library, an open-source computer vision library. Future APIs with compatible 3D pipelines will also be supported if a mapping from the pipeline of future APIs to the pipeline of the graphics processor is possible.
[0121] Graphical Pipeline Programming Figure 9A This is a block diagram illustrating a graphics processor command format 900 according to some embodiments. Figure 9B This is a block diagram illustrating a graphics processor command sequence 910 according to an embodiment. Figure 9A Solid lines in the diagram represent components that are typically included in a drawing command, while dashed lines represent components that are optional or only included in a subset of the drawing commands. Figure 9A An exemplary graphics processor command format 900 includes data fields to identify the client 902, a command operation code (opcode) 904, and command data 906. Some commands also include a sub-opcode 905 and a command size 908.
[0122] In some embodiments, client 902 defines a client unit of a graphics device that processes command data. In some embodiments, a graphics processor command parser examines the client field of each command to adjust further processing of the command and routes the command data to the appropriate client unit. In some embodiments, the graphics processor client unit includes a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a client unit receives a command, it reads opcode 904 and sub-opcode 905 (if present) to determine the operation to be performed. The client unit executes the command using information in data field 906. For some commands, an explicit command size 908 is expected to specify the size of the command. In some embodiments, the command parser automatically determines the size of at least some commands based on the command opcode. In some embodiments, commands are aligned via multiple double words.
[0123] Figure 9B The flowchart illustrates an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of a graphics processor uses a version of the illustrated command sequence to establish, execute, and terminate a set of graphics operations. Sample command sequences are shown and described for illustrative purposes only, as embodiments are not limited to these specific commands or this command sequence. Moreover, commands may be issued as a batch of commands in a command sequence, such that the graphics processor will process the command sequence in a manner that occurs at least partially simultaneously.
[0124] In some embodiments, the graphics processor command sequence 910 may begin with a pipeline dump clearing command 912 to cause any active graphics pipeline to complete currently pending commands for the pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate simultaneously. Pipeline dump clearing is performed to cause any pending commands for the active graphics pipeline to complete. In response to pipeline dump clearing, the command parser for the graphics processor suspends command processing until the active painting engine completes pending operations and invalidates the relevant read cache. Optionally, any data marked as 'dirty' in the render cache may be dumped and cleared to memory. In some embodiments, pipeline dump clearing command 912 can be used for pipeline synchronization, or before putting the graphics processor into a low-power state.
[0125] In some embodiments, a pipeline selection command 913 is used when a sequence of commands requires the graphics processor to explicitly switch between pipelines. In some embodiments, the pipeline selection command 913 is required only once within the execution context before a pipeline command is issued, unless the context is issuing commands for two pipelines. In some embodiments, a pipeline dump clearing command 912 is required immediately prior to the pipeline switch via the pipeline selection command 913.
[0126] In some embodiments, pipeline control command 914 configures the graphics pipeline for operation and is used to program the 3D pipeline 922 and the media pipeline 924. In some embodiments, pipeline control command 914 configures the pipeline state for the active pipeline. In one embodiment, pipeline control command 914 is used for pipeline synchronization and to clear data from one or more caches within the active pipeline before processing a batch of commands.
[0127] In some embodiments, the return buffer status command 916 is used to configure a set of return buffers for a given pipeline for writing data. Some pipeline operations require allocating, selecting, or configuring one or more return buffers to which an operation writes intermediate data during processing. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer status 916 includes selecting the size and number of return buffers to be used for a set of pipeline operations.
[0128] The remaining commands in the command sequence differ based on the active pipeline used for the operation. Based on pipeline determination 920, the command sequence is adjusted to either 3D pipeline 922 starting at 3D pipeline state 930, or media pipeline 924 starting at media pipeline state 940.
[0129] The commands for configuring 3D pipeline states 930 include 3D state setting commands for vertex buffer states, vertex element states, constant color states, depth buffer states, and other state variables to be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the specific 3D API being used. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass certain pipeline elements if those elements are not used.
[0130] In some embodiments, the 3D primitive 932 command is used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 command are forwarded to the vertex extraction function in the graphics pipeline. The vertex extraction function uses the 3D primitive 932 command data to generate a vertex data structure. The vertex data structure is stored in one or more return buffers. In some embodiments, the 3D primitive 932 command is used to perform vertex operations on the 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution unit.
[0131] In some embodiments, the 3D pipeline 922 is triggered via the execution of command 934 or an event. In some embodiments, register writes trigger command execution. In some embodiments, execution is triggered via a "go" or "kick" command in a command sequence. In one embodiment, command execution uses pipeline synchronization command triggering to dump and clear the command sequence through the graphics pipeline. The 3D pipeline performs geometric processing on the 3D primitives. Once the operation is complete, the resulting geometry is rasterized, and the pixel engine colors the resulting pixels. Additional commands controlling pixel shading and pixel backend operations may also be included for these operations.
[0132] In some embodiments, when performing media operations, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the specific use and manner of programming the media pipeline 924 depends on the media or computational operation to be performed. Specific media decoding operations may be offloaded to the media pipeline during media decoding. In some embodiments, the media pipeline may also be bypassed, and media decoding may be performed wholly or partially using resources provided by one or more general-purpose processing cores. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor performs SIMD vector operations using computational shader programs that are not explicitly associated with the rendering of graphics primitives.
[0133] In some embodiments, media pipeline 924 is configured in a manner similar to 3D pipeline 922. A set of commands configuring media pipeline state 940 is dispatched or placed in a command queue before media object commands 942. In some embodiments, commands for media pipeline state 940 contain data configuring media pipeline elements that will be used to process media objects. This includes data configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, commands for media pipeline state 940 also include one or more pointers to "indirect" state elements containing a batch of state settings.
[0134] In some embodiments, media object command 942 provides a pointer to a media object for processing by the media pipeline. The media object contains a memory buffer with video data to be processed. In some embodiments, all media pipeline states must be valid before media object command 942 is issued. Once the pipeline states are configured and media object command 942 is queued, media pipeline 924 is triggered via execution command 944 or an equivalent execution event (e.g., register write). The output from media pipeline 924 can then be post-processed by operations provided by 3D pipeline 922 or media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a similar manner to media operations.
[0135] Graphical software architecture Figure 10 The illustration depicts an exemplary graphics software architecture for a data processing system 1000 according to some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.
[0136] In some embodiments, the 3D graphics application 1010 contains one or more shader programs that include shader instructions 1012. The shader language instructions can be in a high-level shader language, such as High-Level Shading Language (HLSL) or OpenGL Shading Language (GLSL). The application also includes executable instructions 1014 in machine language suitable for execution by a general-purpose processor core 1034. The application also includes graphics objects 1016 defined by vertex data.
[0137] In some embodiments, the operating system 1020 is from Microsoft Corporation. The operating system 1020 is a proprietary Unix-like operating system or a variant of an open-source Unix-like operating system using the Linux kernel. The operating system 1020 supports graphics APIs 1022, such as the Direct3D API, OpenGL API, or Vulkan API. When the Direct3D API is used, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. Compilation can be just-in-time (JIT) compilation, or the application can perform shader pre-compilation. In some embodiments, high-level shaders are compiled into low-level shaders during the compilation of the 3D graphics application 1010. In some embodiments, shader instructions 1012 are provided in an intermediate form, such as a version of the standard Portable Intermediate Representation (SPIR) used by the Vulkan API.
[0138] In some embodiments, the user-mode graphics driver 1026 includes a back-end shader compiler 1027 to translate shader instructions 1012 into a hardware-specific representation. When the OpenGL API is used, shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 communicates with the kernel-mode graphics driver 1029 using operating system kernel-mode functionality 1028. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.
[0139] IP core implementation One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit (such as a processor). For example, the machine-readable medium may contain instructions representing various logics within a processor. When read by a machine, the instructions enable the machine to construct logic that performs the techniques described herein. Such a representation, referred to as an "IP core," is a reusable unit of logic of an integrated circuit that can be stored on a tangible machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model may be provided to various customers or manufacturing facilities that load the hardware model onto fabrication machines that manufacture integrated circuits. The integrated circuit may be fabricated such that the circuit performs the operations described in connection with any embodiment described herein.
[0140] Figure 11AThis is a block diagram illustrating an IP core development system 1100 that can be used to manufacture integrated circuits that perform operations, according to one embodiment. The IP core development system 1100 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to construct entire integrated circuits (e.g., SOC integrated circuits). Design facility 1130 can generate software simulations 1110 of the IP core design using a high-level programming language (e.g., C / C++). The software simulation 1110 can be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 may contain functional, behavioral, and / or timing simulations. Register transfer level (RTL) designs 1115 can then be created or synthesized from the simulation model 1112. The RTL design 1115 is a behavioral abstraction of an integrated circuit that models the flow of digital signals between hardware registers (containing associated logic executed using modeled digital signals). In addition to the RTL design 1115, lower-level designs at the logic or transistor levels can also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation can be modified.
[0141] The RTL design 1115 or an equivalent can be further synthesized into a hardware model 1120 by the design facility, which may take the form of a hardware description language (HDL) or some other representation of the physical design data. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored using non-volatile memory 1140 (e.g., hard disk, flash memory, or any non-volatile storage medium) for delivery to a third-party fabrication facility 1165. Alternatively, the IP core design can be transmitted over a wired connection 1150 or a wireless connection 1160 (e.g., via the Internet). The fabrication facility 1165 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operation according to at least one embodiment described herein.
[0142] Figure 11BA cross-sectional side view of an integrated circuit package assembly 1170 according to some embodiments described herein is illustrated. The integrated circuit package assembly 1170 illustrates an implementation of one or more processor or accelerator devices as described herein. The package assembly 1170 includes multiple units of hardware logic 1172, 1174 connected to a substrate 1180. Logic 1172, 1174 may be implemented at least partially in configurable logic or fixed-function logic hardware and may include one or more portions of any of a processor core(s), a graphics processor(s), or other accelerator devices described herein. Each unit of logic 1172, 1174 may be implemented within a semiconductor die and coupled to the substrate 1180 via an interconnect structure 1173. The interconnect structure 1173 may be configured to route electrical signals between logic 1172, 1174 and the substrate 1180 and may include interconnects such as, but not limited to, bumps or pillars. In some embodiments, interconnect structure 1173 may be configured to route electrical signals, such as, for example, input / output (I / O) signals and / or power or ground signals associated with the operation of logic 1172, 1174. In some embodiments, substrate 1180 is an epoxy-based laminated substrate. In other embodiments, package substrate 1180 may comprise other suitable types of substrates. Package assembly 1170 can be connected to other electrical devices via package interconnect 1183. Package interconnect 1183 may be coupled to the surface of substrate 1180 to route electrical signals to other electrical devices, such as a motherboard, other chipsets, or multi-chip modules.
[0143] In some embodiments, cells of logic cells 1172, 1174 are electrically coupled to bridge 1182, which is configured to route electrical signals between logic cells 1172, 1174. Bridge 1182 may be a dense interconnect structure that provides routing for electrical signals. Bridge 1182 may include a bridge substrate made of glass or a suitable semiconductor material. Electrical wiring features can be formed on the bridge substrate to provide chip-to-chip connections between logic cells 1172, 1174.
[0144] Although two logic units 1172, 1174 and bridge 1182 are illustrated, the embodiments described herein may include more or fewer logic units on one or more dies. One or more dies may be connected by zero or more bridges, as bridge 1182 can be excluded when logic is contained on a single die. Alternatively, multiple dies or logic units may be connected by one or more bridges. Furthermore, multiple logic units, dies, and bridges may be connected together in other possible configurations, including three-dimensional configurations.
[0145] Demonstration System-on-Chip Integrated Circuit Figure 12-1Figure 4 illustrates exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to those illustrated, other logic and circuitry may be included, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0146] Figure 12 This is a block diagram illustrating an exemplary system-on-a-chip integrated circuit 1200 fabricated using one or more IP cores according to an embodiment. The exemplary integrated circuit 1200 includes one or more application processors 1205 (e.g., a CPU), at least one graphics processor 1210, and may further include an image processor 1215 and / or a video processor 1220, any of which can be modular IP cores from the same or more different design facilities. The integrated circuit 1200 includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I2S / I2C controller 1240. Furthermore, the integrated circuit may include a display device 1245 coupled to one or more of a High Definition Multimedia Interface (HDMI) controller 1250 and a Mobile Industry Processor Interface (MIPI) Display Interface 1255. Storage may be provided by a flash memory subsystem 1260, including flash memory and a flash memory controller. A memory interface may be provided via a memory controller 1265 for accessing SDRAM or SRAM memory devices. Some integrated circuits also include an embedded security engine 1270.
[0147] Figures 13A-13B This is a block diagram illustrating an exemplary graphics processor for use within a SoC according to embodiments described herein. Figure 13A An exemplary graphics processor 1310 of a system-on-a-chip integrated circuit that can be made using one or more IP cores according to one embodiment is illustrated. Figure 13B An additional exemplary graphics processor 1340 of a system-on-a-chip integrated circuit that can be made using one or more IP cores according to one embodiment is illustrated. Figure 13A The graphics processor 1310 is an example of a low-performance graphics processor core. Figure 13B The 1340 graphics processor is an example of a higher-performance graphics processor core. Each of the 1310 and 1340 graphics processors can be... Figure 12 A variant of the 1210 graphics processor.
[0148] like Figure 13AAs shown, the graphics processor 1310 includes a vertex processor 1305 and one or more fragment processors 1315A-1315N (e.g., 1315A, 1315B, 1315C, 1315D to 1315N-1 and 1315N). The graphics processor 1310 can execute different shader programs via separate logic, such that the vertex processor 1305 is optimized to perform operations for the vertex shader program, while the one or more fragment processors 1315A-1315N perform fragment (e.g., pixel) shading operations for fragments or vertex shader programs. The vertex processor 1305 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. The fragment processors (one or more) 1315A-1315N use the primitive and vertex data generated by the vertex processor 1305 to generate frame buffers displayed on a display device. In one embodiment, one or more fragment processors 1315A-1315N are optimized to execute fragment shader programs provided in the OpenGL API, which can be used to perform operations similar to those of pixel shader programs provided in the OpenGL API.
[0149] The graphics processor 1310 further includes one or more memory management units (MMUs) 1320A-1320B, one or more caches 1325A-1325B, and one or more circuit interconnects 1330A-1330B. The one or more MMUs 1320A-1320B provide a virtual-to-physical address mapping for the graphics processor 1310 (including vertex processors 1305 and / or one or more fragment processors 1315A-1315N), which, in addition to referencing vertex or image / texture data stored in the one or more caches 1325A-1325B, can also reference vertex or image / texture data stored in memory. In one embodiment, the one or more MMUs 1320A-1320B may be synchronized with other MMUs within the system, including... Figure 12 The one or more application processors 1205, image processor 1215, and / or video processor 1220 are associated with one or more MMUs, enabling each processor 1205-1220 to participate in a shared or unified virtual memory system. According to an embodiment, one or more circuit interconnects 1330A-1330B enable the graphics processor 1310 to interface with other IP cores within the SoC either via the SoC's internal bus or via a direct connection.
[0150] like Figure 13B As shown, the graphics processor 1340 includes Figure 13AThe graphics processor 1310 includes one or more MMUs 1320A-1320B, caches 1325A-1325B, and interconnects 1330A-1330B. The graphics processor 1340 includes one or more shader cores 1355A-1355N (e.g., 1455A, 1355B, 1355C, 1355D, 1355E, 1355F to 1355N-1 and 1355N), providing a unified shader core architecture where a single core or type or core can execute all types of programmable shader code, including shader program code implementing vertex shaders, fragment shaders, and / or compute shaders. The exact number of shader cores present can vary between embodiments and implementations. In addition, the graphics processor 1340 includes an inter-core task manager 1345 that acts as a thread dispatcher to assign execution threads to one or more shader cores 1355A-1355N and tiling units 1358 to accelerate tiling operations for slice-based rendering, wherein the rendering operations for a scene are subdivided in the image space, for example to take advantage of local spatial coherence within the scene or to optimize the use of internal caches.
[0151] Figures 14A-14B Additional exemplary graphics processor logic according to embodiments described herein is illustrated. Figure 14A The diagram shows what can be included. Figure 12 The graphics processor 1210 contains a graphics core 1400, and it can be such as Figure 13B The unified shader cores 1355A and 1355N in the system. Figure 14B The illustration shows a highly parallel general-purpose graphics processing unit 1430 suitable for deployment on a multi-chip module.
[0152] like Figure 14AAs shown, the graphics core 1400 includes a shared instruction cache 1402, texture units 1418, and cache / shared memory 1420, which are common to the execution resources within the graphics core 1400. The graphics core 1400 can contain multiple slices 1401A-1401N or partitions for each core, and the graphics processor can contain multiple instances of the graphics core 1400. Slices 1401A-1401N can contain supporting logic, including local instruction caches 1404A-1404N, thread schedulers 1406A-1406N, thread dispatchers 1408A-1408N, and a set of registers 1410A. To perform logical operations, slices 1401A-1401N can include a set of additional functional units (AFU 1412A-1412N), floating-point units (FPU 1414A-1414N), integer arithmetic logic units (ALU 1416-1416N), address calculation units (ACU 1413A-1413N), double-precision floating-point units (DPFPU1415A-1415N), and matrix processing units (MPU 1417A-1417N).
[0153] Some computational units perform operations with specific precision. For example, the FPU 1414A-1414N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, while the DPFPU 1415A-1415N performs double-precision (64-bit) floating-point operations. The ALU 1416A-1416N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. The MPU 1417A-1417N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. The MPU 1417-1417N can perform various matrix operations to accelerate machine learning application frameworks, including implementations supporting accelerated General Matrix-to-Matrix Multiplication (GEMM). The AFU 1412A-1412N can perform additional logical operations not supported by floating-point or integer units, including trigonometric operations (e.g., sine, cosine, etc.).
[0154] like Figure 14BAs shown, the General Purpose Processing Unit (GPGPU) 1430 can be configured to enable highly parallel computing operations to be performed by the graphics processing unit array. Furthermore, the GPGPU 1430 can be directly linked to other instances of the GPGPU to create multi-GPU clusters to improve the training speed, specifically for deep neural networks. The GPGPU 1430 includes a host interface 1432 that implements connectivity to the host processor. In one embodiment, the host interface 1432 is a PCI Express interface. However, the host interface can also be a vendor-specific communication interface or communication architecture. The GPGPU 1430 receives commands from the host processor and uses a global scheduler 1434 to distribute the execution threads associated with those commands to a set of compute clusters 1436A-1436H. The compute clusters 1436A-1436H share a cache memory 1438. The cache memory 1438 can act as a higher-level cache for the cache memory within the compute clusters 1436A-1436H.
[0155] The GPGPU 1430 includes memory 1434A-1434B coupled to the computing cluster 1436A-1436H via a set of memory controllers 1442A-1442B. In various embodiments, memory 1434A-1434B can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory.
[0156] In one embodiment, each of the computing clusters 1436A-1436H contains a set of graphics cores, such as... Figure 14A The graphics core 1400 can contain various types of integer and floating-point logic units that can perform computational operations at a range of precisions, including operations suitable for machine learning computations. For example, and in one embodiment, at least a subset of the floating-point units in each of the computing clusters 1436A-1436H can be configured to perform 16-bit or 32-bit floating-point operations, while different subsets of the floating-point units can be configured to perform 64-bit floating-point operations.
[0157] Multiple instances of GPGPU 1430 can be configured to operate as a computing cluster. The communication mechanisms used by the computing cluster for synchronization and data exchange vary across embodiments. In one embodiment, multiple instances of GPGPU 1430 communicate over host interface 1432. In one embodiment, GPGPU 1430 includes I / O hub 1439, which couples GPGPU 1430 to GPU links 1440 that implement direct connections to other instances of GPGPU. In one embodiment, GPU link 1440 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1430. In one embodiment, GPU link 1440 is coupled to a high-speed interconnect for sending and receiving data to and from other GPGPUs or parallel processors. In one embodiment, multiple instances of GPGPU 1430 reside in a separate data processing system and communicate via network devices accessible through host interface 1432. In one embodiment, GPU link 1440 can be configured to implement a connection to a host processor as an addition to or alternative to host interface 1432.
[0158] While the illustrated configuration of the GPGPU 1430 can be configured to train neural networks, one embodiment provides an alternative configuration of the GPGPU 1430 that can be configured for deployment within a high-performance or low-power inference platform. In the inference configuration, the GPGPU 1430 includes fewer compute clusters 1436A-1436H compared to the training configuration. Furthermore, the memory technology associated with the memories 1434A-1434B can differ between the inference and training configurations, with a higher-bandwidth memory technology dedicated to the training configuration. In one embodiment, the inference configuration of the GPGPU 1430 can support inference-specific instructions. For example, the inference configuration can provide support for one or more 8-bit integer dot product instructions, which are typically used during inference operations on the deployed neural network.
[0159] Overview of Immersive Video Figure 15A The illustrations depict various forms of immersive video. Immersive video can be presented in multiple formats, depending on the degrees of freedom available to the viewer. Degrees of freedom refer to the number of different directions an object can move in three-dimensional (3D) space. Immersive video can be viewed via head-mounted displays and includes tracking of position and orientation. Example forms of immersive video include 3DoF 1502, 3DoF+1504, and full 6DoF 1506. In addition to immersive video in full 6DoF 1506, 6DoF immersive video also includes omnidirectional 6DoF 1507 and windowed 6DoF 1508.
[0160] For videos in 3DoF 1502 (e.g., 360-degree videos), viewers can change their orientation (e.g., yaw, pitch, rotate) but not their position. For videos in 3DoF+1504, viewers can change their orientation and make minor changes to their position. For videos in 6DoF 1506, viewers can change both their orientation and position. More limited forms of 6DoF videos are also available. Omnidirectional 6DoF 1507 videos allow viewers to take multiple steps within a virtual scene. Windowed 6DoF 1508 videos allow viewers to change both their orientation and position, but viewers are confined to a limited viewing area. Increasing the available degrees of freedom in immersive videos generally involves increasing the amount of data and complexity involved in video generation, encoding, decoding, and playback.
[0161] Figure 15B The illustration depicts image projection and texture planes for immersive video. A 3D view 1510 can be generated using data from multiple cameras to display video content. Multiple projection planes 1512 can be used to generate geometric data for the video content. Multiple texture planes 1514 can be derived from the projection planes 1512 used to generate the geometric data. The texture planes 1514 can be applied to a pre-generated or generated 3D model based on a point cloud derived from the video data. The multiple projection planes 1512 can be used to generate multiple two-dimensional (2D) projections, each associated with a projection plane.
[0162] Figure 16The illustration depicts a client-server system through which immersive video content can be generated and encoded by server 1620 infrastructure for delivery to one or more client 1630 devices. The client 1630 devices can then decompress and render the immersive video content. In one embodiment, one or more server 1620 devices can contain input from one or more optical cameras 1601 with a depth sensor 1602. Parallel computing resources 1604 can decompose the video and depth data into point clouds 1605 and / or textured triangles 1606. Data for generating textured triangles 1606 can also be provided using a pre-generated 3D model 1603 of the scene. The point clouds 1605 and / or textured triangles 1606 can be compressed for delivery to one or more client devices, where they can render the content locally. In one embodiment, various compression units 1607, 1608 using various compression algorithms can compress the generated content for delivery from server 1620 to one or more client 1630 devices via a delivery medium. The decompression units 1609 and 1610 on the client device 1630 can decompress and decode the incoming bitstream into video / texture and geometric data. For example, the decompression unit 1609 can decode the compressed point cloud data and provide the decompressed point cloud data to the viewpoint interpolation unit 1611. The interpolated viewpoint data can be used to generate bitmap data 1613. The decompressed point cloud data can be provided to the geometry reconstruction unit 1612 to reconstruct the geometry data of the scene. The reconstructed geometry data can be textured using the decoded texture data (textured triangles 1614) to generate a 3D rendering 1616 for viewing by the client 1630.
[0163] Figures 17A-17B The diagram illustrates systems 1700 and 1710 for encoding and decoding 3DoF+ content. System 1700 can be implemented using the hardware and software of server 1620 infrastructure, for example, as... Figure 16 As in the example. System 1710 can be implemented by the hardware and software of client 1630, such as... Figure 16 Like in China.
[0164] like Figure 17AAs shown, system 1700 can be used to encode video data 1702 for a base view 1701 and video data 1705A-1705C for an additional view 1704. Multiple cameras can provide input data containing video data and depth data, where each frame of video data can be converted into texture. A set of reprojection 1706 and occlusion detection 1707 units can operate on the received video data and output the processed data to patch forming 1708 unit. The patch formed by patch forming 1708 unit can be provided to patch packing 1709 unit. For example, the video data 1702 for the base view 1701 can be encoded via a high-efficiency video decoding (HEVC) encoder 1703A. Variants of HEVC encoder 1703A can also be used to encode the patch video data output from patch packing 1709 unit. Metadata reconstructed from the encoded patches can be encoded by metadata encoding 1703B unit. Multiple encoded video and metadata streams can then be sent to a client device for viewing.
[0165] like Figure 17B As shown, multiple video data streams can be received, decoded, and reconstructed into immersive video by system 1710. The multiple video streams include a stream of base video, along with a stream containing packaged data for additional views. Encoded metadata is also received. In one embodiment, the multiple video streams can be decoded via an HEVC 1713A decoder. The metadata can be decoded via a metadata 1713B decoder. The decoded metadata is then used to unpack the decoded additional views via patch unpacking logic 1719. The decoded texture and depth data (video 01712, videos 1-31714A-1715C) of base view 1701 and additional view 1704 are received by a client (e.g., as shown in the image). Figure 16 The view generation logic 1718 on the client (1630) is reconstructed. The decoded video 1712, 1715A-1715C can be provided as texture and depth data to the intermediate view renderer 1714, which can be used to render an intermediate view for the head-mounted display 1711. Head-mounted display position information 1716 is provided as feedback to the intermediate view renderer 1714, which can render an updated view for the viewport of the display presented via the head-mounted display 1711.
[0166] Figures 18A-18B The diagram illustrates a system for encoding and decoding 6DoF textured geometry data. Figure 18A The 6DoF textured geometry encoding system 1800 is shown. Figure 18BA 6DoF textured geometry decoding system 1820 is illustrated. 6DoF textured geometry encoding and decoding can be used to implement variations of 6DoF immersive video, where video data is applied as texture to geometric data, allowing the rendering of new intermediate views based on the position and orientation of the head-mounted display. Data recorded by multiple cameras can be combined with 3D models, particularly for static objects.
[0167] like Figure 18A As shown, the 6DoF textured geometry encoding system 1800 can receive video data 1802 for a base view and video data 1805A-1805C for additional views. The video data 1802, 1805A-1805C contain texture and depth data that can be processed by the reprojection and occlusion detection unit 1806. The output from the reprojection and occlusion detection unit 1806 can be provided to the patch decomposition unit 1807 and the geometry image generator 1808. The output from the patch decomposition unit 1807 is provided to the patch packing unit 1809 and the auxiliary patch information compressor 1813. The auxiliary patch information (patch information) provides information about the patches used to reconstruct the video texture and depth data. The patch packing unit 1809 outputs the packed patch data to the geometry image generator 1808, the texture image generator 1810, the attribute image generator 1811, and the occupancy map compressor 1812.
[0168] The geometry image generator 1808, texture image generator 1810, and attribute image generator 1811 output data to the video compressor 1814. The geometry image generator 1808 receives input from the reprojection and occlusion detection unit 1806, the patch decomposition unit 1807, and the patch packing unit 1809, and generates geometry image data. The texture image generator 1810 receives packed patch data from the patch packing unit 1809 and video texture and depth data from the reprojection and occlusion detection unit 1806. The attribute image generator 1811 generates an attribute image based on the video texture and depth data received from the reprojection and occlusion detection unit 1806 and the patched patch data received from the patch packing unit 1809.
[0169] Occupancy map can be generated by occupancy map compressor 1812 based on packed patch data output from patch packing unit 1809. Auxiliary patch information can be generated by auxiliary patch information compressor 1813. The compressed occupancy map and auxiliary patch information data can be multiplexed together with compressed and / or encoded video data output from video compressor 1814 by multiplexer 1815 into compressed bitstream 1816. The compressed video data output from video compressor 1814 includes compressed geometric image data, texture image data, and attribute image data. Compressed bitstream 1816 can be stored or provided to client devices for decompression and viewing.
[0170] like Figure 18B As shown, the 6DoF textured geometry decoding system 1820 can be used for... Figure 18A The encoding system 1800 decodes the 6DoF content generated. The compressed bitstream 1816 is received and demultiplexed by the demultiplexer 1835 into multiple video decoded streams, occupancy maps, and auxiliary patch information. Video decoders 1834A-1834B decode / decompress the multiple video streams. Occupancy map decoder 1832 decodes / decompresses the occupancy map data. The decoded video data and occupancy map data are output by video decoders 1834A-1834B and occupancy map decoder 1832 to the unpacking unit 1829. The unpacking unit decompiles the data generated by the demultiplexer 1816 into multiple video decoded streams, occupancy maps, and auxiliary patch information. Figure 18A The video patch data packaged by the patch packing unit 1809 is unpacked. Auxiliary patch information from the auxiliary patch information decoder 1833 is provided to the occlusion filling unit 1826, which can be used to fill in the patch from occluded parts of objects that may be lost from the specific view of the video data. The corresponding video streams 1822, 1825A-1825C with texture and depth data are output from the occlusion filling unit 1826 and provided to the intermediate view renderer 1823, which can render the view to be displayed on the head-mounted display 1824 based on the position and orientation information provided by the head-mounted display 1824.
[0171] Figures 19A-19B The diagram illustrates a system used for encoding and decoding 6DoF point cloud data. Figure 19A The diagram illustrates the 6DoF point cloud encoding system 1900. Figure 19B The diagram illustrates a 6DoF point cloud decoding system 1920. Point clouds can be used to represent 6DoF video, where, for a point cloud video sequence, new point cloud frames exist at regular time intervals (e.g., 60Hz). Each point in the point cloud data frame is represented by six parameters: (X, Y, Z) geometric position and (R, G, B or Y, U, V) texture data. Figure 19A In the coding system 1900, point cloud frames are projected onto several two-dimensional (2D) planes, each corresponding to a projection angle. The projection planes can be similar to... Figure 15B The projection plane is 1512. In some implementations, six projection angles are used in the PCC standard test model, where each projection angle corresponds to an angle pointing to the center of the six faces of the cuboid that defines the object represented by the point cloud data. Although six projection angles are described, other numbers of angles may be used in different implementations.
[0172] Texture and depth 2D image patch representations are formed at each projection angle. A 2D patch image representation for a projection angle can be created by projecting only those points whose projection angle has the closest normal. In other words, a 2D patch image representation is used for points that maximize the dot product of the point normal and the plane normal. Texture patches from individual projections are combined into a single texture image, called a geometry image. The metadata representing the patches and how they are packaged into frames is described in the occupancy map and auxiliary patch information. The occupancy map metadata contains an indication of which image sample locations are empty (e.g., do not contain corresponding point cloud information). The auxiliary patch information indicates the projection plane to which the patch belongs and can be used to determine the projection plane associated with a given sample location. Texture and depth images are encoded using a standard 2D video encoder such as a High Efficiency Video Decoding (HEVC) encoder. Metadata can be compressed separately using metadata encoding logic. In the test model decoder, the texture and depth images are decoded using an HEVC video decoder. The point cloud is reconstructed using the decoded texture and depth images along with the occupancy map and auxiliary patch information metadata.
[0173] like Figure 19A As shown, the input frame of point cloud data can be decomposed into patch data. The point cloud data and the decomposed patch data can be compared with... Figure 18A The video texture and depth data are encoded in a similar manner. Input data containing point cloud frame 1906 can be provided to patch decomposition unit 1907. The input point cloud data and its decomposed patches can be processed by packing unit 1909, geometry image generator 1908, texture image generator 1910, attribute image generator 1911, occupancy map compressor 1912, and auxiliary patch information compressor 1913 using a method similar to that used by... Figure 18A The reprojection and occlusion detection unit 1806 and the patch decomposition unit 1807 process the texture depth and video data output by the processing techniques. The video compressor 1914 can encode and / or compress geometric image, texture image and attribute image data. The compressed and / or encoded video data from the video compressor 1914 is multiplexed with occupancy map and auxiliary patch information data by the multiplexer 1915 into a compressed bitstream 1916, which can be stored or transmitted for display.
[0174] Depend on Figure 19A The compressed potential energy output by the system 1900 is from Figure 19B The point cloud decoding system 1920 shown performs the decoding. For example... Figure 19BAs shown, the compressed bitstream 1916 can be multiplexed into multiple encoded / compressed video streams, occupancy map data, and auxiliary patch information. The video streams can be decoded / decompressed by the multi-stream video decoder 1934, which outputs texture and geometry data. The occupancy map and auxiliary patch information can be decompressed / decoded by the occupancy map decoder 1932 and the auxiliary patch information decoder 1933.
[0175] Then it can perform geometry reconstruction, smoothing, and texture reconstruction to reconstruct the provided... Figure 19A The point cloud data of the 6DoF point cloud encoding system 1900. The geometry reconstruction unit 1936 can reconstruct geometric information based on the geometric data decoded from the video stream of the multi-stream video decoder 1934, as well as the outputs of the occupancy map decoder 1932 and the auxiliary patch information decoder 1933. The reconstructed geometric data can be smoothed by the smoothing unit 1937. The smoothed geometric and texture image data decoded from the video stream output by the multi-stream video decoder 1934 is provided to the texture reconstruction unit 1938. The texture reconstruction unit 1938 can output the reconstructed point cloud 1939, which is provided to... Figure 19A A variant of the input point cloud frame 1926 of the 6DoF point cloud encoding system 1900.
[0176] Figure 20 The illustration shows a computing device 2000 according to an embodiment of a managed prediction and correlation mechanism 2010. The computing device 2000 represents a communication and data processing device, including (but not limited to) smart wearable devices, smartphones, virtual reality (VR) devices, head-mounted displays (HMDs), mobile computers, Internet of Things (IoT) devices, laptop computers, desktop computers, server computers, etc., and is connected to... Figure 1 The processing device 100 is similar to or the same as that used; therefore, for the sake of simplicity, clarity and ease of understanding, the above references... Figures 1-19B Many of the details described are not discussed or repeated below.
[0177] The computing device 2000 may further include (but is not limited to) autonomous machines or artificial intelligence agents, such as mechanical agents or machines, electronic agents or machines, virtual agents or machines, electromechanical agents or machines, etc. Examples of autonomous machines or artificial intelligence agents may include (but are not limited to) robots, autonomous vehicles (e.g., self-driving cars, self-flying aircraft, self-propelled ships, etc.), autonomous devices (self-operating construction vehicles, self-operating medical devices, etc.) and / or the like. Throughout this document, "computing device" may be interchangeably referred to as "autonomous machine" or "artificial intelligence agent" or simply "robot".
[0178] It is anticipated that although “autonomous vehicles” and “autonomous driving” are mentioned throughout this document, the embodiments are not limited thereto. For example, “autonomous vehicles” are not limited to automobiles, but may include any number and type of autonomous machines, such as robots, autonomous devices, home autonomous devices, etc., and any one or more tasks or operations associated with such autonomous machines may be referred to interchangeably with autonomous driving.
[0179] The computing device 2000 may further include (but is not limited to) large-scale computing systems, such as server computers, desktop computers, etc., and may further include set-top boxes (e.g., Internet-based cable TV set-top boxes, etc.), GPS-based devices, etc. The computing device 2000 may include mobile computing devices that act as communication devices, such as cellular phones including smartphones, personal digital assistants (PDAs), tablet computers, laptop computers, e-readers, smart TVs, television platforms, wearable devices (e.g., glasses, watches, bracelets, smart cards, jewelry, clothing, etc.), media players, etc. For example, in one embodiment, the computing device 600 may include a mobile computing device employing a computer platform hosted on a single chip on which various hardware and / or software components of the computing device 2000 are integrated (“IC”), such as a system-on-a-chip (“SoC” or “SOC”).
[0180] As illustrated, in one embodiment, computing device 2000 may include any number and type of hardware and / or software components, such as (but not limited to) a graphics processing unit (“GPU” or “graphics processor”) 2014, a graphics driver (also referred to as a “GPU driver,” “graphics driver logic,” “driver logic,” user-mode driver (UMD), UMD, user-mode driver framework (UMDF), UMDF, or simply “driver”) 2016, a central processing unit (“CPU” or “application processor”) 2012, memory 2008, network devices, drivers, etc., and one or more input / output (I / O) sources 2004, such as a touchscreen, touchpad, touchpad, virtual or regular keyboard, virtual or regular mouse, ports, connectors, etc. Computing device 2000 may include an operating system (OS) 2006 that serves as an interface between the hardware and / or physical resources of computing device 2000 and the user. It is contemplated that the graphics processor 2014 and application processor 2012 may be... Figure 1 One or more of the processors 102.
[0181] It should be recognized that systems with fewer or more features than the examples described above may be preferred for certain implementations. Therefore, the configuration of computing device 2000 can vary from implementation to implementation depending on numerous factors, such as price constraints, performance requirements, technological improvements, or other circumstances.
[0182] Implementations may be carried out as any one or a combination of the following: one or more microchips or integrated circuits interconnected using a motherboard, hard-wired logic, software stored in a memory device and executed by a microprocessor, firmware, application-specific integrated circuits (ASICs) and / or field-programmable gate arrays (FPGAs). Throughout this document, the terms “logic,” “module,” “component,” “engine,” “mechanism,” “tool,” “line,” and “circuit” are used interchangeably and, by way of example, include software, hardware, firmware, and / or any combination thereof.
[0183] In one embodiment, as illustrated, the prediction and correlation mechanism 2010 may be managed by the memory 2008 of the computing device 2000. In another embodiment, the prediction and correlation mechanism 2010 may be managed by the operating system 2006 or the graphics driver 2016. In yet another embodiment, the prediction and correlation mechanism 2010 may be part of or managed by the firmware of the graphics processor 2014 or the graphics processing unit (“GPU” or simply “graphics processor”) 2014. For example, the prediction and correlation mechanism 2010 may be embedded in or implemented as part of the processing hardware of the graphics processor 2014. Similarly, in yet another embodiment, the prediction and correlation mechanism 2010 may be part of or managed by the central processing unit (“CPU” or simply “application processor”) 2012. For example, the prediction and correlation mechanism 2010 may be embedded in or implemented as part of the processing hardware of the application processor 2012.
[0184] In yet another embodiment, the prediction and correlation mechanism 2010 may be part of or hosted by any number and type of components of the computing device 2000. For example, one part of the prediction and correlation mechanism 2010 may be part of or hosted by the operating system 2006, another part may be part of or hosted by the graphics processor 2014, another part may be part of or hosted by the application processor 2012, and one or more parts of the prediction and correlation mechanism 2010 may be part of or hosted by any number and type of devices and / or operating system 2006 of the computing device 2000. It is contemplated that the embodiments are not limited to any implementation or hosting of the prediction and correlation mechanism 2010, and one or more parts or components of the prediction and correlation mechanism 2010 may be adopted or implemented as hardware, software, or any combination thereof, such as firmware.
[0185] The computing device 2000 may host one or more network interfaces to provide access to networks such as LANs, WANs, MANs, PANs, Bluetooth, cloud networks, mobile networks (e.g., 3G, 4G, etc.), intranets, the Internet, etc. The network interfaces may include, for example, a wireless network interface with an antenna, which may represent one or more antennas. The network interfaces may also include, for example, a wired network interface for communicating with remote devices via a network cable, such as an Ethernet cable, coaxial cable, fiber optic cable, serial cable, or parallel cable.
[0186] Examples of embodiments may be provided as computer program products, which may include one or more machine-readable media on which machine-executable instructions are stored, which, when executed by one or more machines (such as computers, computer networks, or other electronic devices), cause the one or more machines to perform operations according to the embodiments described herein. Machine-readable media may include, but are not limited to, floppy disks, optical disks, CD-ROMs (compact disc read-only memory) and magneto-optical disks, ROMs, RAMs, EPROMs (erasable programmable read-only memory), EEPROMs (electrically erasable programmable read-only memory), magnetic cards or optical cards, flash memory, or other types of media / machine-readable media suitable for storing machine-executable instructions.
[0187] Furthermore, the embodiments can be downloaded as a computer program product, wherein the program can be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a computer) via a communication link (e.g., a modem and / or a network connection) as one or more data signals implemented on and / or modulated by a carrier or other propagation medium.
[0188] Throughout this document, the term "user" may be used interchangeably with "viewer," "observer," "person," "individual," "end user," and / or similar terms. It should be noted that throughout this document, terms like "graphics domain" may be used interchangeably with "graphics processing unit," "graphics processor," or simply "GPU," and similarly, "CPU domain" or "host domain" may be used interchangeably with "computer processing unit," "application processor," or simply "CPU."
[0189] It should be noted that terms such as "node," "computing node," "server," "server device," "cloud computer," "cloud server," "cloud server computer," "machine," "host machine," "device," "computing device," "computer," and "computing system" can be used interchangeably throughout this document. Furthermore, terms such as "application," "software application," "program," "software program," "package," and "software package" can be used interchangeably throughout this document. Also, terms such as "work," "input," "request," and "message" can be used interchangeably throughout this document.
[0190] Figure 21 The illustration shows an embodiment. Figure 20 The prediction and correlation mechanisms (2010). For the sake of brevity, references have been made. Figures 1-20 Many details of the discussion may be omitted or repeated below. In one embodiment, the prediction and correlation mechanism 2010 may include any number and type of components, such as (but not limited to): detection and selection logic 2101; calculation and analysis logic 2103; construction and assignment logic 2105; communication / compatibility logic 2107; and encoding logic 2109.
[0191] As illustrated, in one embodiment, a server device 2000 (also referred to as a "server computer" or "construction device") on the capture or construction side communicates with another computing device 2150 (also referred to as a "client computer") on the rendering side. The client device 2150 includes a memory 2158 that hosts a prediction / correlation-based rendering mechanism ("rendering mechanism") 2160, which has any number and type of components, such as (but not limited to): detection and reception logic 2161; decoding logic 2163; interpretation and prediction logic 2165; application and rendering logic 2167; and communication / display logic 2169. The client device 2150 further includes a graphics processor 2154 and an application processor 2152 communicating with the memory 2158 to execute the rendering mechanism 2160. The client device 2150 further includes a user interface 2159 and one or more I / O components 2156, which include one or more display devices or screens for displaying immersive media such as 3DoF+ video, 6DoF video, etc.
[0192] It is conceivable that although server device 2000 and client device 2150 are shown as two separate devices communicating with each other via one or more communication media 2025, the embodiments are not limited thereto. For example, in some embodiments, server device 2000 and client device 2150 may not be two separate devices, but a single device having all the functionality and components of the two computing devices 2000, 2150. Similarly, the embodiments are not limited to one or two computing devices, and there may be three or more devices performing one or more of the aforementioned functionalities, such as cameras (one or more) A2181, B 2183, C 2185, D 2187, which may be regarded as separate devices or hosted on one or more computing devices communicating with server device 2000 and / or client device 2150.
[0193] The server device 2000 is further shown to communicate with one or more repositories, datasets, and / or databases (such as one or more databases 2130, e.g., cloud storage, non-cloud storage, etc.), wherein the one or more databases 2130 may reside on a local storage device or a remote storage device via one or more communication media 2025 (such as one or more networks, e.g., cloud networks, proximity networks, mobile networks, intranets, the Internet, etc.).
[0194] It is anticipated that software applications running on server device 2000 can be responsible for using one or more components of computing device 2000 (e.g., GPU 2014, graphics driver 2016, CPU 2012, etc.) to perform or facilitate the execution of any number and type of tasks. When performing such tasks, as defined by the software application, one or more components such as GPU 2014, graphics driver 2016, CPU 2012, etc., can communicate with each other to ensure that those tasks are processed and completed accurately and in a timely manner.
[0195] Adaptive resolution The embodiments provide a novel technique for adaptive resolution of a point cloud based on 1) object metadata such as the edges of large volumes or highly detailed regions (e.g., human fingers, faces, etc.); and 2) the location of objects in a scene such as camera distance, field of view, and focus, while adjusting the fidelity of the point cloud to prioritize execution time for objects or object regions that have or are expected to have the greatest impact on the user's quality perception.
[0196] In one embodiment, the prediction and correlation mechanism 2010 can be used to conservatively predict viewpoint updates for each frame, such that the adaptive multi-frequency shader determines which texture data to send, as facilitated by the computation and prediction logic 2103 and the construction and assignment logic 2105. Alternatively, concave rendering techniques can be applied, for example. In one embodiment, viewport / user prediction is performed on the server side, as facilitated by the computation and prediction logic 2103, such that only relevant data is separated and assigned for rendering, as facilitated by the construction and assignment logic 2105, and sent to the client device 2150 via communication medium 2025, as facilitated by communication / compatibility logic 2107, wherein the transmitted data is then prepared by the client device 2150 and rendered on the renderer side, as facilitated by one or more of the interpretation and prediction logic 2165, application and rendering logic 2167, and communication / display logic 2169.
[0197] It is foreseeable that latency may be a concern, for example, in situations where each frame of data depends on the viewport and thus has occasional updates on the viewport from client device 2150 to server device 2000, where first and second derivatives can be used so that the prediction and correlation mechanism 2010 of server device 2000 can conservatively predict at each frame and for each frame's viewport. Additionally, for example, detection and selection logic 2101 and calculation and prediction logic 2103 can be used to perform Adaptive Multi-Frequency Shading (AMFS) rendering by determining which texture data to send to client device 2150 for rendering, where, for example, AMFS is facilitated by calculation and prediction logic 2103 to conservatively calculate coarse slices while only sending those qualified slices as specified by construction and assignment logic 2105. Furthermore, this can be based on pin pad cylinder compression perimeters, which can be calculated to reduce even more data.
[0198] Further referring to the adaptive resolution of the point cloud, in one embodiment, the detection and selection logic 2101 can be used to detect a scene captured by one or more of cameras A 2181, B 2183, C 2185, and N 2187, wherein the scene contains objects, wherein the detection and selection logic 2101 further detects and selects any metadata related to or associated with the objects, such as metadata related to the edges of large volumes or highly detailed regions (such as human fingers, faces, etc.). Additionally, in one embodiment, the detection and selection logic 2101 further detects the location of objects in the scene, such as in terms of distance from one or more cameras A 2181, B 2183, C 2185, and D 2187.
[0199] In one embodiment, database(s) 2130 may be used to store and maintain raw footage of scenes captured by one or more of cameras 2181, 2183, 2185, 2187, wherein having these raw footage at database(s) 2130 allows prediction and correlation mechanisms 2010 and rendering mechanisms 2160 to perform their tasks or operations (e.g., 3D model generation) offline and / or on demand, as described throughout this document. For example, even when a network and / or system such as server device 2000 and / or client device 2150 is down or offline, any raw footage from database(s) 2130 may be accessed and obtained via detection and selection logic 2101 and / or detection and reception logic 2161. Similarly, such raw footage may be accessed and obtained on demand or as necessary, without having to wait for the live streaming or delivery of camera footage.
[0200] Using computation and prediction logic 2103, the fidelity of the point cloud is adapted to prioritize the execution of objects and / or object regions that have the greatest impact, highest importance, or highest visibility on the user's quality perception. Following such prediction, construction and assignment logic 2105 can be used to classify information based on low-fidelity and high-fidelity of objects and / or object regions, such that the information can then be encoded by encoding logic 2109. The encoded information is then transmitted to client device 250 for rendering purposes, wherein the transmission is accomplished via network 2125, facilitated by communication / compatibility logic 2107 and communication / display logic 2169.
[0201] At client device 2150, on the rendering side, encoded information is received by detection and reception logic 2161 and then decoded by decoding logic 2163 for further processing by interpretation and prediction logic 2165. In one embodiment, the decoded information is interpreted for low-fidelity and high-fidelity classification, and because the encoded information received from server device 2000 contains only relevant information, all information is reconstructed in the rendered version for rendering purposes, as facilitated by interpretation and prediction logic 2165. The rendered version of the information is then rendered by application and rendering logic 2167 and displayed using one or more display devices, as facilitated by communication / display logic 2169, where the rendered information includes immersive media such as 3DoF+ video, 6DoF video, etc.
[0202] For example, refer to Figure 22It illustrates a transaction sequence 2200 for adaptive resolution of a point cloud according to one embodiment. Any process of the transaction sequence 2200 can be executed by processing logic, which may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof, such as... Figure 20 Prediction and correlation mechanisms 2010 and Figure 21 The rendering mechanism 2160 facilitates this. For the sake of simplicity and clarity in presentation, processes associated with a sequence of transactions can be illustrated or explained in a linear order; however, it is to be expected that any number of them can be executed in parallel, asynchronously, or in a different order.
[0203] As here Figure 22 As illustrated and as described above, any data captured by one or more cameras 2181, 2183, 2187, 2189 can be referred to as camera feed, corresponding to various cameras 2181, 2183, 2185, 2187. These feeds contain a scene, which further contains objects (e.g., living and / or inanimate objects, such as people, faces, hands, animals, plants, furniture, stars, mountains, cars, etc.). Figure 22 As illustrated, transaction sequence 2200 begins at 2201 by receiving camera scene data, and at 2203 by determining whether the camera scene data contains any objects detected by detection and selection logic 2101. If no objects or at least no relevant objects exist in the scene, the scene is not considered highly relevant, and thus transaction sequence 2200 continues with the scene, and at 2209 its content is flagged and considered low fidelity.
[0204] However, if an object is found in the scene, in one embodiment, at 2207, another determination is made regarding the relevance of the object or any area of the object, such as high relevance or low relevance, based on the probability or proximity of the user viewing the object and / or one or more areas as determined and analyzed by the calculation and prediction logic 2103. In one embodiment, this determination can be made at 2205 based on object metadata, wherein the object metadata may contain any information or type that can help determine the relevance of the object and / or areas to the user, such as historical viewing patterns, known importance, distance from one or more cameras 2181, 2183, 2185, 2187, in-focus or out-of-focus objects / areas in the view, whether the object / area in the view is peripheral or centered, etc., as facilitated by the calculation and prediction logic 2103.
[0205] In one embodiment, based on metadata analysis, if one or more objects and / or regions are considered to be of low relevance or importance, they are marked as low fidelity at 2209, as facilitated by the construction and assignment logic 2105. Conversely, if objects and / or regions are considered to be of high relevance or importance, they are then marked as high fidelity at 2211, as facilitated by the construction and assignment logic 2105. Once the fidelity level is determined, the encoder then encodes the predicted information regarding the fidelity level (such as low fidelity or high fidelity) associated with the objects and / or regions, as facilitated by the encoding logic 2109, and passes it to the client device 2150, as facilitated by the communication / compatibility logic 2107.
[0206] In one embodiment, the fidelity prediction is then received at the client device 2150 by detection and reception logic 2161, decoded by a decoder as facilitated by decoding logic 2163, and then interpreted by interpretation and prediction logic 2165. The interpreted data can then be prepared for rendering, as facilitated by application and rendering logic 2167, and subsequently rendered and displayed on a display device or screen, such as the display device or screen of a wearable device (such as an HMD) or a mobile device display screen, a desktop monitor, etc., as facilitated by communication / display logic 2169.
[0207] Although in this embodiment, fidelity prediction is performed on the server side, such as through calculation and prediction logic 2103 based on the extraction and evaluation of object metadata, it is expected and should be noted that the embodiment is not limited thereto. In some embodiments, fidelity prediction may be performed on the client side, such as through interpretation and prediction logic 2165 based on user observation and evaluation and / or display movement, such as the movement of the HMD relative to the user and / or objects in the environment.
[0208] Additionally, as described with reference to this and other embodiments, the various components of the prediction and correlation mechanism 2010 and the rendering mechanism 2160 are not fixed in their process or location. For example, one or more of these components may be used in different processes or located elsewhere, such as on any one or both of the server device 2000 and the client device 2150.
[0209] Before proceeding further, it is anticipated and noted that the embodiments discussed above and below are not limited to only those objects or views in a fully visible scene. For example, in cases where certain objects may occlude parts of the model (such as virtual objects, virtual trees occluding the viewpoint), the prediction and correlation mechanism 2010 and / or rendering mechanism 2160 can be used to predict the occluded parts using any number of techniques (e.g., raw data or metadata from database(s)2130, artificial intelligence via machine / deep learning models, etc.).
[0210] Similarly, it is expected and noteworthy that, although the use of head poses based on head position and movement is discussed throughout this document for the sake of brevity and clarity, the embodiments are not limited thereto. For example, in a 6DoF view, predictions based on the physical displacement of a person (such as a user) in the real world can also be used, not just measurements of the user's head pitch, yaw, and rotation. In other words, for example, any relevant portion of the immersive media being predicted and used for rendering and display is not simply based on head pose, but can also be based on other relevant factors, such as the physical displacement of a person / user in the real world. For example, based on the user's known or determined physical displacement relative to the real world, future physical displacements or information can be calculated and predicted, and subsequently used for the selection of relevant portions of the immersive media to be rendered and displayed, as facilitated by the prediction and correlation mechanism 2010 and / or the rendering mechanism 2160.
[0211] Viewport and / or user prediction The embodiment further provides a novel technique for constructing or predicting viewports and / or users at server device 2000 to transmit relevant data to the user only on the rendering side using client device 2150. In one embodiment, predicting head motion and / or other physical motion is used to estimate subsequent frames in a streaming video frame. In one embodiment, the novel technique can be performed by using motion vectors in 6DoF, sending the motion vectors to server computer 2000 to estimate the next frame, and calculating the next frame moment for only a portion of the entire 6DoF content.
[0212] As discussed, conventional technologies cannot take into account the user's perspective and the relevance of the content when streaming content, and thus all content (such as the entire cloud information containing objects and / or scenes) is streamed to the user on the rendering side, which leads to inefficiency and waste of resources, such as system and network bandwidth, time, etc.
[0213] Referring to one embodiment of viewport and / or user prediction on the reference server device 2000, sending only relevant data to the client device 2150 for rendering purposes may include using the server device 2000 on the server side to process all immersive media (e.g., 3DoF+ video, 6DoF video, etc.), while using the client computer 2150 on the renderer side to send and render only the relevant immersive media. For example, HMDs, displays, etc., may be used to render the relevant media to provide an enhanced viewing experience from the user. For example, immersive media such as 6DoF video may be game videos, which are rendered or displayed on an HMD for the user to experience by watching and playing.
[0214] As mentioned earlier, point clouds used for immersive media content, as well as current 180 / 360-degree video, require high bandwidth to transmit the entire scene from the server to the client / consumer over the network. This is because, in any given instance, the user is only viewing a portion of the scene and the rest remains unviewed; sending and rendering all the data, regardless of their relevance, can lead to network / bandwidth overhead.
[0215] The embodiments provide for transmitting point clouds by sending only preferred, relevant, or desired or slightly more desired scenes to the user from the client device 2150, thereby saving significant network and bandwidth overhead. This novel technique provides bandwidth reduction, which in one embodiment typically requires delivering immersive media content to the consumer by estimating or predicting the viewport for subsequent frames that the user expects or may see and transmitting only the desired portion of the scene. As mentioned above, such prediction of the viewport may be performed on the server side or the client side, such as using computation and prediction logic 2103 or interpretation and prediction logic 2165, respectively.
[0216] For example, if prediction is performed on the client side, as facilitated by interpretation and prediction logic 2165, any metadata for predicting the viewport is sent to server device 2000 to calculate relevant data to be transmitted for that viewport, as facilitated by calculation and prediction logic 2103. Additionally, any additional data outside the viewport can be transmitted at a lower resolution or can be skipped entirely to save bandwidth. For example, to reduce latency in transmitting the predicted viewport from client device 2150 to server device 2000, the predicted viewport can be calculated by calculation and prediction logic 2103 on the server side when client device 2150 sends relevant header position information for previous frames as calculated and interpreted by interpretation and prediction logic 2165. On the server side, in one embodiment, header position prediction can be performed by calculation and prediction logic 2103 based on first and second derivatives that are lazily (e.g., periodically but not necessarily per frame) updated from client device 2150 to server device 2000. For example, the previous head position can be used to calculate the viewport, and in subsequent frames, if the calculated head position differs from the actual head position, the predicted viewport / head position can be corrected.
[0217] Client-side viewport and / or user prediction For example, as referenced Figure 23A As illustrated, viewport prediction can be performed on the client side, such as by one or more server devices 2301 (e.g., Figure 20 The server device 2000 can communicate with the wearable device / display 2303 (e.g., an HMD, a display at the client device 2150) via a communication medium 2125 (e.g., a cloud network, the Internet, etc.). As previously described, in conventional technology, a complete 360-degree video for the initial frame is transmitted between the server and the client computer. For brevity, the previous reference... Figures 1-22 Many details of the discussion may be omitted or repeated in the following text.
[0218] In the illustrated client-based viewport prediction transaction sequence 2300, an initial frame of the entire point cloud or the entire 360-degree scene is encoded and sent from the server side to the client side at 2311. The initial frame is decoded at 2315, as facilitated by decoding logic 2163, and the decoded information is then used to detect and record the current user head position for each frame of the initial frame obtained from the decoded information, as facilitated by detection and reception logic 2161. This head position, along with previous frame head positions, is then used to predict future positions, and based on the head position, the viewport is predicted at 2317, as facilitated by interpretation and prediction logic 2165. In one embodiment, at 2319, this viewport prediction is then transmitted to the server side as viewport prediction feedback, as facilitated by communication / display logic 2169.
[0219] At this point, one or more server devices 2301 process this feedback received from the client side, as facilitated by the calculation and prediction logic 2103, and then, at 2313, only the relevant data for a specific prediction viewport is sent from the server side to the client side for further processing, such as rendering specific immersive media (e.g., 3DoF+ video, 6DoF video, etc.) using one or more display devices (such as wearable devices / displays 2303), as facilitated by the application and rendering logic 2167. It is anticipated that users and / or broadcasters may choose to send data slightly larger than the prediction viewport, and similarly, to save additional bandwidth, any data outside the viewport may be sent at a low resolution or not sent at all.
[0220] Server side view and / or user prediction For example, now refer to Figure 23B It illustrates server-side viewport prediction, which is then sent to the client side for rendering. For brevity, refer to the previous section. Figures 1-23A Many details of the discussion may be omitted or repeated in the following text. For example, Figure 23B In the transaction sequence 320, in a server-based viewport prediction system, the initial frames may encompass the entire 360-degree scene or the entire point cloud, and in 2321, these initial frames may be passed from the server side to the client side, as facilitated by communication / compatibility logic 2107. In one embodiment, the current user head position may be recorded for the most recently rendered frames along with their frame identifiers on the client side, as facilitated by interpretation and prediction logic 2165, and sent to the server side in 2327, as facilitated by communication / display logic 2169.
[0221] On the server side, detection and selection logic 2101 receives the current header position and frame identifier, and subsequently, calculation and prediction logic 2103 can be triggered at 2333 to predict the viewport using their frame identifiers based on the deviation of each header position from the previous frame to the current frame. For example, the process can be based on viewport prediction calculation at 2329 based on the corrected header position, as facilitated by calculation and prediction logic 2103, and at 2331 incremental correction between the predicted viewport and the desired viewport based on the actual header position of a specific frame identifier, such as correcting the incremental difference between the predicted viewport and the viewport based on the actual header position of a specific frame, as facilitated by calculation and prediction logic 2103.
[0222] Once the prediction data is collected and corrected as needed or desired, construction and assignment logic 2105 is triggered to classify and assign the prediction data as a prediction viewport. Then, at 2323, this prediction viewport, based on subsequent frames received from the client predictor and feedback from the client side, is transmitted from the server side to the client side, as facilitated by communication / compatibility logic 2107, and via communication media(s) 2125. In one embodiment, the prediction viewport is encoded by one or more encoders on the server side, as facilitated by encoding logic 2109, and when received on the client side, the encoded prediction viewport is then decoded at 2325 by one or more decoders, as facilitated by decoding logic 2163.
[0223] Based on this short-term head position data, the calculation and prediction logic 2103 can continue to predict the viewport at 2333, and the communication / compatibility logic 2107 can continue to send the predicted viewport containing only relevant and / or necessary data to the client side at 2323 for rendering purposes, and in doing so, the server saves the viewport / predicted head position along with the frame identifier.
[0224] In one embodiment, once head position feedback from the client side is received on the server side for a specific frame identifier at 2327, the calculation and prediction logic 2103 compares and calculates the increment between its predicted value and the actual expected value. If there is an increment or discrepancy between the future predictions for the head position, the corresponding viewport is updated by compensating for the increment value. Once the increment is the smallest for a specific viewport, data can be sent to the client side at 2323, saving bandwidth by sending only relevant and / or necessary data instead of the entire content. In other words, server-based prediction ensures that the server device (such as server device 2000) does not need to rely on feedback from the client device (such as client device 2150) for each individual frame and can correct its predictions based on the feedback received from client device 2150.
[0225] In one embodiment, regardless of whether the predicted viewport is server-side or client-side, the construction and specification logic 2105 at server device 2000 can be used to construct or reconstruct textures by selecting the correct patches for the corresponding object using one or more of texel space shading (TSS), AMFS, etc. This novel technique identifies the exact texture blocks required by application and rendering logic 2167 at client device 2150 to reconstruct a complete view of the object. Additionally, using calculation and prediction logic 2103, the viewport can be calculated conservatively (e.g., with guard bands), and / or any texel block can be calculated conservatively on, for example, 128B blocks or even larger.
[0226] This sparse texture can then be prepared by construction and specification logic 2105, encoded by encoding logic 2109, and then sent to client device 2150 by communication / compatibility logic 217, instead of sending the entire point cloud. Additionally, the sparse texture can be compressed using one or more compression techniques to fill the correct texture blocks at client device 2150. Furthermore, AMFS can be used to select the correct level of detail (LOD) for a given location (u, v) in object space to further help reduce bandwidth, as per the requirements for... Figure 23C The transaction sequence 2350 for unpacking and viewport selection is illustrated in the diagram. (See reference...) Figure 23C As illustrated, patches in the point cloud are packed into a 2D rectangle at 2351, and then at 2353, the correct patch is selected in the packed point cloud to reconstruct the rendered texture for the selected viewport. Finally, at 2355, the texture composed of the patches is rendered within the packed rectangle.
[0227] Figure 23A , 23B Any process of 23C can be executed by processing logic, which may include hardware (e.g., circuits, special-purpose logic, programmable logic, etc.), software (such as instructions that run on a processing device), or a combination thereof, such as by... Figure 20 Prediction and correlation mechanisms 2010 and Figure 21 The rendering mechanism 2160 facilitates this. For the sake of simplicity and clarity in presentation, processes associated with a sequence of transactions can be illustrated or explained in a linear order; however, it is to be expected that any number of them can be executed in parallel, asynchronously, or in a different order.
[0228] Viewport-related immersive video streaming The embodiments further provide novel techniques for viewport-dependent 6DoF video streaming. For example, as described above, user head pose information and / or physical displacement information can be obtained through detection and reception logic 2161 at client device 2150 and transmitted to server device 2000, where the head pose information and / or physical displacement information is analyzed by calculation and prediction logic 2103, and the results are then used to obtain relevant information, such as information relevant only from the user's perspective, as facilitated by construction and assignment logic 2105. The analysis performed by calculation and prediction logic 2103 may include checking video quality and throughput, frame capture to determine which part of the content is being streamed, etc.
[0229] The relevant viewport-related video is then encoded via encoding logic 2109 and, as facilitated by communication / compatibility logic 2107, transmitted to client device 2150 via communication medium 2125. This novel technique for sending only relevant information based on the user's perspective allows for efficient streaming of immersive media, reducing bandwidth usage while achieving the highest possible fidelity content for rendering purposes at client device 2150, based on throughput limitations.
[0230] Video-based approach to immersive media user viewport For example, now refer to Figure 24A The illustration depicts a video-based method for an immersive media (e.g., 3DoF+ video, 6DoF video) user viewport according to one embodiment. For brevity, the previous reference... Figures 1-23C Many details of the discussion need not be repeated or discussed further below. For example... Figure 24A As illustrated, one or more cameras, such as camera A2181 and camera N2189, can be used to capture a scene of an object, wherein these camera feeds of the scene can be used to reconstruct the geometry of the captured scene and the corresponding position or localization and field of view (FOV) 2401, 2403 (at boxes 2401 and 2403, respectively). This reconstructed geometry can be achieved using one or more occlusion detection techniques at boxes 2405 and 2407, using camera feeds obtained from the corresponding cameras A2181 and N2189, as facilitated by computation and prediction logic 2103.
[0231] Additionally, at frames 2409 and 2411, valid pixel information associated with the corresponding cameras 2181 and 2189 (relative to the areas where the camera's view is occluded) is used to create a mapping of 3D regions for each camera 2181 and 2189. Similarly, occlusion detection technology is used at frame 2415, and a region of user-visible objects or objects is created at frame 2417. Furthermore, in one embodiment, at frame 2431, any information including the user's FOV and the location of or related to the cameras 2181 and 2189 is transmitted from the client side (such as from client device 2150) to the server side (such as server device 2000).
[0232] In one embodiment, a base camera and a patch are selected at frame 2419, wherein the base camera among cameras 2181, 2189 at frame 2421 is selected to match the position and / or orientation closest to the user, or in another embodiment, the base camera is selected to have the highest match with the user's visible area from frame 2417 in terms of overlap area volume. Additionally, in one embodiment, patches from other cameras among cameras 2181, 2189 are selected at frame 2423, wherein patches from these other cameras are selected to fill the remaining user visible area of frame 2417, as facilitated by the construction and assignment logic 2105. This addition of patches can continue until a maximum number of camera views reaches a threshold for the coverage of the user area.
[0233] Furthermore, the use of a practical matching method with a video encoder can encode multiple independent scan lines or independent slices, enabling the mixing and matching of patches to form a new stream for encoding by the encoder at frames 2425, 2247, as facilitated by encoding logic 2109. This technique can be used to facilitate 1:N encoding for users by having an encoded stream that can be formed from different lines or slices rather than streams completely different for each user.
[0234] After the encoded data is transmitted from the server side (such as server device 2000) to the client side (such as client device 2150), it is then received and decoded using a decoder at block 2429, as facilitated by decoding logic 2163. In one embodiment, a user's view can then be created at block 2433 using existing view difference and decoding information from block 2431 for the user's current position and FOV, as facilitated by interpretation and prediction logic 2165. This view can then be rendered to the user, as facilitated by application and rendering logic 2167, and displayed using a display screen.
[0235] Now for reference Figure 24B The illustration shows an encoding system 2440 according to one embodiment of a video-based method for immersive media (e.g., 3DoF+ video, 6DoF video) user viewports. For brevity, the previous reference is... Figures 1-24AMany details of the discussion may be omitted or repeated below. As illustrated, the encoding system 2440 is configured to receive a basic view 2441 from a basic camera, thereby producing HEVC encoding 2443. In one embodiment, an additional view 2441 of video (e.g., 3DoF+ video) is also received, but undergoes a patching process, wherein patch information 2453A, 2453B, 2453B corresponding to the video providing the additional view 2441 is used to patch the video and combine them into a patch package 2455. In the patch package 2455, not only is HEVC encoding 2459 (similar to or equivalent to HEVC encoding 2443) then provided, but also corresponding metadata encoding 2457 (driven from patch information 2453A, 2453, 2453B) is provided.
[0236] Now for reference Figure 24C The illustration shows a video-based decoding system 2450 according to one embodiment of a method for immersive media (e.g., 3DoF+ video, 6DoF video) user viewports. For brevity, the previous reference is... Figures 1-24B Many details of the discussion may be omitted or repeated below. As illustrated, the decoding system 2450 provides decoding of incoming decoding information, such as... Figure 24B HEVC encoding 2443, HEVC encoding 2459, and metadata encoding 2457 are configured to be decoded into HEVC decoding 2451, HEVC decoding 2453, and metadata decoding 2455, respectively. In one embodiment, any information associated with or obtained through HEVC decoding 2453 and metadata decoding 2455 is then configured for patch unpacking 2457, thereby generating a result corresponding to and based on... Figure 24B The video view generation 2459A, 2459B, and 2459C. These view generation 2459A, 2459B, and 2459C undergo view composition 2461, where the output viewport of view composition 2461 is then rendered and displayed on a display device such as HMD 2463.
[0237] Figure 24A , 24B Any process of 24C can be executed by processing logic, which may include hardware (e.g., circuits, special-purpose logic, programmable logic, etc.), software (such as instructions that run on a processing device), or a combination thereof, such as by... Figure 20 Prediction and correlation mechanisms 2010 and Figure 21 The rendering mechanism 2160 facilitates this. For the sake of simplicity and clarity in presentation, processes associated with a sequence of transactions can be illustrated or explained in a linear order; however, it is to be expected that any number of them can be executed in parallel, asynchronously, or in a different order.
[0238] Point cloud-based approach for immersive media user viewport Now for reference Figure 25 The illustration shows a transaction sequence 2500 based on a point cloud-based method for immersive media (e.g., 3DoF+ video, 6DoF video) user viewports according to one embodiment. For brevity, the previous reference is... Figures 1-24C Many details discussed herein need not be repeated or discussed further below. Any process related to transaction sequence 2500 can be executed by processing logic, which may include hardware (e.g., circuits, special-purpose logic, programmable logic, etc.), software (such as instructions running on a processing device), or a combination thereof, such as by... Figure 20 Prediction and correlation mechanisms 2010 and Figure 21 The rendering mechanism 2160 facilitates this. For the sake of simplicity and clarity in presentation, processes associated with a sequence of transactions can be illustrated or explained in a linear order; however, it is to be expected that any number of them can be executed in parallel, asynchronously, or in a different order.
[0239] In the illustrated embodiment, at box 2521, one or more client devices, such as client device 2150, are used on the client side to generate or obtain the user's position and FOV, wherein the position and FOV (FOV limitation, occlusion detection, etc.) are based on what the user can see on or through a display device (such as a display screen, HMD, etc.). The embodiment also provides novel techniques to account for areas or regions that may be occluded for one reason or another, such as virtual objects occluding parts of a constructed model, which can then be covered or detected by occlusion detection techniques, such as those provided by [other methods]. Figure 21 This is facilitated by the detection and reception logic 2161 and / or the detection and selection logic 2101. In one embodiment, some of the features used to collect the user's location and FOV characteristics include one or more of the following: the distance of the object from the camera (e.g., closer objects require full detail), objects at the center of the FOV, predefined objects of interest (e.g., people, significant objects relative to simple background objects, etc.), more detail around the edges of the object, or areas where higher resolution is desired (e.g., a person's face) and / or the like.
[0240] The location and FOV associated with the user are converted into a metadata bitstream and transmitted over a network (such as a cloud network, the Internet, etc.) to a server device (such as server device 2000). On the server side, a full-resolution full scene based on the point cloud is received at box 2501, along with the metadata bitstream received from the client side. In one embodiment, the bitstream, along with the scene content, is then used to provide downsampling at box 2503 and key feature selection at box 2507, wherein the downsampling at box 2503 and the key feature selection at box 2507 result in a low-resolution full-scene point cloud at box 2505 and a high-resolution key feature point cloud at box 2509, respectively.
[0241] In box 2511, point cloud information, including the low-resolution full-screen point cloud 2505 at box 2509 and high-resolution key features, is then compressed in box 2511. This compressed point cloud information is then passed back to the client side, where point cloud decompression is performed at box 2423. In one embodiment, the dataset from the point cloud decompression at box 2523, along with the user position and FOV at box 2521, is used to obtain the scene rendered at box 2527 and displayed on the display device at box 2529.
[0242] It is anticipated that the embodiments are not limited to simple low and high resolutions, but can use any number of varying levels of point cloud resolution based on feature importance. Similarly, varying levels of point cloud compression (such as lossy compression) can be used for different levels of feature interest.
[0243] Viewport-related immersive video streaming Now for reference Figure 26 The illustration is based on a system 2600 for streaming viewport-dependent immersive media (e.g., 3DoF+ video, 6DoF video) according to one embodiment. For simplicity, the previous reference is... Figures 1-25 Many details of the discussion may be omitted or repeated below. Any process related to System 2600 can be executed by processing logic, which may include hardware (e.g., circuits, special-purpose logic, programmable logic, etc.), software (such as instructions that run on a processing device), or a combination thereof, such as by... Figure 20 The construction of pipeline mechanisms 2010 and Figure 21 The rendering pipeline mechanism 2160 facilitates this. For the sake of simplicity and clarity in presentation, processes associated with a sequence of transactions can be illustrated or explained in a linear order; however, it is to be expected that any number of them can be executed in parallel, asynchronously, or in a different order.
[0244] As illustrated, this embodiment of the system 2600 for viewport-dependent video streaming includes a server or construction device 2000 and a client or rendering device 2150, wherein the server device 2000 receives head pose measurements and predictions (and / or physical displacement measurements and predictions about a real-world user) 2617 from the client device 2150. In one embodiment, the head pose measurements and predictions (and / or physical displacement measurements and predictions) can be calculated at the client device 2150 (e.g., HMD) using tracking of the HMD via sensors, cameras, etc., of I / O components 2156, and facilitated by detection and reception logic 2161. Based on this collected information, the future head position (and / or physical displacement) is calculated, as facilitated by interpretation and prediction logic 2165. Upon receiving head pose measurement and prediction (and / or physical displacement measurement and prediction) 2617, as facilitated by computation and analysis logic 2103, texture and / or depth maps 2605 are calculated for video captured by one or more of cameras A 2181, B 2183, N 2187, etc. Additionally, the texture and / or depth map 2605 may be based on a 3D model of the video generated in real time or pre-generated and stored in database 2130.
[0245] In one embodiment, based on head pose measurement and prediction (and / or physical displacement measurement and prediction) 2167, the user-visible area (and, for example, edges) is identified by calculation and analysis logic 2103, and then encoded by encoding logic 2109 and packaged by construction and assignment logic 2105. Alternatively, in one embodiment, only packets containing the user-visible area 2607 are sent to the client device 2150, which reduces resource consumption (such as bandwidth, computing resources, power, time, etc.). As previously stated, it is anticipated and should be noted that the embodiments are not limited to head pose measurement and prediction, but can also track, analyze, consider, and utilize physical displacement of the user in the real world to obtain and render relevant portions of immersive media.
[0246] The client device 2150 receives and unpacks 2611 packets, as facilitated by the detection and reception logic 2161 and the interpretation and prediction logic 2165, respectively, and decodes them 2613 as facilitated by the decoding logic 2163, and then renders 2615 the decoded information to a display device for viewing by a user as facilitated by the application and rendering logic 2167.
[0247] Refer back Figure 21The communication / compatibility logic 2107 can be used to facilitate the necessary communication and compatibility between any number of devices in the server device 2000 and the various components of the prediction and correlation mechanism 2010. Similarly, the communication / display logic 2169 can be used to facilitate communication and compatibility between the various components of the rendering mechanism 2160 and the client device 2150.
[0248] Communication / compatibility logic 2107 can be used to facilitate dynamic communication and compatibility between server device 2000 and the following: any number and type of other computing devices (such as mobile computing devices, desktop computers, server computing devices, etc.); processing devices or components (such as CPUs, GPUs, etc.); capture / sensing / detection devices (such as capture / sensing components, including cameras, depth-sensing cameras, camera sensors, red-green-blue (RGB) sensors, microphones, etc.); display devices (such as output components, including displays, display areas, display projectors, etc.); user / context-aware components and / or identification / authentication sensors / devices (such as... Biosensors / detectors, scanners, etc.; (one or more) databases 2130, such as memory or storage devices, databases and / or data sources (such as data storage devices, hard drives, solid-state drives, hard disks, memory cards or devices, memory circuits, etc.); (one or more) communication media 2125, such as one or more communication channels or networks (e.g., cloud networks, the Internet, intranets, cellular networks, proximity networks such as Bluetooth, Bluetooth Low Energy (BLE), Bluetooth Smart, Wi-Fi proximity, radio frequency identification (RFID), near field communication (NFC), body area network (BAN), etc.); wireless or wired communication and related protocols (e.g. WiMAX, Ethernet, etc.); connectivity and location management technologies; software applications / websites (such as social and / or business networking websites, business applications, games and other entertainment applications, etc.) and programming languages, while ensuring compatibility with changing technologies, parameters, protocols, standards, etc.
[0249] Throughout this document, terms such as “logic,” “component,” “module,” “framework,” “engine,” “mechanism,” “circuit,” and “circuit” are used interchangeably and may, by way of example, include software, hardware, firmware, or any combination thereof. In one example, “logic” may refer to or include a software component that works with one or more of the following: an operating system (e.g., operating system 2006), a graphics driver (e.g., graphics driver 2016), etc., of a computing device (such as server device 2000). In another example, “logic” may refer to or include a hardware component that is capable of being part of or physically installed with one or more system hardware elements (such as an application processor (e.g., CPU 2012), a graphics processor (e.g., GPU 2014), etc.) of a computing device (such as server device 2000). In yet another embodiment, “logic” may refer to or include a firmware component that is capable of being part of the system firmware (such as firmware of an application processor (e.g., CPU 2012) or a graphics processor (e.g., GPU 2014), etc.) of a computing device (such as server device 2000). In other words, "logic" itself can be or contains circuitry or hardware components to perform certain tasks or facilitates certain circuitry or hardware components to perform certain tasks, or is executed by one or more processors (such as a graphics processor 2014, an application processor 2012, etc.) to perform certain tasks.
[0250] Additionally, terms such as "viewport," "viewpoint / user prediction," "adaptive resolution," "point cloud," "immersive media," "3DoF+," "6DoF," "GPU," "GPU domain," "GPGPU," "CPU," "CPU domain," "graphics driver," "workload," "application," "graphics pipeline," "pipeline process," "register," "register file," "RF," "extended register file," "ERF," "execution unit," "EU," "instructions," "API," and "3D API" are also relevant. Any use of specific brands, words, terms, phrases, names and / or acronyms such as “fragment shader,” “YUV texture,” “shader execution,” “existing UAV capability,” “existing backend,” “hardware,” “software,” “agent,” “graphics driver,” “kernel-mode graphics driver,” “user-mode driver,” “user-mode driver framework,” “buffer,” “graphics buffer,” “task,” “process,” “operation,” “software application,” “game,” etc., should not be construed as limiting the embodiments to software or devices labeled thereon in literature or products outside this document.
[0251] It is anticipated that components of any number and type can be added to and / or removed from the prediction and correlation mechanism 2010 and / or rendering mechanism 2160 to facilitate various embodiments, including adding, removing, and / or enhancing certain features. For the sake of brevity, clarity, and ease of understanding of the prediction and correlation mechanism 2010, many standard and / or known components (such as components of a computing device) are not shown or discussed herein. It is anticipated that the embodiments described herein are not limited to any technology, topology, system, architecture, and / or standard, and are dynamic enough to accommodate and adapt to any future changes.
[0252] References to “an embodiment,” “an embodiment,” “an example embodiment,” “various embodiments,” etc., indicate that one or more embodiments described as such may include specific features, structures, or characteristics, but not every embodiment necessarily includes the specific features, structures, or characteristics described therein. Furthermore, some embodiments may have some, all, or none of the features described with respect to other embodiments.
[0253] In the foregoing specification, embodiments have been described with reference to specific exemplary embodiments. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the embodiments set forth in the appended claims. The specification and drawings are therefore to be regarded as illustrative and not restrictive.
[0254] In the following description and claims, the term “coupled” along with its derivatives may be used. “Coupled” is used to indicate that two or more elements cooperate or interact with each other, but they may or may not have a physical or electrical component in between.
[0255] As used in the claims, unless otherwise specified, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to describe common elements merely indicates that different instances of similar elements are mentioned and is not intended to imply that the elements described so are in a given order, in time, in space, in sequence, or in any other way.
[0256] The following statements and / or examples relate to other embodiments or examples. Specific details in the examples can be used anywhere in one or more embodiments. Various features of different embodiments or examples can be combined in various ways with some of the included features and others that are excluded to suit a wide variety of different applications. Examples may include the subject matter of the embodiments and examples described herein: such as methods according to the embodiments and examples described herein, components for performing method actions, at least one machine-readable medium containing instructions that cause a machine to perform method actions when executed by a machine, or devices or systems for facilitating hybrid communication.
[0257] This invention also provides the following technical solutions: Technical Solution 1. A device comprising: One or more processors, said one or more processors being used for: Receive the user's viewing position relative to the display; The relevance of media content is analyzed based on the viewing location, wherein the media content includes immersive video of scenes captured by one or more cameras; Based on the viewing position, a portion of the media content is predicted as a relevant portion; and Transmit the relevant portion to be rendered and displayed.
[0258] Technical Solution 2. The device as described in Technical Solution 1, wherein the one or more processors further analyze the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head posture information and physical displacement information, wherein the head posture information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0259] Technical Solution 3. The device as described in Technical Solution 2, wherein the one or more processors further predict the relevant portion by identifying one or more of the following information: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0260] Technical Solution 4. The device as described in Technical Solution 1, wherein the relevant portion of the media content is predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portion includes one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and a subject-related object or region.
[0261] Technical Solution 5. The device as described in Technical Solution 2, wherein other portions of the media content are predicted as irrelevant portions that are unlikely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions.
[0262] Technical Solution 6. The device as described in Technical Solution 1, wherein the one or more processors are further configured to encode the relevant portion of the media content prior to transmitting the relevant portion, wherein the immersive video comprises one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the one or more processors include a graphics processor, wherein the graphics processor is co-located on a common semiconductor package with an application processor.
[0263] Technical Solution 7. A device comprising: One or more processors, said one or more processors being used for: Track the user's head position and movement relative to the display; Based on the tracked head position and movement, generate head pose information associated with the user; Transmit the head pose information to be evaluated in order to estimate future pose information; and Receive a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future pose information.
[0264] Technical Solution 8. The device as described in Technical Solution 7, wherein the one or more processors are further configured to: Track the user's physical displacement relative to the real world; Generate physical displacement information associated with the user based on the tracked physical displacement of the user; Transmit the physical displacement information to be evaluated in order to estimate the user's future physical displacement; and Receive the relevant portion of media content to be rendered by the display, wherein the relevant portion is further based on the future physical displacement.
[0265] Technical Solution 9. The device of Technical Solution 7, wherein the relevant portion of the media content is predicted to be more likely to be viewed by the user based on one or more of the future head posture information and the future physical displacement, wherein the relevant portion includes one or more of the following objects or areas: a central object or area in the scene, a clear object or area, a focused object or area, and a subject-related object or area, wherein the future head posture information includes one or more of the following: the future head position or movement of the head relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0266] Technical Solution 10. The device of Technical Solution 9, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in a scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the one or more processors include a graphics processor, wherein the graphics processor is co-located on a common semiconductor package with the application.
[0267] Technical Solution 11. A method comprising: Receive the user's viewing position relative to the display; The relevance of media content is analyzed based on the viewing location, wherein the media content includes immersive video of scenes captured by one or more cameras; Based on the viewing position, a portion of the media content is predicted as a relevant portion; and Transmit the relevant portion to be rendered and displayed.
[0268] Technical Solution 12. The method of Technical Solution 11, further comprising: analyzing the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head posture information and physical displacement information, wherein the head posture information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0269] Technical Solution 13. The method of Technical Solution 12 further includes: predicting the relevant portion by identifying one or more of the following: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0270] Technical Solution 14. The method of Technical Solution 11, wherein the relevant portion of the media content is predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portion includes one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and a subject-related object or region.
[0271] Technical Solution 15. The method of Technical Solution 11, wherein other portions of the media content are predicted as irrelevant portions that are unlikely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the one or more processors are further configured to encode the relevant portions of the media content before transmitting the relevant portions, wherein the immersive video includes one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the method is facilitated by one or more processors including a graphics processor and a memory coupled to an application processor, wherein the graphics processor is co-located with the application processor on a common semiconductor package.
[0272] Technical Solution 16. At least one machine-readable medium including instructions, which, when executed by a processing device, cause the processing device to perform an operation including the following: Receive the user's viewing position relative to the display; The relevance of media content is analyzed based on the viewing location, wherein the media content includes immersive video of scenes captured by one or more cameras; Based on the viewing position, a portion of the media content is predicted as a relevant portion; and Transmit the relevant portion to be rendered and displayed.
[0273] Technical Solution 17. The machine-readable medium of Technical Solution 16, wherein the operation further comprises: analyzing the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head posture information and physical displacement information, wherein the head posture information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0274] Technical Solution 18. The machine-readable medium of Technical Solution 17, wherein the operation further comprises: predicting the relevant portion by identifying one or more of the following: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0275] Technical Solution 19. The machine-readable medium as described in Technical Solution 16, wherein the relevant portion of the media content is predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portion includes one or more of the following objects or regions: a central object or region in a scene, a clear object or region, a focused object or region, and a subject-related object or region.
[0276] Technical Solution 20. The machine-readable medium of Technical Solution 16, wherein other portions of the media content are predicted as irrelevant portions that are unlikely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in a scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the one or more processors are further configured to encode the relevant portions of the media content before transmitting the relevant portions, wherein the immersive video includes one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the processing means includes a graphics processor and a memory coupled to an application processor, wherein the graphics processor is co-located with the application processor on a common semiconductor package.
[0277] Some embodiments relate to Example 1, which includes an apparatus for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the apparatus comprising: one or more processors for: receiving a viewing position relative to a display and associated with a user; analyzing the relevance of media content based on the viewing position, wherein the media content comprises immersive video of a scene captured by one or more cameras; predicting portions of the media content as relevant portions based on the viewing position; and transmitting the relevant portions to be rendered and displayed.
[0278] Example 2 includes the subject of Example 1, wherein one or more processors further analyze the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head pose information and physical displacement information, wherein the head pose information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0279] Example 3 incorporates the themes of Examples 1-2, wherein one or more processors further predict the relevant portion by identifying one or more of the following: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0280] Example 4 includes the themes of Examples 1-3, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and an object or region related to the theme.
[0281] Example 5 includes the themes of Examples 1-4, wherein other parts of the media content are predicted as irrelevant parts that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant parts include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and objects or regions unrelated to the theme.
[0282] Example 6 incorporates the themes of Examples 1-5, wherein the one or more processors are further configured to encode the relevant portion of the media content prior to transmission of the relevant portion, wherein the immersive video comprises one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the one or more processors include a graphics processor, wherein the graphics processor is co-located on a common semiconductor package with an application processor.
[0283] Some embodiments relate to Example 7, which includes an apparatus for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the apparatus comprising: one or more processors configured to: track the head position and movement of a user's head relative to a display; generate head pose information associated with the user based on the tracked head position and movement; transmit the head pose information to be evaluated in order to estimate future pose information; and receive a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future pose information.
[0284] Example 8 includes the subject of Example 7, wherein one or more processors are further configured to: track the physical displacement of the user relative to the real world; generate physical displacement information associated with the user based on the tracked physical displacement of the user; transmit the physical displacement information to be evaluated in order to estimate the future physical displacement of the user; and receive the relevant portion of media content to be rendered by the display, wherein the relevant portion is further based on the future physical displacement.
[0285] Example 9 incorporates the themes of Examples 7-8, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or areas: a central object or area in the scene, a sharp object or area, a focused object or area, and an object or area related to the theme, wherein the future head pose information includes one or more of the following: the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0286] Example 10 incorporates the themes of Examples 7-9, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the one or more processors include a graphics processor, wherein the graphics processor is co-located on a common semiconductor package with the application.
[0287] Some embodiments relate to Example 11, which includes a method for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the method comprising: receiving a viewing position relative to a display and associated with a user; analyzing the relevance of media content based on the viewing position, wherein the media content comprises immersive video of a scene captured by one or more cameras; predicting a portion of the media content as a relevant portion based on the viewing position; and transmitting the relevant portion to be rendered and displayed.
[0288] Example 12 includes the subject of Example 11, further comprising: analyzing the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head pose information and physical displacement information, wherein the head pose information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0289] Example 13 includes the subject matter of Examples 11-12, and further includes: predicting the relevant portion by identifying one or more of the following: future head pose information relative to the display based on the head pose information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head pose information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0290] Example 14 includes the themes of Examples 11-13, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and an object or region related to the theme.
[0291] Example 15 includes the themes of Examples 11-14, wherein other parts of the media content are predicted as irrelevant parts that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant parts include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions.
[0292] Example 16 incorporates the subject matter of Examples 11-15, wherein the immersive video comprises one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, and wherein the method is facilitated by a computing device comprising one or more processors, the one or more processors comprising one or more of a graphics processor and an application processor, wherein the graphics processor and the application processor are co-located on a common semiconductor package.
[0293] Some embodiments relate to Example 17, which includes a method for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the method comprising: tracking a user’s head position and movement relative to a display; generating head pose information associated with the user based on the tracked head position and movement; transmitting the head pose information to be evaluated in order to estimate future pose information; and receiving a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future pose information.
[0294] Example 18 includes the subject matter of Example 17, further comprising: tracking the user's physical displacement relative to the real world; generating physical displacement information associated with the user based on the tracked physical displacement; transmitting the physical displacement information to be evaluated in order to estimate the user's future physical displacement; and receiving the relevant portion of media content to be rendered by the display, wherein the relevant portion is further based on the future physical displacement.
[0295] Example 19 incorporates the themes of Examples 17-18, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or areas: a central object or area in the scene, a sharp object or area, a focused object or area, and a topic-related object, wherein the future head pose information includes one or more of the following: the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0296] Example 20 incorporates the themes of Examples 17-19, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the method is facilitated by a computing device comprising one or more processors, the one or more processors including one or more of a graphics processor and an application processor, wherein the graphics processor and the application processor are co-located on a common semiconductor package of the computing device.
[0297] Some embodiments relate to Example 21, which includes a data processing system comprising: a processing device for: receiving a viewing position relative to a display associated with a user; analyzing the relevance of media content based on the viewing position, wherein the media content comprises immersive video of a scene captured by one or more cameras; predicting portions of the media content as relevant portions based on the viewing position; and transmitting the relevant portions to be rendered and displayed; and a memory communicatively coupled to the processing device.
[0298] Example 22 includes the subject of Example 21, wherein the processing device further: analyzes the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head pose information and physical displacement information, wherein the head pose information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0299] Example 23 incorporates the subject matter of Examples 21-22, wherein the processing device further predicts the relevant portion by identifying one or more of the following: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0300] Example 24 includes the themes of Examples 21-23, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and an object or region related to the theme.
[0301] Example 25 incorporates the themes of Examples 21-24, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the one or more processors include a graphics processor, wherein the graphics processor is co-located on a common semiconductor package with the application.
[0302] Example 26 includes the subject of Examples 21-25, wherein the processing device is further configured to encode the relevant portion of the media content prior to transmission of the relevant portion, wherein the immersive video comprises one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the processing device includes a graphics processing unit that is co-located with a central processing unit on a common semiconductor package of the data processing system.
[0303] Some embodiments relate to Example 27, which includes a data processing system comprising: a processing device for: tracking the head position and movement of a user's head relative to a display; generating head pose information associated with the user based on the tracked head position and movement; transmitting the head pose information to be evaluated in order to estimate future pose information; and receiving a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future pose information; and a memory communicatively coupled to the processing device.
[0304] Example 28 incorporates the subject matter of Example 27, wherein the processing apparatus is further configured to: track the physical displacement of the user relative to the real world; generate physical displacement information associated with the user based on the tracked physical displacement of the user; transmit the physical displacement information to be evaluated in order to estimate the future physical displacement of the user; and receive the relevant portion of media content to be rendered by the display, wherein the relevant portion is further based on the future physical displacement.
[0305] Example 29 includes the themes of Examples 26-28, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or areas: a central object or area in the scene, a clear object or area, a focused object or area, and a theme-related object, wherein the future head pose information includes one or more of the following: the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0306] Example 30 includes the themes of Examples 26-29, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the processing device includes a graphics processing unit that is co-located with a central processing unit on a common semiconductor package of the data processing system.
[0307] Some embodiments relate to Example 31, which includes an apparatus for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the apparatus comprising: components for receiving a viewing position relative to a display and associated with a user; components for analyzing the relevance of media content based on the viewing position, wherein the media content comprises immersive video of a scene captured by one or more cameras; components for predicting portions of the media content as relevant portions based on the viewing position; and components for transmitting the relevant portions to be rendered and displayed.
[0308] Example 32 includes the subject of Example 31, further comprising: components for analyzing the relevance of the media content by evaluating location information associated with the viewing position, wherein the location information includes one or more of head pose information and physical displacement information, wherein the head pose information is based on the movement or position of the user's head relative to the display, and wherein the physical displacement information is based on the user's physical displacement in the real world.
[0309] Example 33 includes the subject matter of Examples 31-32, and further includes: a component for predicting the relevant portion by identifying one or more of the following: future head posture information relative to the display based on the head posture information and future physical displacement of the user relative to the real world based on the physical displacement information, wherein the future head posture information includes one or more of the following: future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0310] Example 34 includes the themes of Examples 31-33, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or regions: a central object or region in the scene, a clear object or region, a focused object or region, and an object or region related to the theme.
[0311] Example 35 incorporates the themes of Examples 31-34, wherein other portions of the media content are predicted as irrelevant portions that are unlikely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein one or more processors are further configured to encode the relevant portions of the media content prior to transmitting the relevant portions. Example 36 includes the subject of Examples 31-35, wherein the immersive video includes one or more of three degrees of freedom+ (3DoF+) video and six degrees of freedom (6DoF) video, wherein the device includes one or more processors, the one or more processors including one or more of a graphics processor and an application processor, wherein the graphics processor and the application processor are co-located on a common semiconductor package of the device.
[0312] Some embodiments relate to Example 37, which includes an apparatus for facilitating adaptive resolution and viewpoint prediction for immersive media in a computing environment, the apparatus comprising: components for tracking the head position and movement of a user's head relative to a display; components for generating head pose information associated with the user based on the tracked head position and movement; components for transmitting the head pose information to be evaluated in order to estimate future pose information; and components for receiving a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future pose information.
[0313] Example 38 includes the subject matter of Example 37, further comprising: components for tracking the physical displacement of the user relative to the real world; generating physical displacement information associated with the user based on the tracked physical displacement of the user; transmitting the physical displacement information to be evaluated in order to estimate the future physical displacement of the user; and receiving the relevant portion of media content to be rendered by the display, wherein the relevant portion is further based on the future physical displacement.
[0314] Example 39 includes the themes of Examples 37-38, wherein the relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the relevant portions include one or more of the following objects or areas: a central object or area in the scene, a sharp object or area, a focused object or area, and an object or area related to the theme, wherein the future head pose information includes one or more of the following: the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects and areas of interest to the user in the future.
[0315] Example 40 includes the themes of Examples 37-39, wherein other portions of the media content are predicted as irrelevant portions that are less likely to be viewed by the user based on one or more of the future head pose information and the future physical displacement, wherein the irrelevant portions include one or more of the following objects or regions: peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions, wherein the device includes one or more processors, the one or more processors including one or more of a graphics processor and an application processor, wherein the graphics processor and the application processor are co-located on a common semiconductor package of the device.
[0316] Example 41 includes at least one non-transitory or tangible machine-readable medium comprising a plurality of instructions which, when executed on a computing device, implement or perform one or more methods as claimed in any of claims or Examples 11-20.
[0317] Example 42 includes at least one machine-readable medium comprising a plurality of instructions which, when executed on a computing device, implement or perform one or more methods as claimed in any of the claims or Examples 11-20.
[0318] Example 43 includes a system comprising a mechanism that implements or performs one or more methods as claimed in any of the claims or Examples 11-20.
[0319] Example 44 includes an apparatus comprising components for performing one or more methods as claimed in any of the claims or Examples 11-20.
[0320] Example 45 includes a computing device arranged to implement or perform one or more methods as claimed in any of the claims or Examples 11-20.
[0321] Example 46 includes a communication device arranged to implement or perform one or more methods as claimed in any of the claims or Examples 11-20.
[0322] Example 47 includes at least one machine-readable medium comprising a plurality of instructions which, when executed on a computing device, implement or perform the method as claimed in any of the preceding claims, or the apparatus for implementing the claimed claim in any of the preceding claims.
[0323] Example 48 includes at least one non-transitory or tangible machine-readable medium comprising a plurality of instructions which, when executed on a computing device, implement or perform the method as claimed in any of the preceding claims, or the apparatus as claimed in any of the preceding claims.
[0324] Example 49 includes a system comprising a method for implementing or performing the claims of any of the preceding claims, or a mechanism for implementing the apparatus of any of the preceding claims.
[0325] Example 50 includes a device comprising components that perform the method as described in any of the preceding claims.
[0326] Example 51 includes a computing device arranged to implement or perform the method or apparatus as claimed in any of the preceding claims.
[0327] Example 52 includes a communication device arranged to implement or perform the method or apparatus as claimed in any of the preceding claims.
[0328] Example 53 includes a communication device arranged to implement or perform the method or apparatus as claimed in any of the preceding claims.
[0329] The accompanying drawings and the foregoing description provide examples of embodiments. Those skilled in the art will recognize that one or more of the described elements can be well combined into a single functional element. Alternatively, certain elements may be divided into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, the order of processes described herein may be changed and is not limited to the manner described herein. Moreover, the actions in any flowchart need not be performed in the order shown; nor are all actions necessarily required to be performed. Furthermore, actions unrelated to other actions may be performed in parallel with other actions. The scope of the embodiments is by no means limited to these specific examples. Numerous variations, such as differences in structure, size, and material use, are possible, whether or not explicitly stated in the specification. The scope of the embodiments is at least as broad as that given by the following claims.
Claims
1. An apparatus comprising: One or more processors, said one or more processors being used for: Track the user's head position and movement relative to the display; Based on the head tracking, the head position and movement are used to generate head pose information associated with the user. The head pose information to be evaluated is transmitted to estimate future pose information; as well as Receive a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future posture information, wherein the relevant portion includes a predicted portion of the media content, and wherein the future posture information is determined based on one or more viewing positions of the user.
2. The apparatus according to claim 1, wherein, The one or more processors are also used for: Track the user's physical displacement relative to the real world; Based on the physical displacement tracked by the user, physical displacement information associated with the user is generated; Transmit the physical displacement information to be evaluated in order to estimate the user's future physical displacement; as well as Receive the relevant portion of media content to be rendered by the display, wherein the relevant portion is also based on the future physical displacement.
3. The apparatus according to claim 1, wherein, The relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head posture information and the future physical displacement, wherein the relevant portions include one or more of the following: a central object or area in the scene, a clear object or area, a focused object or area, and a subject-related object or area; wherein the future posture information includes one or more of the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects or areas of the media content that will be of interest to the user in the future.
4. The apparatus according to claim 3, wherein, Other portions of the media content are predicted as irrelevant parts that are unlikely to be seen by the user, based on one or more of the future head pose information and the future physical displacement. The irrelevant parts include one or more of peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions. The one or more processors include a graphics processor, wherein the graphics processor and the application are positioned on a common semiconductor package.
5. A method comprising: The device uses one or more processors to track the user's head position and movement relative to the display. Based on the head tracking, the head position and movement are used to generate head pose information associated with the user. The head pose information to be evaluated is transmitted to estimate future pose information; as well as Receive a relevant portion of media content to be rendered by the display, wherein the relevant portion is based on the future posture information, wherein the relevant portion includes a predicted portion of the media content, and wherein the future posture information is determined based on one or more viewing positions of the user.
6. The method according to claim 5, further comprising: Track the user's physical displacement relative to the real world; Based on the physical displacement tracked by the user, physical displacement information associated with the user is generated; Transmit the physical displacement information to be evaluated in order to estimate the user's future physical displacement; as well as Receive the relevant portion of media content to be rendered by the display, wherein the relevant portion is also based on the future physical displacement.
7. The method according to claim 5, wherein, The relevant portions of the media content are predicted to be more likely to be viewed by the user based on one or more of the future head posture information and the future physical displacement, wherein the relevant portions include one or more of the following: a central object or area in the scene, a clear object or area, a focused object or area, and a subject-related object or area; wherein the future posture information includes one or more of the future head position or movement relative to the display, objects or areas of the media content that will be visible to the user in the future, and objects or areas of the media content that will be of interest to the user in the future.
8. The method according to claim 7, wherein, Other portions of the media content are predicted as irrelevant parts that are unlikely to be seen by the user, based on one or more of the future head pose information and the future physical displacement. The irrelevant parts include one or more of peripheral objects or regions in the scene, out-of-focus objects or regions, blurred objects or regions, and subject-irrelevant objects or regions. The one or more processors include a graphics processor, wherein the graphics processor and the application are positioned on a common semiconductor package.
9. One or more computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 5 to 8.
10. A computing device comprising means for performing the method according to any one of claims 5 to 8.
11. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 5 to 8.
Citation Information
Patent Citations
Reduced rendering of six-degree of freedom video
US20200045292A1