Dual detail video coding with region-specific quality control
By using dual detail coding technology, pixel data is encoded into low-resolution frame copies and high-resolution sub-region copies, solving the problem of data transmission rate limitation in wireless VR game systems and achieving high-quality image display at high frame rates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VALVE CORPORATION
- Filing Date
- 2024-11-20
- Publication Date
- 2026-06-26
AI Technical Summary
The data transmission rate of existing wireless protocols limits the amount of data that can be transmitted at high frame rates in wireless VR gaming systems, especially in graphics-intensive video games, where it is impossible to transmit enough pixel data in time to maintain high-quality image display.
By employing dual detail coding technology, pixel data is encoded into a low-resolution frame copy and a high-resolution sub-region copy, thereby reducing the amount of data transmitted and achieving high-quality image display using existing data rates.
High frame rate image display was achieved under limited data rates, allowing users to perceive high-quality images without reducing the image fidelity on the display device.
Smart Images

Figure CN122295935A_ABST
Abstract
Description
Background Technology
[0001] Wireless streaming technology is widely used both within and outside the video game industry. In the video game industry, some virtual reality (VR) game systems employ wireless streaming technology to leverage the high computing power of the console to run video games, while providing wearers of wireless VR headsets with greater mobility than wired (non-wireless) headsets. Despite these advantages, the data transmission rate of existing wireless protocols limits the amount of data that can be sent in a short period. This poses a challenge to VR game systems running at high frame rates, especially when running graphics-intensive video games. Summary of the Invention
[0002] This paper proposes technical solutions for improving and enhancing the above-mentioned systems and other systems. Attached Figure Description
[0003] Detailed descriptions will be provided with reference to the accompanying drawings. In the drawings, the leftmost numeral of the reference numeral indicates the drawing in which that reference numeral first appears. The same reference numerals are used in different drawings to indicate similar or identical components or features.
[0004] Figure 1 This is a schematic diagram illustrating an example distributed display system including a host and a display device in the form of a head-mounted display (HMD) according to an embodiment of the present disclosure.
[0005] Figure 2 This is a schematic diagram of an image presented on an HMD display panel according to an embodiment of the present disclosure.
[0006] Figure 3A An example stacked layout of a two-dimensional (2D) pixel array is shown according to an embodiment of the present disclosure.
[0007] Figure 3B Another example stacked layout of a 2D pixel array according to an embodiment of the present disclosure is shown.
[0008] Figure 4 An example process for implementing double detail coding in a distributed display system according to an embodiment of the present disclosure is shown.
[0009] Figure 5 This is a schematic diagram of the encoding process according to an embodiment of the present disclosure.
[0010] Figure 6 This is a schematic diagram illustrating a base frame according to an embodiment of the present disclosure and a high-attention sub-region that can be fused with the base frame to construct an image presented to the user's eye.
[0011] Figure 7 This is a schematic diagram illustrating various quality levels of different regions of the base frame and high-concern sub-regions according to embodiments of the present disclosure.
[0012] Figure 8 A flowchart illustrating an example process for generating, encoding, and transmitting video data in a distributed display system according to an embodiment of the present disclosure is shown.
[0013] Figure 9 A flowchart illustrating an example process for generating a 2D pixel array to be encoded and transmitted, according to an embodiment of the present disclosure.
[0014] Figure 10 A flowchart illustrating an example process for selecting a quality level for each of a plurality of compression units of an image, according to an embodiment of the present disclosure.
[0015] Figure 11A , Figure 11B and Figure 11C An alternative setup for a system for streaming data from a host to a display device, according to an embodiment of this disclosure, is shown.
[0016] Figure 12 Example components of a display device and host, such as an HMD (e.g., a VR headset), are shown that can implement the technologies disclosed herein. Detailed Implementation
[0017] This paper describes techniques, apparatus, and systems for streaming pixel data from a host computer to a display device using a dual-detail coding scheme. This dual-detail coding scheme enables timely transmission of pixel data, along with other data, from the host computer to the display device without compromising the subjective fidelity of the image displayed on the display device. This is particularly beneficial for streaming real-time content, such as video game content. For example, unlike pre-recorded content, real-time content cannot be transmitted in advance or cached on the client until it is ready to be displayed. Instead, frame rendering is performed milliseconds before the content is displayed, based on data received from the display device.
[0018] While the descriptions herein are not limited to gaming systems or VR systems, most of the examples described herein relate to VR gaming systems comprising a host and an HMD communicatively connected to the host. It should be understood that an HMD is merely one example of a “display device” that can be used to display streaming content, and other types of display devices are also covered herein. Furthermore, the techniques described herein can be implemented in wirelessly distributed systems and / or wired distributed systems, such as a host and display device connected by a cable. In a distributed system incorporating an HMD, a user can wear the HMD to immerse themselves in a VR or augmented reality (AR) environment, depending on the specific circumstances. One or more display panels of the HMD are configured to present images based on data generated by an application, such as a video game. The application runs on the host and generates pixel data for individual frames in a series of frames. The pixel data is encoded and sent (or transmitted) from the host to the HMD to present an image that is viewed by the user through optical components included in the HMD, allowing the user to perceive the image as if immersed in a VR or AR environment.
[0019] In some embodiments, the HMD is configured to send data to a host, which is configured to use the data in various ways, such as generating pixel data for a given frame. For example, the host can use head-tracking data received from the HMD to generate pose data indicating the predicted pose the HMD will be in at the moment the light-emitting elements of the HMD's display panel are illuminated for that frame. This pose data is fed into a running application (e.g., a video game) to generate pixel data for that given frame. In some examples, as detailed below, eye-tracking data received from the HMD can be used to determine sub-regions in a scene that will ultimately be displayed at a higher quality than the remainder of the displayed image, allowing the user to focus on the more detailed parts of the image, making the image appear sharper. Typically, the HMD receives encoded pixel data from the host and displays an image on the HMD's display panel based on the received data. As detailed below, the HMD can reconstruct and / or modify the pixel data as needed before rendering the final image so that the user perceives the content in the intended perceptual manner based on the user's head orientation and gaze.
[0020] As mentioned earlier, the limited data rates of existing communication protocols (such as wireless protocols) pose a challenge to distributed display systems, especially in distributed wireless VR gaming systems running graphics-intensive video games. For example, according to wireless protocols, typical data rates for wirelessly transmitting data from a host to a display device are 50–300 Mbps. Consider an example where the available data rate is approximately 100 Mbps. In this example, the video game running on the host might render frames at a resolution of 2.5k pixels per eye (e.g., a 2500×2500 pixel image), equivalent to 12.5 million pixels across two images, one image per eye. At 3 bytes per pixel, this means that the amount of data transmitted for a given frame is approximately 300 megabits. In a 120 Hz VR gaming system, to send all the image data in time, the host would need to send pixel data at a rate of 36,000 Mbps (36 gigabits per second). However, in a distributed system limited by a 100 Mbps data rate, this amount of pixel data cannot be transmitted quickly enough.
[0021] This document describes techniques, apparatus, and systems for implementing a dual-detail encoding scheme that enables timely transmission of pixel data and other data from a host to a display device without reducing the perceptual fidelity of the image displayed on the display device. An example process for streaming pixel data from a host to a display device may include: rendering a frame of a scene at a first resolution by a graphics rendering application running on the host; generating a 2D pixel array of the frame; encoding the 2D pixel array to obtain encoded pixel data for the frame; and sending the encoded pixel data to the display device. Specifically, in the above process, the encoded 2D pixel array provides two levels of detail for each eye. That is, for each eye, the 2D pixel array contains a copy of the frame scaled down to a second resolution lower than the first resolution, and a copy of a sub-region of the scene. And since the frame contains sub-regions, for each eye, the 2D pixel array contains two copies of the sub-region (one copy of the sub-region itself, and the other contained in the scaled-down frame copy). In some examples, the pixels representing a sub-region of a scene correspond one-to-one with the original pixels of the frame. This allows the image to be rendered via a display device where the rendering quality of the sub-region is the same or similar to that applied to the original frame. Furthermore, by sending copies of the sub-regions (which have fewer pixels than the full image) and copies of the scaled-down frame, the host can transmit a reduced amount of pixel data at data rates typically available in distributed display systems (including wireless systems). Therefore, the dual-detail encoding technique described herein significantly reduces the amount of data (i.e., the 2D pixel array) transmitted by the host, and the display device is configured to reconstruct the image from the 2D pixel array, so that the user of the display device perceives the presented image as a high-quality image (e.g., an image that appears sharp and clear).
[0022] An example process for displaying an image based on pixel data received from a host may include: receiving encoded pixel data of a frame of a scene from the host; decoding the encoded pixel data to obtain a 2D pixel array of the frame; enlarging a first copy of the frame based at least partially on a first pixel of the 2D pixel array to obtain a first enlarged copy of the frame; and enlarging a second copy of the frame based at least partially on a second pixel of the 2D pixel array to obtain a second enlarged copy of the frame. The process may continue by generating a first image based at least partially on a first enlarged copy of the frame and a third pixel of a 2D pixel array representing a first copy of a sub-region of the scene, wherein a subset of the third pixels at the edges of the sub-region is blended in the first image; and generating a second image based at least partially on a second enlarged copy of the frame and a fourth pixel of a 2D pixel array representing a second copy of the sub-region, wherein a subset of the fourth pixels at the edges of the sub-region is blended in the second image. Subsequently, the first and second images may be presented on respective display panels of a display device, such as left and right display panels, or on a single display panel.
[0023] By using the double detail encoding technique described herein to reduce the amount of data to be sent from the host to the display device, a distributed display system can stream pixel data in a timely manner to display the corresponding image on the display device, and the user perceives the image as a high-fidelity image on the display device. For example, even if parts outside a sub-region of the displayed image may be presented in relatively low quality (e.g., low resolution), the user's gaze may be directed to a sub-region of the scene with relatively high image quality, making the image appear high quality to the user. In other words, the distributed display system can implement the double detail encoding technique described herein to avoid sending pixel data of scene areas that the user is unlikely to view when the corresponding image is displayed, without the user perceiving an image quality lower than the original rendered frame.
[0024] This document also discloses devices, systems, and non-transitory computer-readable media storing computer-executable instructions to implement the techniques and processes disclosed herein. Although this document describes the disclosed techniques and systems using video game applications as an example, and specifically VR game applications, it should be understood that the techniques and systems described herein can be used in other applications to provide benefits, including but not limited to: non-VR applications (such as AR applications, mixed reality (MR) applications), and / or non-gaming applications, such as industrial machinery applications, defense applications, robotics applications, etc.
[0025] Figure 1 This is a schematic diagram illustrating an example distributed display system 100 according to embodiments disclosed herein, the system including a host 102 and a display device in the form of an HMD 104. Figure 1 It depicts user 106 wearing HMD 104. Figure 1 Example implementations of host 102 are also shown, such as in the form of a laptop computer 102 (1) carried in a backpack, or in the form of a personal computer (PC) 102 (N) located in the home of user 106. However, it should be understood that these exemplary types of host 102 do not constitute a limitation of this disclosure. For example, host 102 may be implemented as any type and / or any number of computing devices, including but not limited to PCs, laptops, desktops, PDAs, mobile phones, tablets, set-top boxes, game consoles, portable gaming devices, servers, wearable computers (such as smartwatches), or any other electronic device capable of sending and receiving data to and from other devices. Host 102 may be located in the same environment as HMD 104, such as in the home of user 106 wearing HMD 104. Alternatively, host 102 may be located remotely relative to HMD 104, for example, host 102 in the form of a server computer located in a remote geographical location different from that of HMD 104. In the remote host 102 implementation, host 102 may be communicatively connected to HMD 104 via a wide area network (such as the Internet). In the local host 102 implementation, host 102 may be located in the same environment (e.g., home) as HMD 104, wherein host 102 and HMD 104 may be directly connected to each other or connected to each other via a local area network (LAN) through an intermediate network device.
[0026] It should also be understood that the exemplary HMD 104 is merely one example type of display device that can be implemented in conjunction with the technologies and systems described herein. For example, other types and / or numbers of display devices may be used in place of, or in conjunction with, the HMD 104 described in the various examples herein. Therefore, HMD 104 may be more generally referred to herein as a display device, and it should be understood that other suitable types of display devices may exchange data with host 102. Such display devices may include, but are not limited to, televisions, portable gaming devices with displays, laptops, desktops, PDAs, mobile phones, tablets, wearable computing devices (such as smartwatches), or any other electronic device capable of sending / receiving data and displaying images on a display screen / panel.
[0027] exist Figure 1In the example, HMD 104 is communicatively connected to host 102 and configured to work collaboratively to render a given frame and present a corresponding image on display panel 108 of HMD 104. This collaborative process iterates through a series of frames to render a series of images (such as images from a VR game) on HMD 104. In the illustrated embodiment, HMD 104 includes one or more processors 110 and memory 112 (e.g., computer-readable medium 112). In some embodiments, processor 110 may include a central processing unit (CPU), a graphics processing unit (GPU) 114, both CPU and GPU 114, a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, without limitation, illustrative types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), etc. In addition, each of the processors 110 may have its own local memory, which may also store program modules, program data, and / or one or more operating systems. In some embodiments, each of the processors 110 may have its own encoder and / or decoder hardware.
[0028] Memory 112 may include volatile and non-volatile memory, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Such memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, optical disc ROM (CD-ROM) or other optical storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, redundant array of independent disks (RAID) storage systems, or any other medium that can be used to store required information and is accessible by a computing device. Memory 112 may be implemented as a computer-readable storage medium (CRSM), which may be any available physical medium accessible to processor 110 for executing instructions stored on memory 112. In one basic embodiment, the CRSM may include RAM and flash memory. In other embodiments, the CRSM may include, but is not limited to, ROM, EEPROM, or any tangible medium that can be used to store required information and is accessible by processor 110.
[0029] Typically, HMD 104 may include logic (such as software, hardware, and / or firmware) configured to implement the techniques, functions, and / or operations described herein. Computer-readable medium 112 may include various modules, such as instructions, data storage, etc., configured to run on processor 110 to perform the techniques, functions, and / or operations described herein. An example functional module in the form of compositor 116 is shown stored in computer-readable medium 112 and executable on processor 110, although the same functionality may alternatively be implemented in hardware, firmware, system-on-a-chip (SOC), and / or other logic. Furthermore, additional or different functional modules may be stored in computer-readable medium 112 and executable on processor 110. Compositor 116 is configured to modify pixels (or pixel data) to generate an image and output the modified pixel data representing the image to a frame buffer (such as a stereo frame buffer) so that the image can be rendered on display panel 108 of HMD 104. For example, compositor 116 is configured to reconstruct frames using pixel data received from host 102, as described in more detail below. Furthermore, the synthesizer 116 can apply additional modifications to the pixel data before outputting the final pixel data to the frame buffer. For example, the synthesizer 116 can be configured to make adjustments for geometric distortion, chromatic aberration, reprojection, etc. At least some of these adjustments can compensate for distortions in the near-eye optical subsystems of the HMD 104 (such as lenses and other optics), while other adjustments (such as reprojection adjustments) can compensate for minor errors in the original pose prediction of the HMD 104, and / or compensate for unmet frame rates and / or compensate for delayed packet arrival. For example, a reprojected frame can be generated using pixel data from the applied rendered frame by transforming (e.g., by rotation and reprojection calculations) the applied render frame in a manner adapted to the pose of the HMD 104 and potentially more accurate than the original pose prediction. Thus, the modified pixel data obtained by applying reprojection adjustments (and / or other adjustments) can be used to render an image for a given frame on the display panel 108 of the HMD 104, and this process can be iterated over a series of frames.
[0030] HMD 104 may also include a head tracking system 118 that generates head tracking data. The head tracking system 118 may utilize one or more sensors (e.g., infrared (IR) light sensors mounted on HMD 104) and one or more tracking beacons (e.g., IR light emitters located in the environment co-located with HMD 104) to track the head movements or motions of user 106, including head rotation. This example head tracking system 118 is not limiting, and other types of head tracking systems 118 may be used (e.g., camera-based, inertial measurement unit (IMU)-based, etc.).
[0031] HMD 104 may also include an eye-tracking system 120 that generates eye-tracking data. The eye-tracking system 120 may include, but is not limited to, cameras or other optical sensors within HMD 104 to capture image data (or information) of the user's eyes, and the eye-tracking system 120 may use the captured data / information to determine motion vectors, interpupillary distance, interocular distance, the 3D position of each eye relative to HMD 104 (including the magnitude of torsion and rotation (i.e., rolling, tilting, and swaying)), and the gaze direction of each eye. In one example, infrared light is emitted within HMD 104 and reflected from each eye. The reflected light is received or detected and analyzed by the camera of eye-tracking system 120 to extract eye rotation information from changes in the infrared light reflected by each eye. Eye-tracking system 120 may use various methods for tracking the eyes of user 106.
[0032] Figure 1 The diagram depicts HMD 104 sending data 122 to host 102. This data 122 may include, but is not limited to, the head tracking data and / or eye tracking data described above. Furthermore, data 122 may be sent via communication interface 124 of the running HMD 104, while frames are being rendered on host 102 and corresponding images are being displayed on HMD 104. Communication interface 124 of HMD 104 may include wired and / or wireless components (e.g., chips, ports, etc.) to facilitate wired and / or wireless data transmission / reception with host 102, either directly or via one or more intermediate devices (such as wireless access points (WAPs)). For example, communication interface 124 may include a wireless network interface controller. Communication interface 124 may conform to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard and may include wireless (e.g., having one or more antennas) to facilitate wireless connectivity with host 102 and / or other devices, and to transmit and receive data packets using radio frequency (RF) communication. Communication interface 124 may be built into or connected to HMD 104 (e.g., a wireless adapter, peripheral, accessory, etc.) to adapt HMD 104 for operation as a wireless communication device compliant with the 802.11 standard. In some embodiments, communication interface 124 may include a Universal Serial Bus (USB) wireless adapter (such as a USB dongle) that plugs into a USB port on HMD 104. In other embodiments, communication interface 124 may be connected to electrical components of HMD 104 via a Peripheral Component Interconnect (PCI) interface. It should be understood that communication interface 124 may also include one or more physical ports to facilitate wired connections to host 102 and / or other devices (e.g., network-connected devices communicating with other wireless networks).
[0033] Turning to host 102, host 102 is shown as including one or more processors 126 and memory 128 (e.g., computer-readable medium 128). In some embodiments, processor 126 may include a CPU, GPU 130, both CPU and GPU 130, a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, and not limited to, illustrative types of hardware logic components that may be used include FPGAs, ASICs, ASSPs, SOCs, etc. Furthermore, each of the processors 126 may have its own local memory, which may also store program modules, program data, and / or one or more operating systems. In some embodiments, each of the processors 126 may have its own encoder and / or decoder hardware. Compared to GPU 114 of HMD 104, GPU 130 of host 102 may be a higher-performance GPU with greater computing power. For example, the computing power of the GPU 114 on the HMD 104 may not be sufficient to handle the graphics of some graphics-intensive video games, which is why the system 100 adopts a distributed architecture; it leverages the relatively higher computing power of the host 102 to execute graphics-intensive video games.
[0034] The memory 128 of host 102 may include volatile and non-volatile memory, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Such memory includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM or other optical storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, RAID storage systems, or any other medium that can be used to store required information and is accessible by a computing device. Memory 128 may be implemented as a CRSM, which may be any available physical medium accessible to processor 126 for executing instructions stored on memory 128. In one basic implementation, the CRSM may include RAM and flash memory. In other implementations, the CRSM may include, but is not limited to, ROM, EEPROM, or any tangible medium that can be used to store required information and is accessible by processor 126.
[0035] Typically, host 102 may include logic (such as software, hardware, and / or firmware) configured to implement the techniques, functions, and / or operations described herein. Computer-readable medium 128 may include various modules, such as instructions, data storage, etc., configured to run on processor 126 to perform the techniques, functions, and / or operations described herein. Figure 1Example functional modules in the form of application 132 are shown. Application 132 may represent a graphics rendering application or a graphics-based application, such as video game 132 (1). In some examples, host 102 may include client applications or other software (such as a video game client) configured to execute one or more applications 132. For example, game software may be installed on host 102 to run video game 132 (1).
[0036] Typically, the communication interface 134 of host 102 may include wired and / or wireless components (such as chips, ports, etc.) to facilitate wired and / or wireless data transmission / reception with HMD 104 (or any other type of display device) directly or via one or more intermediate devices (such as WAP). In some examples, communication interface 134 may include a wireless network interface controller. Communication interface 134 may conform to the IEEE 802.11 standard and may include wireless (such as having one or more antennas) to facilitate wireless connectivity with HMD 104 and / or other devices, and to transmit and receive data packets using RF communication. Communication interface 134 may be built into host 102 or connected to host 102 (such as a wireless adapter, peripheral, accessory, etc.) to adapt host 102 for operation as a wireless communication device conforming to the 802.11 standard. In some embodiments, communication interface 134 may include a USB wireless adapter (such as a USB dongle) that is plugged into a USB port on host 102. In other embodiments, communication interface 134 may be connected to electrical components of host 102 via a PCI interface. It should be understood that the communication interface 134 may also include one or more physical ports to facilitate wired connections with the HMD 104 and / or other devices, such as plug-in devices that communicate with other wireless networks.
[0037] In some embodiments, HMD 104 may represent a VR headset used in a VR system, such as for use with a VR gaming system, in which case video game 132(1) may represent VR video game 132(1). However, HMD 104 may additionally or alternatively be implemented as an AR headset used in AR applications, an MR headset used in MR applications, or a headset suitable for VR, AR, and / or MR applications (such as industrial applications) unrelated to gaming. In AR, user 106 sees virtual objects superimposed on the real environment; while in VR, user 106 typically does not see the real environment but is fully immersed in the virtual environment as perceived via the display panel 108 and optical components (such as lenses) of HMD 104. It should be understood that in some VR systems, virtual images may be combined to display a through-image of the user 106's real environment to create an augmented VR environment in the VR system, wherein the VR environment is augmented by real images (e.g., superimposed on a virtual world). The examples described in this article primarily concern VR-based HMD 104, but it should be understood that HMD 104 is not limited to implementation in VR applications.
[0038] Typically, and as described above, the application 132 running on host 102 may represent a graphics-based application 132 (such as a video game 132 (1)). Application 132 is configured to generate pixel data for a series of frames, and the pixel data is ultimately used to render a corresponding image on the display panel 108 of HMD 104. In system 100 having HMD 104, components of host 102 may determine a predicted “light-up time” for a frame. This predicted “light-up time” represents the time it takes for the light-emitting elements of the display panel 108 of HMD 104 to light up for that frame. This prediction may take into account the estimated time required to transmit data over the wireless communication link between host 102 and HMD 104, the predicted rendering time of application 132, and / or the known scan output time of pixels from the frame buffer. In some embodiments, the prediction for the wireless communication link may differ from the prediction for the wired communication link. In some embodiments, the lighting time can be predicted as a future amount of time (e.g., approximately 20-40 milliseconds in the future), and this amount of time can vary depending on the connection type (e.g., wired versus wireless).
[0039] Host 102 may receive data 122 (e.g., head tracking data, eye tracking data, etc.) from HMD 104. Data 122 may be generated and / or transmitted at any suitable frequency, such as a frequency corresponding to the target frame rate and / or refresh rate of HMD 104, or a different (e.g., faster) frequency, such as 360 Hz (or a sensor readout every 2.7 milliseconds). Components of host 102 may determine, at least in part, the predicted pose that HMD 104 will be in at a predicted illumination time based on the head tracking data contained in data 122. Pose data indicating the predicted pose may be provided to an executing application 132 to render a frame based on the predicted pose (e.g., generate pixel data for the frame), and application 132 may output pixel data associated with the frame. The pixel data may correspond to the pixel array of display panel 108 of HMD 104. For example, the pixel data output by application 132 based on the pose data may include a two-dimensional array of per-pixel values (e.g., color values) of the pixel array on display panel 108 of HMD 104. In an illustrative example, HMD 104 may include a pair of stereoscopic display panels 108, with application 132 rendering frames of a scene at a first resolution. For example, the first resolution at which frames of each display panel 108 are rendered may be an array of 2500 × 2500 pixels. In this illustrative example, the pixel data for a single frame of display panel 108 may include 2500 × 2500 pixel values (or 12.5 million pixel values for both display panels 108). In some embodiments, the pixel data may include data per pixel, represented by a set of color and alpha values (e.g., a color value for the red channel, a color value for the green channel, a color value for the blue channel, and one or more values for one or more alpha channels).
[0040] It should be understood that the logic of host 102 may also generate additional data besides the pixel data generated for rendering an image on display panel 108 of HMD 104, and at least some of this additional data may also be sent to HMD 104. In some embodiments, the additional data may be encapsulated together with the pixel data and sent to HMD 104 as data 136. At least some of the additional data encapsulated in data 136 may be used by the logic of HMD 104 to render the image corresponding to the frame on display panel 108 of HMD 104. The additional data may include, but is not limited to, pose data indicating the predicted pose of HMD 104, depth data, motion vector data, parallax occlusion data, additional pixel data, and / or audio data. For example, in addition to generating pixel data corresponding to the image to be rendered, application 132 may also generate depth data (e.g., Z-buffer data) and / or additional pixel data (sometimes referred to herein as “out-of-bounds pixel data” or “additional pixel data”) for the frame. Alternatively or concurrently, motion vector data may be generated at least in part based on head tracking data received from HMD 104 and sent to HMD 104 to assist HMD 104 in processing pixel data. For example, motion vector data may be generated based on a comparison of head tracking data generated at two different time points (e.g., a comparison of head tracking data spaced a few milliseconds apart). The logic of HMD 104 may use some or all of the additional data to modify the pixel data to correct errors in the pose predictions made in advance by host 102. For example, the adjustments made by HMD 104 may include, but are not limited to, adjustments for geometric distortion, chromatic aberration, reprojection, etc.
[0041] Before the pixel data is sent to HMD 104, host 102 may implement the double detail encoding technique disclosed herein, which allows data 136 to be sent to HMD 104 in a timely manner under a data rate limit of approximately 100 Mbps. Figure 1As shown, compared to the number of pixels of the frame rendered by the application (which could be approximately 12.5 million pixels as described in the example above), host 102 may generate a 2D pixel array 138 with a reduced number of pixels for the frame. The 2D pixel array 138 may include a first pixel representing a first copy of frame 140(1) (scaled down to a second resolution lower than the first resolution at which application 132 renders the frame), and a second pixel representing a second copy of frame 140(2) (scaled down to a second resolution). As used herein, “scaled down” means a reduction in resolution and is sometimes referred to as “downscaling” or “downconversion.” In the illustrative example, each copy of frames 140(1), 140(2) of a single display panel 108 of HMD 104 may be scaled down from a first resolution (e.g., 2500 × 2500 pixels) to a second resolution (e.g., 500 × 500 pixels). This is merely exemplary, and any suitable amount or degree of scaling down may be implemented. The 2D pixel array 138 may also include a third pixel representing a first copy of a sub-region 142(1) of the scene, and a fourth pixel representing a second copy of the sub-region 142(2). The sub-region may be a scene region of a given frame.
[0042] In some examples, the sub-regions may be predetermined (or fixed). In VR, the user 106 looks straight ahead most of the time, meaning that the user 106 spends most of the time looking at only a very small portion of the image presented on a near-eye display (such as HMD 104). This portion of the image corresponds to the central sub-region of the scene, approximately 10 to 30 degrees off-center. In any case, the user 106 will hardly look at the edges of the scene (e.g., the corners and far edges of the image), in fact, due to the extreme angles in the near-eye display, the user's eyes cannot see the edges of the scene except through peripheral vision. Therefore, in some examples, the sub-regions of the scene may be predetermined, and the predetermined sub-regions may correspond to the central sub-region of the scene. That is, in some examples, for each frame rendered, the 2D pixel array 138 may include pixels representing copies of the sub-regions 142(1), 142(2) of the scene, which are predetermined as the central sub-region of the scene, such that for each frame generated and transmitted and for each 2D pixel array 138, the sub-region is the same part of the scene, although the scene itself changes frame by frame.
[0043] In some examples, sub-regions can be dynamically determined for each frame. For example, sub-regions for a single frame can be determined at least in part based on eye-tracking data generated by the eye-tracking system 120 of HMD 104 and received by host 102 in data 122 during runtime. The eye-tracking data received by host 102 can indicate the predicted location on display panel 108 of HMD 104 where user 106 will look when an image is presented on display panel 108 based on the coded pixel data sent to HMD 104 in data 136. In some examples, host 102 makes a prediction of where user 106 will look when the image is presented. In some examples, eye-tracking system 120 makes the prediction and sends data representing that prediction to host 102. The advantage of host 102 making the prediction is that it can be made closer to the moment the image is presented on HMD 104, and / or can utilize the relatively higher computing power of host 102 to make the prediction. In any case, in the example of determining sub-regions dynamically (e.g., instantaneously) frame-by-frame, the position of the sub-region can change frame-by-frame such that when user 106's gaze turns to the left, the sub-region is selected as the sub-region slightly to the left of the center in the scene, or when user 106's gaze turns upward, the sub-region is selected as the sub-region above the center of the scene. In other words, eye tracking can be used to guide more detailed (e.g., higher resolution) internal targets, referred to herein as sub-regions of the scene. In this way, more detailed sub-regions can be presented in the image on display panel 108 at the position where user 106 is looking, so that user 106 directly looks at the more detailed (or higher quality) portion of the image in the scene, while the less detailed (or lower quality) image is viewed (if viewed) through the user's peripheral vision.
[0044] In some examples, the pixels of each copy of subregions 142(1), 142(2) in 2D array 138 may correspond one-to-one with the original pixels of the frame representing the subregion. In other words, each copy of subregions 142(1), 142(2) in 2D array 138 may be an exact copy that has not been resampled in any way (e.g., downsampled). This means that each pixel in the scene subregion of each image rendered by the application of the frame may be copied as is to the corresponding copy of subregions 142(1), 142(2) in 2D array 138. In some examples, the copies of subregions 142(1), 142(2) in 2D array 138 have been downsampled to a reduced resolution (i.e., below the resolution at which the application 132 renders the subregion in a given frame). In these examples, the copies of subregions 142(1), 142(2) in 2D array 138 may be downsampled by ideal sampling, such as downsampling the subregion by an integer ratio (e.g., 1:2). This ideal sampling of the sub-region allows for relatively high-quality resampling when the copy of the sub-region is enlarged in the HMD 104. In other words, scaling down the copies of sub-regions 142(1) and 142(2) by an integer ratio can reduce artifacts (such as aliasing, blurring, etc.) in the corresponding parts of the image presented on the HMD 104, while scaling down the copies of sub-regions 142(1) and 142(2) by 10%, 15%, etc., may cause such artifacts to appear in later images of the HMD 104.
[0045] Copies of subregions 142(1) and 142(2) can be of any suitable size. The trade-off of making the subregions as large as possible (so that relatively low-quality parts of the image outside the subregions are less noticeable) is that more data needs to be transferred for the pixels of larger subregions. Therefore, the optimal size of the subregions may vary due to data rate limitations that may differ from implementation, and there may be practical limitations to increasing the size of the subregions. Furthermore, in some examples, application 132 may change the resolution of the rendered frames frame by frame. In other words, application 132 may render a first frame at a first resolution in a series of frames, and then application 132 may render a second frame (such as the next frame) at a second resolution different from the first resolution in a series of frames. This means that the percentage of the scene (or image) corresponding to the subregion can vary frame by frame, even if the size of the subregion is fixed at M×N pixels. In the current example, the subregion may be an 800×800 pixel subregion. Therefore, in Figure 1In the example, the 2D pixel array 138 can be an array of 800×3200 pixels (or 2.56 million pixel values for two display panels 108). Another example is a 2D pixel array 138 of 960×3840 pixels (or 3.6864 million pixel values for two display panels 108). In other words, the rendering target of the 2D array 138 can have an aspect ratio of 4:1. Therefore, generating the 2D pixel array 138 significantly reduces the number of pixels ultimately sent to the HMD 104 compared to the 12.5 million pixel values of a frame rendered at the first resolution by application 132.
[0046] Host 102 can encode a 2D pixel array 138 to obtain encoded pixel data for the frame, and host 102 can send the encoded pixel data to HMD 104 as part of data 136 sent to HMD 104, such as... Figure 1 As shown. In some examples, host 102 may utilize the encoder hardware of GPU 130 to encode 2D pixel array 138. In addition, data 136, including encoded pixel data (and possibly additional data), is sent to HMD 104 at runtime. In this way, host 102 is configured to encode and transmit images smaller than the image rendered by the application, these smaller images include redundant data and provide different levels of detail for each eye of user 106, enabling HMD 104 to reconstruct the image presented on display panel 108 of HMD 104, thereby perceiving it as a high-quality image to user 106, even though, as described above, some pixel data of a given frame is omitted (e.g., not sent) by downscaling copies of frames 140(1), 140(2) in 2D pixel array 138.
[0047] At HMD 104, after receiving data 136 including encoded pixel data (and possibly additional data), HMD 104 can decode the encoded pixel data to obtain a 2D pixel array 138 of the frame. In some examples, HMD 104 can utilize the decoder hardware of GPU 114 to decode the encoded pixel data received from host 102. Subsequently, compositor 116 of HMD 104 can modify the pixels of the 2D pixel array 138 to reconstruct the image of the frame and make adjustments for geometric distortion, chromatic aberration, reprojection, etc.
[0048] Figure 2 This is a schematic diagram illustrating images 200(1) and 200(2) respectively displayed on display panels 108(1) and 108(2) of HMD 104 according to embodiments disclosed herein. As described above, HMD 104 is configured to decode encoded pixel data received from host 102 to obtain Figure 1The frame is shown in a 2D pixel array 138, and the compositor 116 modifies the pixels of the 2D pixel array 138 to reconstruct the images 200(1), 200(2) of the frame. In order to reconstruct the right image 200(2) (e.g., the right display panel 108(2) of a pair of stereoscopic display panels 108), the compositor 116 may enlarge a first copy of frame 140(1) based at least in part on the corresponding pixels of the 2D pixel array 138 to obtain a first enlarged copy of frame 202(1), and may generate the right image 200(2) based at least in part on the pixels of the first enlarged copy of frame 202(1) and the first copy of the sub-region 142(1) representing the scene in the 2D pixel array 138, wherein a subset of the pixels of the edge 204(1) of the sub-region is blended in the right image 200(2). To reconstruct the left image from the two images, the synthesizer 116 may enlarge a second copy of frame 140(2) at least partially based on corresponding pixels of the 2D pixel array 138 to obtain a second enlarged copy of frame 202(2), and may generate a left image 200(1) at least partially based on the second enlarged copy of frame 202(2) and the pixels in the 2D pixel array 138 representing a second copy of sub-region 142(2), wherein a subset of pixels of the edge 204(2) of the sub-region is fused in the left image 200(1). The left image 200(1) and the right image 200(2) may then be presented on display panels 108(1) and 108(2) of HMD 104, respectively.
[0049] Since the copies of frames 140(1), 140(2) in the 2D pixel array 138 have been scaled down, the magnified copies of frames 202(1), 202(2) may appear blurry in the displayed images 200(1), 200(2). Specifically, the areas of images 200(1), 200(2) outside the respective sub-regions 142(1), 142(2) will appear blurry if directly viewed, but since the user 106's eyes are likely to be directly viewing the sub-regions 142(1), 142(2), the relatively low-quality images outside the sub-regions 142(1), 142(2) are likely to go unnoticed because these image parts are in the user 106's peripheral vision. By using eye tracking to dynamically determine the sub-regions (as described above), the likelihood of the user 106 noticing the relatively low-quality images outside the sub-regions 142(1), 142(2) can be reduced, but even with predetermined (or fixed) sub-regions, it is difficult for an ordinary user 106 to notice any image quality degradation caused by the double detail encoding scheme disclosed herein.
[0050] By merging edges 204(1), 204(2), the boundary between sub-regions 142(1), 142(2) and the magnified copies of frames 202(1), 202(2) becomes indistinct. That is, without merging the edges 204(1), 204(2) of the sub-regions, user 106 may notice a sharp boundary (or border) around sub-regions 142(1), 142(2) in the presented images 200(1), 200(2). The merging at the edges 204(1), 204(2) of sub-regions 142(1), 142(2) may include a gradient between the overlapping pixel data of the two regions of each image 200(1), 200(2). For example, fusion may involve interpolating between the pixel values of subregions 142(1), 142(2) and the overlapping pixel values of enlarged copies of frames 202(1), 202(2), wherein at points closer to the center of each image 200(1), 200(2), the interpolation is closer to the pixel values of subregions 142(1), 142(2), while at points farther from the image center, the interpolation gradually changes to be closer to the pixel values of enlarged copies of frames 202(1), 202(2). The specific shape of subregions 142(1), 142(2) may result in more or less fusion at edges 204(1), 204(2). The final result of the image reconstruction of each image 200 is an image 200 with three regions, including: a sharp sub-region 142, a slightly blurred region outside the sub-region 142, and a blended region between the two regions that makes the image 200 look smooth (i.e., the user 106 cannot perceive the existence of three different regions of the image 200).
[0051] Figure 3A The first example stacked layout 300A of the aforementioned 2D pixel array 138 is shown, while Figure 3B A second example stacked layout 300B of a 2D pixel array 138 is shown. These are merely exemplary stacked layouts 300, and other layouts may be implemented. Figure 3AIn the example, stacking layout 300A represents a vertical stacking layout 300A, which has pixels representing a first copy of sub-region 142(1) of the scene at the top of the 2D pixel array 138, pixels representing a first copy of frame 140(1) (scaled down to a reduced resolution) below the pixels representing the first copy of sub-region 142(1), pixels representing a second copy of sub-region 142(2) below the pixels representing the first copy of frame 140(1), and pixels representing a second copy of frame 140(2) (scaled down to a reduced resolution) at the bottom of the 2D pixel array 138. This vertical stacking layout 300A allows host 102 to send data packets carrying a portion of encoded pixel data while other pixel data is still being encoded. For example, using the encoder of GPU 130, host 102 can encode the pixels of the first copy of sub-region 142(1) and can begin sending the encoded pixel data of the first copy of sub-region 142(1) while the pixels representing the first copy of frame 140(1) are being encoded. This reduces latency in encoded pixel data transmission compared to waiting for all pixels in the 2D pixel array 138 to be encoded before sending the encoded pixel data to the HMD 104. Furthermore, it should be understood that... Figure 3A The arrangement of the four subarrays of the 2D pixel array 138 shown is merely exemplary. For example, there can be up to 24 different vertical stacking layouts and 24 different sorting combinations for these four subarrays.
[0052] Figure 3A It is also shown that each of the four groups of pixels in the 2D pixel array 138 has a dimension of M×N pixels. In other words, the first copy of subregion 142(1) is a rectangle with a horizontal dimension of M pixels and a vertical dimension of N pixels. Figure 3A As shown, the first copy of frame 140(1) is reduced to a rectangle of M pixels horizontally and N pixels vertically. The second copy of subregion 142(2) is similarly reduced to the second copy of frame 140(2). In some examples, M equals N. In some examples, M is different from N. In the example where M equals N, M and N can be 500 pixels, but this is only one example dimension that can be implemented. Therefore, as mentioned above, the 2D pixel array 138 can be 500×3200 pixels, which is far fewer pixels than the frame rendered by the application.
[0053] Figure 3BThe stacked layout 300B shown is another possible layout for pixel groups in the 2D pixel array 138. This exemplary stacked layout 300B has pixels representing a first copy (scaled down to a reduced resolution) of frame 140(1) at the top of the 2D pixel array 138, pixels representing a second copy (scaled down to a reduced resolution) of frame 140(2) below the pixels representing the first copy of frame 140(1), and pixels representing the first and second copies of sub-regions 142(1), 142(2) arranged side-by-side at the bottom of the 2D pixel array 138. Sub-arrays 142(1) and 142(2) in the stacked layout 300B may have... Figure 3A The diagram shows a horizontal dimension of M pixels and a vertical dimension of N pixels.
[0054] The processes described herein are represented as sets of blocks in a logic flowchart, which represent sequences of operations that can be implemented in hardware, software, firmware, or a combination thereof (i.e., logic). In a software context, a block represents computer-executable instructions that perform the operations when executed by one or more processors. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific abstract data type. The order in which the operations are described should not be construed as limiting; any number of the blocks can be combined in any order and / or in parallel to implement the process.
[0055] Figure 4 A flowchart illustrating an example process 400 for implementing dual-detail coding in a distributed display system 100 according to embodiments disclosed herein is shown. For ease of discussion, process 400 is described with reference to the foregoing figures.
[0056] At 402, the logic of host 102 may execute application 132 (e.g., a graphics rendering application), such as video game 132 (1), which is responsible for rendering frames in a series of frames to create video game visual content to be displayed on a display device (such as HMD 104). The following blocks 404-438 of process 400 may represent the sub-operations involved in rendering frames in a series of frames and presenting the corresponding image on the display panel 108 of HMD 104.
[0057] At 404, HMD 104 may send data 122, to be used by host 102 for rendering frames, to host 102 communicatively connected to HMD 104. In the case of HMD 104, the data sent at block 404 may include head tracking data generated by head tracking system 118, eye tracking data generated by eye tracking system 120, and / or other data. In the case of other types of display devices (such as portable gaming devices with displays), the data 122 sent at block 404 may include control input data (e.g., data corresponding to controls (such as buttons) that are pressed, touched, or otherwise operated on the gaming device). In the case of handheld controllers, controller position data and / or auxiliary input data (e.g., touch sensing data, finger tracking data, etc.) may be sent to host 102 at block 404. The position data of one or more handheld controllers may be forward-predicted to the position of the controller when the frame is displayed to user 106. In an implementation where host 102 is wirelessly connected to HMD 104, data 122 can be wirelessly transmitted from HMD 104 to host 102 at block 404. Although any suitable wireless communication protocol can be used, in some examples, the 802.11 protocol can be used to transmit data 122, 136 between HMD 104 and host 102.
[0058] At 406, host 102 may receive data 122 from a display device (such as HMD 104). Although data 122 may be received in a variety of ways depending on the implementation (e.g., wirelessly, via a wired connection, via a wide area network, etc.), in at least one embodiment, the data may be received wirelessly at block 406 via communication interface 134 of host 102.
[0059] At 408, application 132 running on host 102 can render frames of the scene at a first resolution. In some examples, frames are rendered based on a predicted on-time, which represents the time it will take for the light-emitting elements of the display panel 108 of a display device (such as HMD 104) to illuminate for a given frame. That is, the logic of host 102 can determine the actual time when photons associated with the image presented for a given frame will reach the eye of user 106. This predicted on-time is a future time (e.g., 20-40 milliseconds in the future) because it takes time for application 132 to generate the pixel data for the frame, and time for the pixel data to be encoded and transmitted from host 102 to HMD 104. It also takes time for the pixel data to be decoded, modified, and scanned out on HMD 104, and finally for the corresponding image to be presented on the display panel 108 of the display device (e.g., HMD 104).
[0060] In some examples, at block 408, a frame is rendered based on the predicted pose of HMD 104 at the predicted lighting time. For example, the head tracking system 118 of HMD 104 may be configured to track six degrees of freedom of HMD 104 (e.g., 3D position, roll, tilt, sway), which may be sent in data 122 to host 102 to determine the predicted pose of HMD 104 (e.g., considering the predicted head movement leading to the future pose of HMD 104). Therefore, pose data indicating the predicted pose may be provided to application 132 to render the frame at block 408. For example, application 132 may call a function to receive pose data, and in response to the function call, the requested pose data (predicted to the target lighting time to a given frame and predicted at least in part based on head tracking data received from HMD 104) may be provided to application 132, allowing application 132 to render the given frame based on the pose data corresponding to the virtual camera pose used for rendering the scene. Rendering at block 408 may involve application 132 generating pixel data for a frame rendered at a first resolution (e.g., the 2.5k resolution described above). The pixel data generated at block 408 may include the pixel values of individual pixels in the pixel array described herein.
[0061] At 410, if a sub-region of the scene is not predetermined (or fixed), process 400 may follow a "No" path from block 410 to block 412, where the logic of host 102 determines the sub-region based at least in part on eye-tracking data received from a display device, such as HMD 104. This eye-tracking data may have already been sent in data 122 at block 404 and may indicate where user 106 will be looking at at the predicted illumination time of that frame. Thus, if it is predicted that user 106 will be looking at the center of the display panel 108 of HMD 104, the central sub-region of the scene (or the center of the scene) may be determined as the sub-region at block 412. If it is predicted that user 106 will be looking at the left side of the display panel 108 of HMD 104, a sub-region slightly to the left of the center in the scene may be determined as the sub-region at block 412. If the sub-region of the scene is predetermined (i.e., follows the "yes" path starting at block 410), or after the sub-region is determined at block 412, process 400 may proceed to block 414.
[0062] At 414, the logic of host 102 may generate a 2D pixel array 138 for the frame. As shown in subblocks 416 and 418, the 2D pixel array 138 may include two copies of subregions 142(1), 142(2) of the scene, and two copies of frames 140(1), 140(2), which are scaled down to a second resolution lower than the first resolution at which the frame is rendered at block 408. For example, the 2D pixel array 138 may include a first pixel (e.g., a first pixel group / subarray) representing the first copy of frame 140(1) (scaled down to a reduced second resolution), a second pixel (e.g., a second pixel group / subarray) representing the second copy of frame 140(2) (scaled down to a reduced second resolution), a third pixel (e.g., a third pixel group / subarray) representing the first copy of subregion 142(1) of the scene, and a fourth pixel (e.g., a fourth pixel group / subarray) representing the second copy of subregion 142(2). As described herein, the pixels in the 2D pixel array 138 can be arranged in any suitable stacking layout, such as stacking layout 300A, stacking layout 300B, or any other suitable stacking layout. In the vertical stacking layout 300A, the third pixel is located at the top of the 2D pixel array 138, the first pixel is located below the third pixel, the fourth pixel is located below the first pixel, and the second pixel is located at the bottom of the 2D pixel array, but this is merely an exemplary stacking layout 300A that can be implemented at block 414.
[0063] As described above, the pixels of each copy of subregions 142(1), 142(2) in the 2D pixel array 138 can correspond one-to-one with the original pixels of the frame rendered at block 408. In other words, each copy of subregions 142(1), 142(2) in the 2D pixel array 138 can be an exact copy of a subregion of the rendered image of the frame, without any resampling (e.g., without being scaled down) of the applied rendered subregion. In the example where the copies of subregions 142(1), 142(2) in the 2D pixel array 138 are scaled down, the scaling down of the subregions can be performed by ideal sampling, such as scaling down the subregions by an integer ratio (e.g., 1:2), as described above.
[0064] At 420, the logic of host 102 can encode 2D pixel array 138 to obtain encoded pixel data of a frame. In some embodiments, encoding and / or decoding may be performed at least partially in hardware (e.g., in a silicon chip built into the corresponding GPU and / or CPU of the device). Thus, host 102 may utilize the encoder hardware of GPU 130 at block 420 to encode 2D pixel array 138. In some examples, encoding performed at block 420 may include compressing and / or serializing data to be sent to a display device (e.g., HMD 104) to render an image associated with the frame. In some embodiments, encoding performed at block 420 includes changing the format of the pixel data. For example, encoding at block 420 may utilize video coding standards such as High Efficiency Video Coding (HEVC) and / or extensions of HEVC, sometimes referred to as H.265 and / or MPEG-H Part 2. HEVC is an example of a standard that can be used at block 420, but it should be understood that other suitable video coding / compression standards, such as H.264 or VP9, may be used at block 420. HEVC utilizes motion estimation to compress pixel data, with the goal of transmitting less data than uncompressed pixel data in bandwidth-constrained systems, while ensuring that the compressed data is sufficient to approximate the original uncompressed data at the receiving end (e.g., in HMD 104). In some examples, the operation of generating a 2D pixel array 138 for the frame, performed at block 414, is performed as part of the encoding operation at block 420.
[0065] In some examples, the encoding at block 420 may include uniformly or variably allocating bits to encode each pixel group / subarray in the 2D pixel array 138. For example, the encoder of GPU 130 may allocate fewer bits to encode the first pixel (e.g., the first pixel group / subarray) representing the first copy of frame 140(1) (scaled down to a reduced second resolution) and the second pixel (e.g., the second pixel group / subarray) representing the second copy of frame 140(2) (scaled down to a reduced second resolution). At the same time, the encoder of GPU 130 may allocate more bits to encode the third pixel (e.g., the third pixel group / subarray) representing the first copy of subregion 142(1) of the scene and the fourth pixel (e.g., the fourth pixel group / subarray) representing the second copy of subregion 142(2). In other words, a first number of bits may be allocated to encode the first and second pixels, and a larger second number of bits may be allocated to encode the third and fourth pixels, such that more data is used to encode the copies of subregions 142(1), 142(2) than is used to encode the copies of scaled-down frames 140(1), 140(2). In some examples, the logic of host 102 can determine the subset of pixels in images 200(1), 200(2) that are occluded by pixels in subregions 142(1), 142(2) in a copy of the scaled-down frames 140(1), 140(2), and the encoder can assign zero bits to the occluded subset of pixels in images 200(1), 200(2) at block 420.
[0066] See below for reference Figures 5-10 As further described, in at least some embodiments, quality parameters are provided to the encoder at the compression unit level (e.g., per macroblock in H.264, per coded tree block in HEVC, etc.) to indicate which regions of the frame the encoder should prioritize processing.
[0067] At block 422, host 102 can transmit encoded pixel data to a display device (e.g., HMD 104) via communication interface 134. In some examples, one or more data packets carrying the encoded pixel data are wirelessly transmitted at block 422 via a transceiver of host 102. In some examples, the transmission at block 422 may use a transmission rate of approximately 100 Mbps. In other examples, the encoded pixel data may be transmitted via a wired connection.
[0068] At 424, a display device (e.g., HMD 104) can receive encoded pixel data of frames (scenes) from host 102 via communication interface 124. In some examples, one or more packets carrying the encoded pixel data are wirelessly received at block 424 via a transceiver of the display device (such as HMD 104). In some examples, reception at block 424 may use a reception rate of approximately 100 Mbps.
[0069] At 426, the logic of a display device (such as HMD 104) decodes the encoded pixel data to obtain a 2D pixel array 138 of the frame. In some embodiments, the display device (e.g., HMD 104) may utilize the decoder hardware of GPU 114 at block 426 to decode the encoded pixel data. In some examples, the decoding performed at block 426 may include decompressing and / or deserializing data to render an image associated with the frame.
[0070] At 428, the logic of a display device (such as HMD 104) (such as compositor 116) can modify the pixels in the 2D pixel array 138 to obtain modified pixel data. As shown in sub-blocks 430 and 432, at least part of the modification at block 428 may be for the purpose of reconstructing the image 200(1), 200(2) of that frame. Typically, at block 430, two copies of the previously scaled-down frames 140(1), 140(2) can be scaled up to obtain two scaled-up copies of frames 202(1), 202(2). At 432, images 200(1), 200(2) are generated based on these scaled-up copies of frames 202(1), 202(2) and the two copies of sub-regions 142(1), 142(2), and are blended at the edges 204(1), 204(2) of sub-regions 142(1), 142(2). For example, in order to reconstruct the right image 200 (2) (e.g., the right display panel 108 (2) of a pair of stereoscopic display panels 108), the synthesizer 116 may enlarge a first copy of frame 140 (1) at least partially based on the corresponding pixels of the 2D pixel array 138 in block 430 to obtain a first enlarged copy of frame 202 (1), and may generate the right image 200 (2) at least partially based on the pixels of the first enlarged copy of frame 202 (1) and the first copy of the sub-region 142 (1) representing the scene in the 2D pixel array 138 in block 432, wherein a subset of pixels at the edge 204 (1) of the sub-region is fused in the right image 200 (2). In order to reconstruct the left image in the two images, the synthesizer 116 can enlarge the second copy of frame 140 (2) at least partially based on the corresponding pixels of the 2D pixel array 138 in block 430 to obtain the second enlarged copy of frame 202 (2), and generate the left image 200 (1) at least partially based on the second enlarged copy of frame 202 (2) and the pixels in the 2D pixel array 138 representing the second copy of sub-region 142 (2), wherein a subset of pixels at the edge 204 (2) of the sub-region is fused in the left image 200 (1).
[0071] As shown in subblock 434, other modifications can be performed at block 428, such as modifications for geometric distortion, chromatic aberration, reprojection, etc. In the example, other modifications applied to the pixels at block 434 may include reprojection adjustments based at least in part on the updated pose of HMD 104 at the illumination time of a given frame. For example, the logic of HMD 104 may determine the updated pose of HMD 104 at the illumination time of a given frame based at least in part on updated head tracking data generated by HMD 104’s head tracking system 118, which represents the time during which the light-emitting elements of HMD 104’s display panel 108 are illuminated for the given frame. Because this determination is closer to the illumination time of the given frame, the pose prediction of HMD 104 may be more accurate (e.g., with smaller error) than the pose prediction determined by host 102 before the illumination time. Therefore, the display device (e.g., HMD 104) can compare the original predicted pose (based on head tracking data received by host 102 at block 406) with the updated pose, revealing the difference between the compared poses, and the reprojection adjustment can include rotation calculations to compensate for this difference (e.g., shifting and / or rotating pixel data in one or another direction based on the difference between the two poses). In other words, the head tracking data is used to "correct" the projection of the image rendered to display panel 108.
[0072] In 436, the logic of a display device (such as HMD 104) (such as compositor 116) can output modified pixels (or modified pixel data) to the frame buffer of the display device. Because the modified pixel representation... Figure 2 The images 200(1), 200(2) shown herein are also referred to herein as outputting images 200(1), 200(2) to a frame buffer. For a display device (such as HMD104) having a pair of display panels 108(1), 108(2), the modified pixel data may correspond to frames representing a pair of images to be displayed on the pair of display panels 108(1), 108(2) and are output to the stereo frame buffer accordingly.
[0073] In some examples, images 200(1), 200(2) can be output to the frame buffer via a single write operation or multiple write operations. In a single write implementation, the pixel values output to the frame buffer are output once and the pixel values are not overwritten. This single write implementation may include logic of the display device (e.g., HMD 104) such as compositor 116 performing a lookup for each pixel to determine the output pixel value based on the pixel modification performed at block 428. In a multiple write implementation, compositor 116 may output pixel values corresponding to an enlarged copy of frames 202(1), 202(2) to the frame buffer in the first instance (even those pixel values that are eventually occluded or blended in the overlap with sub-regions 142(1), 142(2),) and then output pixel values corresponding to copies of sub-regions 142(1), 142(2) to the frame buffer in the second instance, blending at the edges of sub-regions 142(1), 142(2), which may overwrite some of the previously written pixel values in the first instance. In some examples, merging is performed during the third (or third) write operation.
[0074] At 438, the logic of the display device (e.g., HMD 104) may, based on the pixel values (or pixel data) output to the frame buffer at block 436, render images 200(1), 200(2) on the corresponding display panels 108(1), 108(2) of the display device (e.g., HMD 104). This may involve scanning the pixel data output to the display panels 108(1), 108(2) of the display device (e.g., HMD 104) and illuminating the light-emitting elements of the display panels 108(1), 108(2) to illuminate the pixels on the display panels 108(1), 108(2).
[0075] Therefore, process 400 is an example technique for implementing a double detail encoding scheme in a distributed display system 100. This double detail encoding scheme reduces the amount of data transmitted from the host 102 to the display device (e.g., HMD 104), allowing pixel data to be streamed in a timely manner for displaying the corresponding image on the display device (e.g., HMD 104), and in such a way that the user 106 perceives the image as a high-fidelity image on the display device (e.g., HMD 104). For example, even if the portion outside sub-regions 142(1), 142(2) of the displayed images 200(1), 200(2) is presented in a relatively low quality (e.g., low resolution), the user 106's gaze is likely to be directed to the sub-regions 142(1), 142(2) of the scene with relatively high image quality, making it appear to the user 106 as a high-quality image.
[0076] Figures 5-10Various example processes for encoding pixel data are illustrated, where quality parameters (such as quantization parameters (QP)) are specified on a per-compression-unit basis, such as per macroblock in H.264 or per coded-tree unit in HEVC. This instructs the encoder which regions should be prioritized and which regions attract less viewer attention. This feature advantageously achieves higher compression without significantly impacting the user experience.
[0077] Figure 5 This is a block diagram of an example video encoder system or encoder 500 for a host (e.g., host 102 described above) according to embodiments disclosed herein. For example, encoder 500 can be used to implement the above-described... Figure 4 Block 420. In the example shown, encoder 500 is designed to optimize video compression by using a variable fundamental quantization parameter (QP) determined by rate control logic and an incremental QP offset value for each block of video compression unit or pixel data. This method provides efficient compression while maintaining video quality, especially in areas that attract more viewer attention. Areas that attract more attention can be the central region or video regions that attract user attention based on content, or can be dynamically determined using eye-tracking data from an eye-tracking system (e.g., eye-tracking system 120). Figure 5 The encoder 500 includes nine functional blocks that implement the encoding process from the initial video input to the final encoded output. These blocks include an input 2D pixel array block 502 (such as the 2D pixel array 138 described above), a preprocessing block 504, a block partitioning block 506, a rate control logic block 508, a basic QP determination block 510, a quantization block 512, a attention analysis logic block 514, a QP offset calculation logic block 516, and an encoded pixel data block 518 representing the encoded data of the output. In some embodiments, fewer or more blocks may be provided to achieve the desired functionality. Furthermore, the logic of some blocks may be omitted, combined with other blocks, or replaced by other functions to provide video compression and achieve specific performance for a particular application.
[0078] The input 2D pixel array block 502 is where the raw video data enters the encoder 500. This block is the starting point of the encoding process, where pixel data is first introduced into the encoder for subsequent processing and compression. As described in more detail above, the input 2D pixel array may include a left-eye base frame (including a left-eye frame scaled down to a second resolution below the first resolution), a right-eye base frame (including a right-eye frame scaled down to the second resolution), a left-eye high-attention sub-region (including a portion of the left-eye frame), and a right-eye high-attention sub-region (including a portion of the right-eye frame), according to reference... Figure 3A or Figure 3B The arrangement described herein, or any other suitable arrangement.
[0079] Preprocessing block 504 may be provided in some embodiments to perform initial processing on pixel data. This preprocessing may include noise reduction, color correction, or other preparatory operations to prepare video data for more efficient encoding. This block can improve quality and consistency before video compression.
[0080] Rate control logic block 508 analyzes various factors (such as bandwidth requirements, current transmission rate, and frame complexity) to determine the optimal base QP for the video. This logic ensures that encoder 500 balances video quality and compression efficiency, dynamically adapting to the characteristics of the video content. Base QP determination block 510 receives information from rate control logic block 508 and sets the base QP accordingly. Base QP determination block 510 establishes the quantization base level that will be applied to the entire video frame, affecting the overall compression ratio and quality. A higher base QP means more quantization, which means higher compression and lower quality.
[0081] In block partitioning block 506, pixel data is divided into smaller compression units or blocks. This partitioning allows for variable compression to be applied to different regions of the pixel data, which enables finer and more efficient coding, especially when different parts of the frame have different levels of importance, as described further below. The type and size of the blocks or compression units used can vary and depend on the specific compression algorithm used. For example, H.264 and other algorithms use macroblocks, while HEVC uses coding tree units.
[0082] Attention analysis logic block 514 evaluates which blocks in 2D pixel data are likely to attract more or less viewer attention. This analysis determines where to apply more aggressive compression (e.g., higher incremental QP offset) to effectively allocate coding resources to maintain the quality of the most needed areas. For example, specific regions in high-attention sub-regions that are considered the viewer's central focus area may be assigned zero or relatively low incremental QP, while other regions considered less attention-grabbing or having a smaller impact on image quality (e.g., edges) may be assigned higher incremental QP values. Similarly, specific regions in the base frame close to high-attention sub-regions may be assigned zero or relatively low incremental QP to provide a smooth transition, and such regions are more likely to be viewed by the user, while other regions considered less attention-grabbing (such as edges) may be assigned higher incremental QP to provide greater compression. See below. Figure 6 A- Figure 10 The determination and application of incremental QP offset are described in more detail.
[0083] Based on the output of attention analysis logic block 514, incremental QP offset calculation block 516 calculates the incremental QP offset for each block (compression unit). This offset is used to adjust the compression level of each block, increasing compression (and thus reducing quality) block by block for areas considered to attract less viewer attention.
[0084] Quantization block 512 performs the actual quantization of the video frames. Quantization block 512 applies a base QP and incremental QP offset to each block, effectively encoding the video data. For example, if the base QP is 7 and the incremental QP is 2, the quantization block applies a total QP value of 9 to that block. This block 512 balances the need for reduced data size with the need to preserve the quality of priority regions of the video stream.
[0085] Encoded pixel data block 518 is the endpoint of the encoding process performed by encoder 500. Its output compressed video stream is ready for transmission to a display device such as an HMD, as shown in the reference. Figure 4 As described in block 422.
[0086] Figure 6 This is a schematic diagram illustrating a base frame 600 according to an embodiment of the present disclosure and a high-attention sub-region 601 that can be fused with the base frame 600 to construct an image presented to the user's eye. For example, the base frame 600 may be a left-eye base frame generated from a left-eye frame of a scene generated by a graphics application, and the high-attention sub-region 601 may be a left-eye high-attention sub-region generated from a portion of the left-eye frame. Similarly, the base frame 600 may be a right-eye base frame, and the high-attention sub-region 601 may be a right-eye high-attention sub-region. The base frame 600 includes a background region 602, an overlapping region 604 that overlaps with the high-attention sub-region 601 (e.g., with a cross-gradient), an occluded region 606 covered by the high-attention sub-region 601 in the final display of the image, and invisible regions 608a-608d located in areas invisible to the user of the display device (e.g., those regions outside the area that can be imaged by the lenses of the HMD). The high-attention sub-region 601 includes a central focus region 610 and an edge region 612 surrounding the central focus region, which overlaps with the overlapping region 604 of the base frame 600.
[0087] The quality level of the compression unit for each of the base frame 600 and the high-attention sub-region 601 can be selected to provide a more pleasing visual experience to the user while offering increased compression in areas of lower attention. The following is an example procedure for adjusting the quality level of these regions using the incremental QP algorithm. It should be understood that in other implementations, different schemes can be used to influence which image regions are prioritized during the encoding process. As a non-limiting example, incremental QP may not be used, or, in addition to incremental QP, emphasis map features provided by NVIDIA or Region of Interest (ROI) features provided by AMD may be used to specify regions to be encoded at different quality levels at the compression unit level granularity.
[0088] In at least some implementations, a minimum incremental QP (e.g., 0) may be applied to the central focal region 610 of sub-region 601, which is the region that attracts the most user attention. As described elsewhere herein, the location of sub-region 601 may be static (e.g., at the center of the base frame 600 as shown) or may be dynamically determined using content from the base frame or data from the eye-tracking system. In the edge regions 612 of sub-region 601 that will intersect with the overlapping region 604 of the base frame 600, the encoder may fuse to a slightly higher incremental QP, which provides slightly higher compression because the transparency gradually increases outward from the central focal region 610. In the occluded region 606 of the base frame, which is completely covered by the highly-attention sub-region 601, a very high incremental QP (e.g., the maximum incremental QP) may be selected because this region is never visible to the user. In the overlapping region 604 of the base frame 600, a lower incremental QP (i.e., higher quality) may be selected. In the background region 602 outside the overlapping region 604, the encoder can fuse outwards from the edge of the background region 602 from the overlapping region to a higher incremental QP. Furthermore, for user-invisible corners 608a-608d in the image, the encoder can rapidly and progressively increase to a very high incremental QP value to provide the highest level of compression. Similarly, when a high-attention sub-region 601 is located near one of the corners 608a-608d based on eye-tracking data, the user-invisible portion of the high-attention sub-region 601 can rapidly and progressively increase to a very high incremental QP to provide the highest level of compression.
[0089] Figure 7 This is a schematic diagram illustrating the quality levels of different regions in a base frame 700 and a high-concern sub-region 701 that will be fused with the base frame 700 in an image displayed according to an embodiment of this disclosure. Figure 7 In this context, the density of the dots is used to represent different quality levels of blocks or compressed units that can be applied to different regions of the base frame 700 and sub-region 701, where sparse dots or no dots indicate little or no quality degradation (e.g., low incremental QP or zero incremental QP), and more dots indicate greater quality degradation (e.g., higher incremental QP).
[0090] The base frame 700 includes a background region 702, an overlapping region 704 that will intersect with the high-attention sub-region 701 for use in the displayed image, an occluded region 706 covered by the sub-region 701 in the final display of the image, and invisible regions 708a-708d located in areas invisible to the user of the display device (e.g., those regions outside the area that can be imaged by the lens of the HMD). The high-attention sub-region 701 includes a central focal region 710 and an edge region 712 surrounding the central focal region and overlapping with the overlapping region 704 of the base frame 700.
[0091] The quality level of each of the base frame 700 and the high-attention sub-region 701 can be selected to provide a comfortable visual experience for the user while offering increased compression in areas where the user might view with lower attention. In at least some embodiments, a small incremental QP or a minimum incremental QP (e.g., 0) can be applied to the central focus region 710 of the high-attention sub-region 701, which is the region receiving the most attention from the user. As described elsewhere herein, the location of the high-attention sub-region 701 can be static (e.g., at the center of the base frame 700 as shown) or can be dynamically determined based on data from the eye-tracking system or the content of the base frame. In the edge regions 712 of the sub-region 701 that will overlap (e.g., cross-gradient) with the overlapping region 704 of the base frame 700, the encoder can fuse to a slightly higher incremental QP, which provides slightly higher compression. In the illustrated example, to provide a smooth transition, a small incremental QP (e.g., QP 1 or 2) can be assigned to the inner portion of edge region 712, while a slightly higher incremental QP (e.g., QP = 2 or 3) can be assigned to the outer portion of edge region 712. In the occluded region 706 of the base frame, which is completely covered by the high-interest sub-region 701, a very high incremental QP can be selected because this region is never visible to the user. In the illustrated example, point drawing is not used for the occluded region 706 for clarity, although a very high incremental QP can be selected in this region. In the overlapping region 704 of the base frame 700, a lower incremental QP (i.e., higher quality) can be selected near the occluded region 706. In the background region 702 outside the overlapping region 704, the encoder can blend from the overlapping region 704 towards the edge of the background region 702 to a higher incremental QP. Furthermore, for corners 708a-708d in the image that are invisible to the user, the encoder can rapidly and incrementally increase the incremental QP to provide the highest level of compression. Similarly, when a high-attention sub-region 701 is located near one of the corners 708a-708d based on eye-tracking data, the user-invisible portion of the high-attention sub-region 701 can rapidly and incrementally increase the incremental QP to provide the highest level of compression.
[0092] Figure 8 A flowchart illustrating an example process 800 for generating, encoding, and transmitting video data in a distributed display system 100 according to an embodiment of the present disclosure is shown. For ease of discussion, process 800 is described with reference to the foregoing drawings.
[0093] In step 802, after the start block, the computing system receives the left-eye and right-eye frames of the scene from the graphics rendering application. The left-eye and right-eye frames are then rendered at a first resolution.
[0094] In 804, the computing system can generate a two-dimensional (2D) pixel array of the scene. An example procedure for generating a 2D pixel array is described below. Figure 9 Describe it.
[0095] In the 806, the computing system can select a quality level for each of the multiple compression units in a 2D pixel array. An example process for selecting a quality level for each of the multiple compression units is described below. Figure 10 Describe it.
[0096] At 808, the computing system can encode the 2D pixel array according to the selected quality level, and at 810, the computing system can send the encoded 2D pixel array to the display device.
[0097] Figure 9 An example process 900 for generating a 2D pixel array according to an embodiment of the present disclosure is shown. For ease of description, process 900 is described with reference to the foregoing drawings.
[0098] At 902, after the start block, the computing system generates a left-eye base frame, which includes a left-eye frame scaled down to a second resolution lower than the first resolution. At 904, the computing system generates a right-eye base frame, which includes a right-eye frame scaled down to the second resolution. At 906, the computing system generates a left-eye high-attention sub-region, which includes a portion of the left-eye frame. The left-eye high-attention sub-region is configured to be blended with the left-eye base frame when displayed by a display device. At 908, the computing system generates a right-eye high-attention sub-region, which includes a portion of the right-eye frame. The right-eye high-attention sub-region is configured to be blended with the right-eye base frame when displayed by a display device. The left-eye and right-eye high-attention sub-regions may have the same resolution (e.g., the first resolution) as the left-eye and right-eye frames received from the graphics application, or they may have different resolutions.
[0099] Figure 10 This illustration shows an embodiment of the present disclosure for providing Figure 9 The flowchart illustrates an example process 1000 for selecting a quality level for each of the high-interest sub-regions and multiple compression units in the base frame. For ease of description, process 1000 is described with reference to the aforementioned figures.
[0100] At 1002, after the start block, for each of the high-attention sub-regions of the left and right eyes, the computing system may select a first quality level for the compression units in the central focusing region of the high-attention sub-region. At 1004, the computing system may select a second quality level for the compression units in the edge regions of the high-attention sub-region surrounding the central focusing region, wherein the second quality level is lower than the first quality level. In some examples, for each of the high-attention sub-regions of the left and right eyes, the computing system may select a quality level such that there is a smooth transition between the first and second quality levels.
[0101] In 1006, for each of the left and right eye base frames, the computing system may select a third quality level for compression units in the overlapping region of the base frame that will overlap with the edge region of the corresponding high-interest sub-region. In 1008, the computing system may select a fourth quality level for compression units in the background region of the overlapping region surrounding the base frame. In some examples, for each of the left and right eye base frames, the computing system may select a quality level such that there is a smooth transition between the third and fourth quality levels.
[0102] Furthermore, as mentioned above, for occluded areas of the base frame, since these areas are invisible to the user, the computational system can select a very low quality level. Additionally, for corners of the image that are invisible to the user, the encoder can rapidly and incrementally increase to a very high incremental QP value to provide the highest level of compression. Similarly, when a high-attention sub-region is located near one of the corners based on eye-tracking data or other data, the user-invisible portion of that sub-region can rapidly and incrementally increase to a very high incremental QP to provide the highest level of compression.
[0103] In some embodiments, the encoder may determine whether the location of the high-attention sub-region is at least partially based on eye-tracking data received from the display device. In response to determining that the location of the high-attention sub-region is at least partially based on eye-tracking data received from the display device, the encoder may select a lower quality level for the background region of the base frame than would be selected if the location of the high-attention sub-region is not at least partially based on eye-tracking data received from the display device. This is because when eye tracking is used to control the location of the high-attention sub-region, it is known that the user will not be directly looking at the background region, and therefore a higher compression level can be used. This is the opposite of the case where eye tracking is not used, and although the user is likely to focus on the high-attention sub-region, they may occasionally focus on the background region of the base frame. Therefore, a lower level of compression can be used when eye-tracking data is not used to determine the high-attention sub-region.
[0104] As described above, in at least some implementations, the selection of a quality level may include selecting a corresponding offset value (e.g., an incremental QP value) of a quality parameter (e.g., QP) adaptively controlled by the rate control logic of the computing system. Furthermore, although the incremental QP feature is used to describe an example process for adjusting the quality level, it should be understood that in other implementations, different schemes may be used to influence which regions of the image are preferentially processed during the encoding process. Examples include the emphasis map feature provided by NVIDIA and the Region of Significance (ROI) feature provided by AMD.
[0105] Figure 11A , Figure 11B , Figure 11C An alternative configuration of a system for streaming data from host 102 to a display device according to an embodiment of this disclosure is shown. Brief Reference Figure 1 An example implementation is as follows: the host 102 and the display device in the form of an HMD 104 worn by the user 106 are located in the same environment. For example, when the user 106 is using the HMD 104 at home, the host 102 may be located in the user 106's home, regardless of whether the host 102 and the HMD 104 are in the same room or different rooms. Alternatively, the host 102 in the form of a mobile computing device (e.g., a tablet or laptop) may be carried (e.g., in a backpack on the user 106's back), thereby achieving greater mobility. For example, when using such a system, the user 106 may be in a park.
[0106] Figure 11A An alternative implementation is shown, in which host 102 represents one or more server computers located geographically remotely relative to HMD 104. In this case, HMD 104 may be communicatively connected to host 102 via access point (AP) 1100 such as a wireless AP (WAP), base station, USB wireless adapter, etc. In the illustrative example, data is exchanged between host 102 and HMD 104 via AP 1100 (e.g., streaming), such as streaming data over the Internet. In this implementation, host 102 and / or HMD 104 may implement one or more of the double detail coding techniques described herein.
[0107] Figure 11B Another alternative implementation is shown, in which host 102 is communicatively connected to HMD 104 via intermediate computing device 1104 (such as a laptop or tablet). Figure 11A and Figure 11B The difference is: Figure 11A AP 1100 can be used as a data routing device relative to HMD 104 and a remote server computer 102, while Figure 11B The intermediate computing device 1104 can be used as a data routing device for the host 102, which is located in the same environment as the HMD 104. The intermediate computing device 1104 can even perform a portion of the rendering workload (such as reconstructing pixel data and / or modifying pixel data, as described herein) to keep the HMD 104 as "lightweight" as possible.
[0108] Figure 11C Another alternative implementation is shown, in which the HMD 104 is replaced by another type of display device (in the form of a portable gaming device 1102 with a display). Figure 11CIn this setup, when user 106 operates the controls of portable gaming device 1102 with his / her hand, portable gaming device 1102 can send control input data to host 102, and host 102 can generate pixel data based at least in part on the control input data and send the encoded pixel data to portable gaming device 1102, as described herein primarily with respect to HMD 104. Therefore, traditional game streaming systems, which face challenges in the unfavorable environments of most home networks, can benefit from the dual-detail encoding techniques and systems described herein. For example, a home PC running game platform / application 132 can allow continuous playback of video streams and game interaction under significant rate limitations. Using game streaming services according to the techniques and systems described herein can achieve a higher quality gaming experience in bandwidth-constrained environments. Furthermore, although this document primarily refers to VR games in describing various technologies, some or all of these technologies are equally applicable to the 2D / 3D gaming domain.
[0109] Figure 12 Example components of a wearable device, such as HMD 104 (e.g., a VR headset), and host 102, which can implement the technologies disclosed herein according to embodiments of this disclosure, are shown. However, it should be understood that the relevant components described with respect to HMD 104 may be implemented in other types of display devices (e.g., portable gaming devices 1102) unless irrelevant components can be omitted from these other types of display devices. HMD 104 may be implemented as a connected device and / or a standalone device communicatively connected to host 102 during operation. In either operating mode, HMD 104 will be worn by user 106 (e.g., worn on user 106's head). In some embodiments, HMD 104 may be head-mountable, for example, by allowing user 106 to secure HMD 104 to his / her head using a fixation mechanism (e.g., an adjustable strap) sized to surround user 106's head. In some embodiments, HMD 104 includes a VR, AR, and / or MR headset with a near-eye or approximate-eye display. Therefore, the terms "wearable device," "wearable electronic device," "VR headset," "AR headset," "MR headset," and "head-mounted display (HMD)" are used interchangeably in this document to refer to... Figure 12 Device 104. However, these types of devices are merely examples of HMD 104, and it should be understood that HMD 104 can be implemented in a variety of other form factors. It should also be understood that... Figure 12 Some or all of the components shown can be implemented on HMD 104. Therefore, in some embodiments, a subset of the components shown implemented in HMD 104 may be implemented on host 102 or another computing device separate from HMD 104.
[0110] In the illustrated embodiment, the HMD 104 includes the processor 110 (which may include one or more GPUs 114), memory 112 (which stores synthesizers 116 that can be executed by the processor 110), display panel 108, head tracking system 118, eye tracking system 120, and communication interface 124.
[0111] HMD 104 may include a single display panel 108 or multiple display panels 108, such as a left display panel 108 (1) and a right display panel 108 (2) in a pair of stereoscopic display panels. One or more display panels 108 of HMD 104 are used to present a series of image frames (referred to herein as “frames”) that can be viewed by a user 106 wearing HMD 104. It should be understood that HMD 104 may include any number of display panels 108 (e.g., more than two display panels, a pair of display panels, or a single display panel). Therefore, the term “display panel” as used herein in the singular may refer to either display panel 108 of a pair of display panels in a dual-panel HMD 104, or to a single display panel 108 of an HMD 104 having any number of display panels (e.g., a single-panel HMD 104 or a multi-panel HMD 104). In a dual-panel HMD 104, a stereoscopic frame buffer can render pixels on both display panels of HMD 104. In a single-panel HMD 104, the HMD 104 may include a single display panel 108 and a pair of lenses, each lens being used for each eye to view a corresponding image displayed on a portion of the display panel 108.
[0112] The display panel 108 of the HMD 104 can employ any suitable type of display technology, such as an emissive display that emits light on the display panel 108 during frame presentation using light-emitting elements (e.g., light-emitting diodes (LEDs)) or laser illumination. As an example, the display panel 108 of the HMD 104 may include a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, an inorganic light-emitting diode (ILED) display, or any other suitable type of display technology for HMD applications.
[0113] The display panel 108 of the HMD 104 can operate at any suitable refresh rate, such as 90 Hz, 120 Hz, etc., which can be a fixed refresh rate or a variable refresh rate that dynamically changes within a refresh rate range. The "refresh rate" of a display is the number of times the display redraws the screen per second. If a fixed refresh rate is used, the number of frames displayed per second may be limited by the display's refresh rate. Therefore, a series of frames can be processed (e.g., rendered) and displayed as an image on the display, such that each screen refresh displays a single frame from the series. That is, in order to present a series of images on the display panel 108, the display panel 108 can transition frame-by-frame through the series according to the display's refresh rate, lighting up pixels with each screen refresh.
[0114] The HMD 104 display system can implement any suitable type of display driving scheme, such as a global flicker type, a scrolling band type, or any other suitable type. In a global flicker type display driving scheme, the display's light-emitting element array lights up simultaneously with each screen refresh, thus flickering globally at the refresh rate. In a scrolling band type display driving scheme, subsets of the display's light-emitting elements can be independently and sequentially lit in a scrolling band manner during the lighting period. These types of display driving schemes can be implemented by allowing individual addressability of the light-emitting elements. If the pixel array and light-emitting element array on the display panel 108 are arranged in rows and columns (but not necessarily with a one-pixel correspondence for each light-emitting element), then the rows and / or columns of the light-emitting elements can be addressed sequentially, and / or groups of consecutive rows and / or groups of consecutive columns of the light-emitting elements can be addressed sequentially for a scrolling band type display driving scheme.
[0115] Generally, as used herein, "lighting a pixel" means lighting the light-emitting element corresponding to that pixel. For example, an LCD lights the light-emitting elements of its backlight to light the corresponding pixel of the display. Furthermore, as used herein, a "subset of pixels" may include a single pixel or multiple pixels (e.g., a group of pixels). To drive the display panel 108, the HMD 104 may include a display controller, display driving circuitry, and similar electronic devices for driving the display panel 108. The display driving circuitry may be connected to the array of light-emitting elements of the display panel 108 via conductive paths (e.g., metal traces) on a flexible printed circuit. In this example, the display controller may be communicatively connected to the display driving circuitry and configured to provide signals, information, and / or data to the display driving circuitry. The signals, information, and / or data received by the display driving circuitry may cause the display driving circuitry to light the light-emitting elements in a specific manner. That is, the display controller may determine which light-emitting elements will be lit, when the elements will be lit, and the level of light output emitted by the light-emitting elements, and transmit appropriate signals, information, and / or data to the display driving circuitry to achieve this purpose.
[0116] Computer-readable medium 112 may store additional functional modules that can be executed on processor 110, although the same functionality may alternatively be implemented in hardware, firmware, SOC, and / or other logic. For example, operating system module 1200 may be configured to manage the hardware internal to HMD 104 and connected to HMD 104 for the benefit of other modules. Furthermore, in some cases, HMD 104 may include one or more applications 1202 stored in memory 112 or accessible to HMD 104. For example, applications 1202 may include, but are not limited to, video game applications (e.g., basic video games with low computational requirements for graphics processing), video playback applications (e.g., applications that access a library of video content stored on HMD 104 and / or in the cloud), etc. HMD 104 may include any number or type of applications 1202, and is not limited to the specific examples described herein.
[0117] Typically, HMD 104 has an input device 1204 and an output device 1206. Input device 1204 may include control buttons. In some embodiments, one or more microphones may be used as input device 1204 to receive audio input, such as user voice input. In some embodiments, one or more cameras or other types of sensors, such as inertial measurement units (IMUs), may be used as input device 1204 to receive gesture input, such as hand and / or head movements of user 106. In some embodiments, additional input devices 1204 in the form of a keyboard, keypad, mouse, touchscreen, joystick, etc., may be provided. In other embodiments, HMD 104 may omit a keyboard, keypad, or other similar mechanical input. In contrast, HMD 104 may be implemented with a relatively simple form of input device 1204, network interface (wireless or wired-based), power supply, and processing / memory capabilities. For example, a limited set of one or more input components (e.g., dedicated buttons for initiating configuration, power on / off, etc.) may be employed to make HMD 104 subsequently usable. In one embodiment, the input device 1204 may include control mechanisms such as basic volume control buttons for increasing / decreasing volume, as well as power and reset buttons.
[0118] Output device 1206 may include display panel 108, which may include one or more display panels 108 (e.g., a pair of stereoscopic display panels 108), as described herein. Output device 1206 may also include, but is not limited to, light-emitting elements (e.g., LEDs), vibrators that generate tactile sensations, speakers (e.g., headphones), etc. Simple light-emitting elements (e.g., LEDs) may also be present to indicate states such as, for example, when the power is on.
[0119] HMD 104 may also include a communication interface 124, which includes, but is not limited to, one or more antennas 910 (e.g., transceiver antennas) to facilitate wireless connectivity with a network and / or a second device (such as host 102 as described herein). Communication interface 124 may implement one or more of a variety of wireless technologies, such as Wi-Fi, Bluetooth, radio frequency (RF), etc. It should be understood that communication interface 124 of HMD 104 may also include physical ports to facilitate wired connectivity with a network and / or a second device (such as host 102).
[0120] HMD 104 may also include an optical subsystem 1212 that uses one or more optical elements to direct light from display panel 108 to the user's eyes. Optical subsystem 1212 may include various types and combinations of different optical elements, including but not limited to apertures, lenses (e.g., Fresnel lenses, convex lenses, concave lenses, etc.), filters, etc. In some embodiments, one or more optical elements in optical subsystem 1212 may have one or more coatings, such as an anti-reflective coating. The amplification of image light by optical subsystem 1212 allows display panel 108 to be physically smaller, lighter, and consume less power than larger displays. Furthermore, the amplification of image light increases the field of view (FOV) of the displayed content (e.g., an image). For example, the FOV of the displayed content allows for the use of almost the entire user FOV (e.g., 120-150 degrees diagonally), and in some cases, the entire user FOV, to present the displayed content. AR applications may have a narrower FOV (e.g., approximately 40 degrees FOV). The optical subsystem 1212 may be designed to correct one or more optical errors, such as, but not limited to, barrel distortion, pincushion distortion, longitudinal chromatic aberration, lateral chromatic aberration, spherical aberration, coma, field curvature, astigmatism, etc. In some embodiments, the content provided to the display panel 108 for display is pre-distorted (e.g., by applying geometric distortion adjustment and / or chromatic aberration adjustment as described herein), and the optical subsystem 1212 corrects the distortion when receiving image light generated based on the content from the display panel 108.
[0121] HMD 104 may also include one or more sensors 1214, such as sensors for generating motion, position, and orientation data. These sensors 1214 may be or include gyroscopes, accelerometers, magnetometers, cameras, color sensors, or other motion, position, and orientation sensors. Sensor 1214 may also include sub-parts of the sensor, such as a series of active or passive markers, which can be viewed externally by a camera or color sensor to generate motion, position, and orientation data. For example, a VR headset may include multiple markers on its exterior, such as reflectors or lights (e.g., infrared or visible light), which can provide one or more reference points for software to resolve to generate motion, position, and orientation data when viewed by an external camera or illuminated by light (e.g., infrared or visible light). HMD 104 may include a light sensor sensitive to light (e.g., infrared or visible light) projected or broadcast by a base station in the environment of HMD 104.
[0122] In the example, sensor 1214 may include inertial measurement unit (IMU) 1216. IMU 1216 may be an electronic device that generates calibration data based on measurement signals received from accelerometers, gyroscopes, magnetometers, and / or other sensors suitable for detecting motion, correcting errors associated with IMU 1216, or some combination thereof. Based on measurement signals from such motion-based sensors (such as IMU 1216), calibration data indicating an estimated position of HMD 104 relative to its initial position can be generated. For example, multiple accelerometers may measure translational motion (forward / backward, up / down, left / right), and multiple gyroscopes may measure rotational motion (e.g., tilting, rocking, rolling). IMU 1216 may, for example, rapidly sample the measurement signals and calculate the estimated position of HMD 104 from the sampled data. For example, IMU 1216 may integrate the measurement signals received from the accelerometers over time to estimate a velocity vector, and integrate the velocity vector over time to determine the estimated position of a reference point on HMD 104. A reference point is a point used to describe the location of HMD 104. While a reference point can generally be defined as a point in space, in various embodiments, it is defined as a point within HMD 104 (e.g., the center of IMU 1216). Alternatively, IMU 1216 provides sampled measurement signals to an external console (or other computing device) to determine calibration data. Sensor 1214 may include sensors of one or more handheld controllers as part of the HMD system. Thus, in some embodiments, the controller may transmit sensor data to host 102, and host 102 may fuse sensor data received from HMD 104 and the handheld controllers.
[0123] Sensor 1214 can operate at a relatively high frequency to provide sensor data at a high rate. For example, sensor data can be generated at a rate of 1000 Hz (or once per millisecond). In this way, one thousand reads are made per second. When the sensor generates such a large amount of data at this rate (or higher), the dataset used to predict motion is quite large, even within a relatively short time period of tens of milliseconds.
[0124] As described above, in some embodiments, sensor 1214 may include a light sensor sensitive to light emitted by a base station in the environment of HMD 104, in order to track the position and / or orientation, attitude, etc. of HMD 104 in 3D space. The position and / or orientation may be calculated based on the temporal characteristics of the light pulses and the presence or absence of light detected by sensor 1214.
[0125] HMD 104 may also include the head tracking system 118 described above. Head tracking system 118 may utilize one or more of the sensors 1214 to track the head movements of user 106, including head rotation, as described above. For example, head tracking system 118 may track six degrees of freedom of HMD 104 (i.e., 3D position, roll, tilt, and sway). These calculations may be performed at each frame of a series of frames, allowing application 132 to determine how to render the scene in the next frame based on the head position and orientation. In some embodiments, head tracking system 118 is configured to generate head tracking data used to predict the future pose (position and / or orientation) of HMD 104 based on current and / or past data, and / or on known / implicit scan output delays of individual subsets of pixels in the display system. This is because application 132 is required to render frames before user 106 actually sees the light on display panel 108 (and thus sees the image). Therefore, the next frame may be rendered based on future predictions of head position and / or orientation made at an earlier point in time. The rotation data provided by the head tracking system 118 can be used to determine both the direction of rotation of the HMD 104 and the amount of rotation of the HMD 104 in any suitable unit of measurement. For example, the direction of rotation can be simplified and output as a positive or negative horizontal direction and a positive or negative vertical direction (corresponding to left, right, up, and down). The amount of rotation can be in degrees, radians, etc. Angular velocity can be calculated to determine the rotation rate of the HMD 104.
[0126] HMD 104 may also include the aforementioned eye-tracking system 120, which generates eye-tracking data. Eye-tracking system 120 may include, but is not limited to, cameras or other optical sensors within HMD 104 to capture image data (or information) of the user's eyes, and eye-tracking system 120 may use the captured data / information to determine motion vectors, interpupillary distance, interocular distance, the 3D position of each eye relative to HMD 104, including the magnitude of twisting and rotation (i.e., rolling, tilting, and swaying), and the gaze direction of each eye. In one example, infrared light is emitted within HMD 104 and reflected from each eye. The reflected light is received or detected by the camera of eye-tracking system 120 and analyzed to extract eye rotation from changes in the infrared light reflected by each eye. Eye-tracking system 120 may use various methods for tracking the eyes of user 106. Therefore, the eye-tracking system 120 can track up to six degrees of freedom (i.e., 3D position, roll, tilt, and yaw) for each eye and can combine at least a subset of the tracking data from the two eyes of user 106 to estimate the gaze point (i.e., the 3D position or orientation in the virtual scene the user is viewing), which can be mapped to a position on display panel 108 to predict which subset (e.g., rows) or which consecutive subset (e.g., a consecutive set of rows) of pixels on display panel 108 the user 106 will view. For example, the eye-tracking system 120 can integrate measurements from the past, measurements identifying the user 106's head position, and 3D information describing the scene presented by display panel 108. Thus, information about the position and orientation of user 106's eyes is used to determine the gaze point in the virtual scene presented by HMD 104 that user 106 is viewing and to map the gaze point to a position on display panel 108 of HMD 104.
[0127] In the illustrated embodiment, the host 102 includes the processor 126 described above. The processor 126 may include one or more GPUs 130, a memory 128 for storing application 132, and a communication interface 134.
[0128] The memory 128 may also store an operating system module 1218, configured to manage the hardware inside the host 102 and the hardware connected to the host 102 for the benefit of other modules. The memory 128 may also store a video game client 1220, configured to run one or more video games within a video game library 1222, such as video game 132 (1). Video games in the video game library 1222 can be accessed and run by loading the video game client 1220. In the example, user 106 can select to play one of several video games they have purchased and downloaded to the video game library 1222 by loading the video game client 1220 and selecting video game 132 (1) to begin running video game 132 (1). The video game client 1220 may allow users to log in to the video game service using credentials such as a user account, password, etc.
[0129] The host device 102 may also include a communication interface 134, which includes, but is not limited to, one or more antennas 1224 (e.g., transceiver antennas) to facilitate wireless connectivity with a network and / or a second device (such as the HMD 104). The communication interface 134 may implement one or more of a variety of wireless technologies, such as Wi-Fi, Bluetooth, radio frequency (RF), etc. It should be understood that the communication interface 134 of the host device 102 may also include a physical port to facilitate wired connectivity with a network and / or a second device (such as the HMD 104 or other types of display devices (such as the portable gaming device 1102)).
[0130] Typically, the host 102 has an input device 1226 and an output device 1228. The input device 1226 may be a keyboard, keypad, mouse, touch screen, joystick, control button, microphone, camera, etc. The output device 1228 may include, but is not limited to, a display, light-emitting elements (such as LEDs), vibrators that generate tactile sensations, speakers (such as headphones), etc.
[0131] Although the subject matter has been described using language specific to structural features, it should be understood that the subject matter defined in the appended claims is not limited to the specific features described. Rather, these specific features are disclosed as illustrative forms of implementing the claims.
[0132] The various embodiments described above can be combined to provide further embodiments. All U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non-patent publications mentioned in this specification and / or listed in the application data list are incorporated herein by reference in their entirety. Where necessary, aspects of the embodiments may be modified to incorporate concepts from various patents, applications, and publications to provide further embodiments.
Claims
1. A computing system, comprising: One or more processors; as well as One or more memories collectively store computer-executable instructions that, when executed by the one or more processors collectively, cause the computing system to: Receive left-eye and right-eye frames of a scene from a graphics rendering application, the left-eye and right-eye frames being rendered at a first resolution; Generate a two-dimensional 2D pixel array of the scene, the 2D pixel array comprising: The left-eye base frame includes the left-eye frame scaled down to a second resolution lower than the first resolution; The right-eye base frame includes the right-eye frame scaled down to the second resolution; A left-eye high-attention sub-region, the left-eye high-attention sub-region including a portion of the left-eye frame, the left-eye high-attention sub-region being configured to merge with the left-eye base frame when displayed by a display device; and A right-eye high-attention sub-region, the right-eye high-attention sub-region including a portion of the right-eye frame, the right-eye high-attention sub-region being configured to merge with the right-eye base frame when displayed by the display device; A quality level is selected for each of the plurality of compression units in the 2D pixel array; The 2D pixel array is encoded according to the selected quality level; and The encoded 2D pixel array is sent to the display device.
2. The computing system according to claim 1, wherein, In order to select a quality level for each of the multiple compression units of the 2D pixel array, the computing system: For each of the left eye high attention sub-region and the right eye high attention sub-region, a first quality level is selected for the compression unit in the central focusing region of the high attention sub-region, and a second quality level is selected for the compression unit in the edge region of the high attention sub-region surrounding the central focusing region, wherein the second quality level is lower than the first quality level; as well as For each of the left-eye base frame and the right-eye base frame, a third quality level is selected for the compression units in the overlapping region of the base frame that will overlap with the edge region of the corresponding high-interest sub-region, and a fourth quality level is selected for the compression units in the background region of the base frame surrounding the overlapping region.
3. The computing system according to claim 2, wherein, For each of the left-eye high-attention sub-region and the right-eye high-attention sub-region, the calculation system selects a quality level such that there is a smooth transition between the first quality level and the second quality level.
4. The computing system according to claim 2, wherein, For each of the left-eye base frame and the right-eye base frame, the computing system selects a quality level such that there is a smooth transition between the third quality level and the fourth quality level.
5. The computing system according to claim 1, wherein, Selecting a quality level for the compression unit includes providing a corresponding offset value to the quality parameters adaptively controlled by the rate control logic of the computing system.
6. The computing system according to claim 1, wherein, For each of the left-eye base frame and the right-eye base frame, the computing system selects a low quality level for the area of the base frame that is determined to be invisible to the user of the display device.
7. The computing system according to claim 1, wherein, For each of the left-eye high-attention sub-region and the right-eye high-attention sub-region, the computing system selects a low-quality level for the high-attention sub-region that is determined to be an area invisible to the user of the display device.
8. The computing system according to claim 1, wherein, The computing system determines each of the left-eye high-attention sub-region and the right-eye high-attention sub-region based at least in part on eye-tracking data received from the display device.
9. The computing system according to claim 1, wherein, When executed by the one or more processors, the computer-executable instructions also cause the computing system to: Whether the determination of the corresponding positions of the left eye high attention sub-region and the right eye high attention sub-region is based at least in part on eye-tracking data received from the display device; as well as In response to the determination of the positions of the left-eye high-attention sub-region and the right-eye high-attention sub-region being based at least in part on eye-tracking data received from the display device, the quality level of the background region of the frame is selected to be lower than the quality level selected when the determination of the positions of the left-eye high-attention sub-region and the right-eye high-attention sub-region is not based at least in part on eye-tracking data received from the display device.
10. The computing system according to claim 1, wherein: The pixels of the left eye high-attention sub-region correspond one-to-one with the original pixels of the portion of the left eye frame; as well as The pixels in the right eye high-attention sub-region correspond one-to-one with the original pixels of the portion of the right eye frame.
11. The computing system according to claim 1, wherein, Generating the 2D pixel array includes arranging the pixels in a vertically stacked layout, the vertically stacked layout having: Located at the top of the 2D pixel array, in one of the left-eye high-attention sub-region and the right-eye high-attention sub-region; The left eye base frame and the right eye base frame located below one of the left eye high attention sub-region and the right eye high attention sub-region; The other of the left-eye high-attention sub-region and the right-eye high-attention sub-region, located below one of the left-eye base frame and the right-eye base frame; and The other of the left-eye base frame and the right-eye base frame, located at the bottom of the 2D pixel array.
12. The computing system according to claim 1, further comprising a transceiver, wherein, Sending the encoded 2D pixel array to the display device includes: wirelessly sending one or more data packets carrying the encoded 2D pixel array via the transceiver.
13. A method comprising: Receive left-eye and right-eye frames of a scene from a graphics rendering application, the left-eye and right-eye frames being rendered at a first resolution; Generate a two-dimensional 2D pixel array of the scene, the 2D pixel array comprising: The left-eye base frame includes the left-eye frame scaled down to a second resolution lower than the first resolution; The right-eye base frame includes the right-eye frame scaled down to the second resolution; A left-eye high-attention sub-region, the left-eye high-attention sub-region including a portion of the left-eye frame, the left-eye high-attention sub-region being configured to merge with the left-eye base frame when displayed by a display device; and A right-eye high-attention sub-region, the right-eye high-attention sub-region including a portion of the right-eye frame, the right-eye high-attention sub-region being configured to merge with the right-eye base frame when displayed by the display device; A quality level is selected for each of the plurality of compression units in the 2D pixel array; The 2D pixel array is encoded according to the selected quality level; and The encoded 2D pixel array is sent to the display device.
14. The method according to claim 13, wherein, Selecting a quality level for each of the plurality of compression units in the 2D pixel array includes: For each of the left-eye high-attention sub-region and the right-eye high-attention sub-region, a first quality level is selected for the compression unit in the central focusing region of the high-attention sub-region, and a second quality level is selected for the compression unit in the edge region of the high-attention sub-region surrounding the central focusing region, the second quality level being lower than the first quality level; and For each of the left-eye base frame and the right-eye base frame, a third quality level is selected for the compression units in the overlapping region of the base frame that will overlap with the edge region of the corresponding high-interest sub-region, and a fourth quality level is selected for the compression units in the background region of the base frame surrounding the overlapping region.
15. The method according to claim 14, wherein, For each of the left eye high attention sub-region and the right eye high attention sub-region, selecting a quality level includes selecting a quality level such that there is a smooth transition between the first quality level and the second quality level.
16. The method of claim 14, wherein, For each of the left-eye base frame and the right-eye base frame, selecting a quality level includes selecting a quality level such that there is a smooth transition between the third quality level and the fourth quality level.
17. The method according to claim 13, wherein, Selecting a quality level for the compression unit includes providing a corresponding offset value to the quality parameters adaptively controlled by the rate control logic of the method.
18. The method of claim 13, further comprising: For each of the left-eye base frame and the right-eye base frame, a low quality level is selected for the area of the base frame that is determined to be invisible to the user of the display device.
19. The method of claim 13, further comprising: For each of the left-eye high-attention sub-region and the right-eye high-attention sub-region, a low-quality level is selected for the high-attention sub-region that is determined to be an area invisible to the user of the display device.
20. The method of claim 13, further comprising: Each of the left-eye high-attention sub-region and the right-eye high-attention sub-region is determined based at least in part on eye-tracking data received from the display device.
21. The method of claim 13, further comprising: Whether the determination of the corresponding positions of the left eye high attention sub-region and the right eye high attention sub-region is based at least in part on eye-tracking data received from the display device; as well as In response to the determination of the positions of the left-eye high-attention sub-region and the right-eye high-attention sub-region being based at least in part on eye-tracking data received from the display device, the quality level of the background region of the frame is selected to be lower than the quality level selected when the determination of the positions of the left-eye high-attention sub-region and the right-eye high-attention sub-region is not based at least in part on eye-tracking data received from the display device.
22. The method according to claim 13, wherein: The pixels of the left eye high-attention sub-region correspond one-to-one with the original pixels of the portion of the left eye frame; and The pixels in the right eye high-attention sub-region correspond one-to-one with the original pixels of the portion of the right eye frame.
23. The method according to claim 13, wherein, Generating the 2D pixel array includes arranging the pixels in a vertically stacked layout, the vertically stacked layout having: Located at the top of the 2D pixel array, in one of the left-eye high-attention sub-region and the right-eye high-attention sub-region; The left eye base frame and the right eye base frame located below one of the left eye high attention sub-region and the right eye high attention sub-region; The other of the left-eye high-attention sub-region and the right-eye high-attention sub-region, located below one of the left-eye base frame and the right-eye base frame; and The other of the left-eye base frame and the right-eye base frame, located at the bottom of the 2D pixel array.
24. The method according to claim 13, wherein, Sending the encoded 2D pixel array to the display device includes: wirelessly sending one or more data packets carrying the encoded 2D pixel array.
25. A non-transitory processor-readable storage medium storing computer instructions that, when executed jointly by one or more processors, cause the one or more processors to perform an action, the action comprising: Receive left-eye and right-eye frames of a scene from a graphics rendering application, the left-eye and right-eye frames being rendered at a first resolution; Generate a two-dimensional 2D pixel array of the scene, the 2D pixel array comprising: The left-eye base frame includes the left-eye frame scaled down to a second resolution lower than the first resolution; The right-eye base frame includes the right-eye frame scaled down to the second resolution; A left-eye high-attention sub-region, the left-eye high-attention sub-region including a portion of the left-eye frame, the left-eye high-attention sub-region being configured to merge with the left-eye base frame when displayed by a display device; and A right-eye high-attention sub-region, the right-eye high-attention sub-region including a portion of the right-eye frame, the right-eye high-attention sub-region being configured to merge with the right-eye base frame when displayed by the display device; A quality level is selected for each of the plurality of compression units in the 2D pixel array; The 2D pixel array is encoded according to the selected quality level; and The encoded 2D pixel array is sent to the display device.