Methods, media, and devices for extended reality video coding
By dividing XR video frames into virtual and real regions and adjusting quantization parameters according to region complexity and user focus, the problem of insufficient bit allocation in regions of interest in existing technologies is solved, achieving efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-08-26
- Publication Date
- 2026-05-12
AI Technical Summary
Existing video coding systems cannot effectively guarantee that the region of interest receives more bits than the background region when allocating bit resources, and the computation is relatively expensive and time-consuming.
By dividing XR video frames into virtual and real regions and determining different quantization parameters for each region, the quantization parameters are adjusted according to the complexity of the region and the user's focused region input, thus achieving high bit allocation and high-resolution encoding of the virtual region.
It improves the encoding quality of regions of interest, reduces resource consumption in the background region, and lowers computational complexity and time cost.
Smart Images

Figure CN115733976B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to image processing. More specifically, but not in a limiting sense, this disclosure relates to techniques and systems for video coding. Background Technology
[0002] Some video coding systems use bitrate control algorithms to determine how many bits to allocate to specific regions of a video frame to ensure uniform picture quality for a given video coding standard and reduce the bandwidth required to transmit encoded video frames. Some bitrate control algorithms use frame-level and macroblock-level content statistics (such as complexity and contrast) to determine quantization parameters and corresponding bit allocation. Quantization parameters are integers mapped to quantization steps and control the amount of compression for each region of a video frame. For example, an 8×8 pixel region is multiplied by the quantization parameter and divided by the quantization matrix. The resulting value is then rounded to the nearest integer. Larger quantization parameters correspond to higher quantization, more compression, and lower image quality compared to smaller quantization parameters, which correspond to lower quantization, less compression, and higher image quality. Bitrate control algorithms can use constant or varying quantization parameters to adapt to a target average bitrate, constant bitrate, constant image quality, etc. However, many bitrate control algorithms are objective and cannot guarantee that more bits will be allocated to the region of interest than to the background. Some bitrate control algorithms can identify the region of interest and allocate more bits to it than to the background, but these are typically computationally expensive and time-consuming to operate. An improved technique is needed to encode video frames. Attached Figure Description
[0003] Figure 1 An example image of an extended reality (XR) video frame is shown.
[0004] Figure 2 An exemplary process for encoding extended reality video frames based on an adaptive quantization matrix is illustrated in flowchart form.
[0005] Figure 3 An example diagram of an extended reality video frame divided into virtual and real regions is shown.
[0006] Figure 4 An exemplary process for encoding extended reality video frames based on an adaptive quantization matrix and input from a gaze-tracking user interface is illustrated in flowchart form.
[0007] Figures 5A to 5C An exemplary process for encoding extended reality video frames based on an adaptive quantization matrix and a first complexity criterion and a second complexity criterion is illustrated in the form of a flowchart.
[0008] Figure 6An example diagram of extended reality video frames is shown, which are divided into regions based on the first complexity criterion and the second complexity criterion.
[0009] Figures 7A to 7C An exemplary process for encoding extended reality video frames based on an adaptive quantization matrix, a first complexity criterion, a second complexity criterion, and an adjusted region size is illustrated in flowchart form.
[0010] Figure 8 An example diagram of the middle region of an extended reality video frame is shown, which is divided into regions based on a first complexity criterion, a second complexity criterion, and an adjusted region size.
[0011] Figure 9 An exemplary system for encoding extended reality video streams is shown in block diagram form.
[0012] Figure 10 Exemplary systems used in various video encoding systems are shown, including those for encoding extended reality video streams. Detailed Implementation
[0013] This disclosure relates to systems, methods, and computer-readable media for video encoding extended reality (XR) video streams. Specifically, an XR video frame comprising a background image and at least one virtual object can be obtained. A first region of the background image to be overlaid with the at least one virtual object can be obtained from an image renderer. The XR video frame can be divided into at least one virtual region and at least one real region. The at least one virtual region includes the first region of the background image and the at least one virtual object. The at least one real region includes a second region of the background image. For each of the at least one virtual region, a corresponding first quantization parameter can be determined based on an initial quantization parameter associated with the virtual region. For each of the at least one real region, a corresponding second quantization parameter can be determined based on an initial quantization parameter associated with the real region. Each of the at least one virtual region can be encoded based on the corresponding first quantization parameter, and each of the at least one real region can be encoded based on the corresponding second quantization parameter.
[0014] Various examples of electronic systems and technologies used in connection with encoding extended reality video streams are described.
[0015] The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0016] Conversely, extended reality (XR) environments refer to fully or partially simulated environments that people perceive and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of characteristics of virtual objects in an XR environment can be done in response to a representation of physical motion (e.g., a voice command).
[0017] Humans can use any of their senses to sense and / or interact with XR objects, including sight, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, audio objects can enable audio transparency, which selectively introduces ambient sound from the physical environment, with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.
[0018] A virtual reality (VR) environment is a simulated environment designed to provide one or more senses entirely based on computer-generated sensory input. A VR environment includes multiple virtual objects that a person can sense and / or interact with. Examples of virtual objects include trees, buildings, and computer-generated images representing human avatars. A person can sense and / or interact with virtual objects in a VR environment through the simulation of their presence within the computer-generated environment and / or through the simulation of a subgroup of physical movements of a person within the computer-generated environment.
[0019] Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments are simulated environments designed to incorporate sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between a purely physical environment as one end and a virtual reality environment as the other end, but not including either end.
[0020] In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present the MR environment can track position and / or orientation relative to the physical environment, enabling virtual objects to interact with real objects (i.e., physical objects or representations of them from the physical environment). For example, the system can cause movement so that virtual trees appear stationary relative to the physical ground.
[0021] Augmented reality (AR) environments are simulated environments in which one or more virtual objects are overlaid on a physical environment or its representation. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via images or videos of the physical environment and perceive the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto a physical environment, such as as a hologram or on a physical surface, so that a person can use the system to perceive the virtual objects superimposed on the physical environment.
[0022] Augmented reality environments also refer to simulated environments where the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, a system can transform images from one or more sensors to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. Alternatively, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portion is a representative but not realistic version of the original captured image. Furthermore, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.
[0023] Augmented virtual (AV) environments are simulated environments in which a virtual or computer-generated environment is combined with one or more sensory inputs from a physical environment. Sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could adopt shadows that correspond to the sun's position within the physical environment.
[0024] Many different types of electronic systems enable people to sense and / or interact with a variety of XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped like lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may combine one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection technology, which projects graphic images onto the human retina. Projection systems can also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.
[0025] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the disclosed concepts. As part of this description, some of the accompanying drawings of this disclosure are block diagrams representing structures and devices to avoid obscuring the novel aspects of the disclosed concepts. For clarity, not all features of actual specific embodiments may be described. Additionally, as part of this specification, some of the drawings of this disclosure are provided in the form of flowcharts. The blocks in any particular flowchart may be presented in a specific order. However, it should be understood that the specific order of any given flowchart is only for illustrative purposes of one embodiment. In other embodiments, any of the various elements depicted in the flowcharts may be omitted, or the illustrated sequence of operations may be performed in a different order, or even simultaneously. Furthermore, other embodiments may include additional steps not shown as part of the flowcharts. Moreover, the language used in this disclosure has been primarily chosen for readability and instructional purposes and may not have been chosen to define or limit the subject matter of the invention, thereby resorting to the necessary claims to determine such inventive subject matter. In this disclosure, reference to “an implementation” or “implementation” means that a particular feature, structure or characteristic described in connection with that implementation is included in at least one implementation of the disclosed subject matter, and the repeated references to “an implementation” or “implementation” should not be construed as necessarily referring to all of the same implementation.
[0026] It should be understood that in any actual implementation of development (as in any software and / or hardware development project), numerous decisions must be made to achieve the developer's specific goals (e.g., compliance with system and business-related constraints), and these goals may differ between different implementations. It should also be understood that such development work can be complex and time-consuming, but nevertheless, it remains routine work for those of ordinary skill in the art who design and implement video coding systems in benefit from this disclosure.
[0027] Figure 1An example diagram of an XR video frame 100 is shown. The XR video frame 100 includes a background image 140 and a virtual object 150. The background image shows real objects, such as a dressing table 110, a rug 120, and a table 130. The virtual object overlaps with the background image 140 such that the virtual object 150 appears on top of the table 130. The background image 140 is described as a "background image" to indicate that it is behind the virtual object 150 and may have foreground and background areas. With XR video, viewers typically focus on the virtual object and the area directly surrounding it, rather than the background environment. For example, a viewer watching XR video frame 100 might focus on the virtual object 150 and a portion of the table 130 and rug 120 directly surrounding it, rather than the dressing table 110. Instead of performing computationally expensive and time-consuming image analysis on every frame of the XR video to determine the region of interest based on the image content of each frame, the video encoding system can use the known area of interest of the virtual object 150 and the background image 140 on which it is placed to determine the region of interest. Based on the virtual object 150 and its position on the background image 140, the video coding system can allocate more bits to the viewer's region of interest than to the rest of the background image 140.
[0028] Figure 2 An exemplary process 200 for encoding XR video frame 100 based on an adaptive quantization matrix is illustrated in flowchart form. For illustrative purposes, the following steps are described as being performed by a specific component. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added. For ease of explanation, refer to... Figure 1 The XR video frame 100 shown is used to describe process 200.
[0029] The flowchart begins at step 210, where the electronic device acquires an XR video frame 100 comprising a background image 140 and at least one virtual object 150. In step 220, the electronic device obtains from an image renderer a first region of the background image 140 over which the virtual object 150 is overlaid. For example, the first region of the background image 140 may indicate portions of a carpet 120 and a table 130 on which the virtual object 150 is located. In step 230, the electronic device divides the XR video frame 100 into at least one virtual region and at least one real region based on the first region of the background image 140. The virtual region includes at least a portion of the virtual object. The virtual region may further include the entire virtual object and exclude the background image or a portion thereof. For example, the virtual region may include the virtual object 150 and portions of the carpet 120 and table 130, while the real region may include the remaining portion of the background image 140 (e.g., a dressing table 110) and other portions of the carpet 120 and table 130.
[0030] In step 240, the electronic device determines a corresponding first quantization parameter for each of at least one virtual region based on an initial quantization parameter associated with the virtual region. For example, the electronic device may determine that the image complexity of a particular virtual region is greater than the image complexity of a reference virtual region associated with the initial quantization parameter of the virtual region, and proportionally decrease the initial quantization parameter. In step 250, the electronic device determines a corresponding second quantization parameter for each of at least one real region based on an initial quantization parameter associated with the real region. For example, the electronic device may determine that the image complexity of a particular real region is less than the image complexity of a reference real region associated with the initial quantization parameter of the real region, and proportionally increase the initial quantization parameter. The initial quantization parameter associated with the virtual region may be less than the initial quantization parameter associated with the real region to indicate that the detail and complexity in the virtual region are greater than the detail and complexity in the real region. That is, the initial quantization parameters associated with the virtual and real regions can be selected such that during video encoding of XR video frame 100, more bits are allocated to the virtual region corresponding to the viewer's region of interest than to the real regions outside the region of interest. In step 260, the electronic device encodes the at least one virtual region based on a first quantization parameter and encodes the at least one real region based on a second quantization parameter. The resulting encoded XR video frame allocates more bits to the at least one virtual region based on the first quantization parameter than it allocates to the at least one real region based on the second quantization parameter.
[0031] Figure 3 The diagram shows the area divided into virtual region 310 and real region 320. Figure 1An example diagram of XR video frame 100 is shown. In step 230 of process 200, the electronic device divides XR video frame 100 into a virtual region 310 and a real region 320. The virtual region 310 includes a virtual object 150 and a portion of a background image 140 surrounding the virtual object 150, showing a portion of the surface of table 130 and carpet 120. In this example, the virtual region 310 includes the entire virtual object 150 and a portion of the background image 140, but in other specific implementations, the virtual region 310 may include the entire virtual object 150 but omit a portion of the background image 140, or include a portion of the virtual object 150 and a portion of the background image 140, or include a portion of the virtual object but omit a portion of the background image 140. Negative space in the real region 320 indicates where the virtual region 310 is located. The virtual region 310 and the real region 320 may be divided into one or more additional smaller regions to allow for further refinement of quantization parameters based on complexity, contrast, etc., in different portions of regions 310 and 320.
[0032] Figure 4 An exemplary process 400 for encoding XR video frames based on an adaptive quantization matrix and input from a gaze-tracking user interface is illustrated in flowchart form. For illustrative purposes, the following steps are described as being performed by a specific component. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added. For ease of explanation, refer to the references herein. Figure 2 The process described in process 200 is used to describe process 400.
[0033] Flowchart 400 begins with steps 210 and 220, as shown above. Figure 2 The process of dividing the XR video frame into at least one virtual region and at least one real region in step 230 may optionally include steps 410 and 420. In step 410, the electronic device receives input indicating the focus area, for example via an eye-tracking user interface, a cursor-based user interface, etc. For example, in the case where the XR video frame includes multiple virtual objects, the input indicating the focus area via the eye-tracking user interface can indicate which specific virtual object the user is viewing among the multiple virtual objects.
[0034] In step 420, the electronic device divides the XR video frame into at least one virtual region and at least one real region based on the focus area. The electronic device may divide a specific virtual object and a corresponding portion of the background image over which the virtual object is superimposed into a unique virtual region, and divide the remaining virtual objects among multiple virtual objects into one or more additional virtual regions. Similarly, the electronic device may divide the remaining portion of the background image not included in the real region into one or more additional smaller regions to further refine quantization parameters based on complexity, contrast, etc., in different regions of the remaining portion of the background image.
[0035] Step 240, which determines a corresponding first quantization parameter for each of the virtual regions based on the initial quantization parameters associated with the virtual regions, may optionally include step 430. In step 430, the electronic device determines the corresponding first quantization parameter based on the focus region indicated by input from the eye-tracking user interface. For example, the first quantization parameter of the virtual region including the focus region may be smaller than the first quantization parameters of other virtual regions. That is, the virtual region including the focus region may be allocated more bits and encoded at a higher resolution than other virtual regions. The electronic device proceeds to steps 250 and 260, as referenced above. Figure 2 The method described above is based on the regions of the XR video frames divided in step 420 and the corresponding first quantization parameters determined in step 430.
[0036] Figures 5A to 5C An exemplary process 500 for encoding XR video frames based on an adaptive quantization matrix and a first complexity criterion and a second complexity criterion is illustrated in flowchart form. For illustrative purposes, the following steps are described as being performed by a specific component. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added. For ease of explanation, refer to the references herein. Figure 2 The process described in 200 and the references in this article Figure 1 The XR video frame 100 is used to describe process 500.
[0037] Flowchart 500 Figure 5A The above is a reference for China and Israel. Figure 2Steps 210, 220, and 230 begin. After dividing the XR video frame into at least one virtual region and at least one real region, the electronic device proceeds to step 510 and determines whether the at least one virtual region satisfies a first complexity criterion. The first complexity criterion may represent a threshold amount of image complexity, contrast, etc., such that a virtual region that satisfies the first complexity criterion is more complex than a virtual region that does not satisfy the first complexity criterion and is considered a complex virtual region. For example, a complex virtual region that satisfies the first complexity criterion may include highly detailed virtual objects, such as the face of a user avatar, while a virtual region that does not satisfy the first complexity criterion includes relatively simple virtual objects, such as a sphere. In response to determining that at least one virtual region satisfies the first complexity criterion, the electronic device proceeds to step 520 and determines a corresponding first quantization parameter for each of the virtual regions that satisfy the first complexity criterion (i.e., complex virtual regions) based on an initial quantization parameter associated with the complex virtual region.
[0038] The corresponding first quantization parameter can be further determined based on an upper and lower threshold associated with the complex virtual region. In response to the first quantization parameter reaching the upper or lower threshold associated with the complex virtual region, the electronic device stops determining the corresponding first quantization parameter. The upper and lower thresholds associated with the complex virtual region can be selected based on the complexity of the virtual object 150 and the background image 140, the image quality requirements associated with a given video coding standard, the time allocated to the video coding process, etc. For example, a specific video coding standard can set a valid range of values for the quantization parameter, and the upper and lower thresholds can define the boundaries of the valid range according to the specific video coding standard. As another example, the first quantization parameter can be determined during iteration, and the upper and lower thresholds can represent the maximum and minimum number of iterations that can be performed within the time allocated to the video coding process, respectively. As yet another example, the upper and lower thresholds can represent the image quality standard associated with the complex virtual region. In other words, the upper threshold can represent the maximum image quality of the complex virtual region at a specific bit rate, ensuring that the bit rate is not slowed down by additional details in the complex virtual region, and the lower threshold can represent the minimum image quality of the complex virtual region at a specific bit rate, ensuring that the minimum image quality of the complex virtual region is maintained at that specific bit rate. In step 530, the electronic device encodes each of the virtual regions that satisfy the first complexity criterion based on the corresponding first quantization parameter.
[0039] Returning to step 510, in response to determining that at least one virtual region does not satisfy the first complexity criterion, the electronic device proceeds to... Figure 5BThe process 500B shows step 550. In step 550, the electronic device determines a corresponding second quantization parameter for each of the virtual regions (i.e., simple virtual regions) that do not satisfy the first complexity criterion, based on the initial quantization parameter associated with the intermediate region. The intermediate region may include relatively simple virtual regions that do not satisfy the first complexity criterion and relatively complex real regions that satisfy the second complexity criterion. The initial quantization parameter associated with the intermediate region may be greater than the initial quantization parameter associated with the complex virtual region, such that the intermediate region is encoded using fewer bits and at a lower resolution than the number of bits used to encode the complex virtual region.
[0040] A second quantization parameter can be further determined based on an upper and lower threshold associated with the intermediate region. In response to the second quantization parameter reaching either the upper or lower threshold associated with the intermediate region, the electronic device stops determining the corresponding second quantization parameter. The upper and lower thresholds associated with the intermediate region can be selected based on the complexity of the virtual object 150 and the background image 140, the image quality requirements associated with a given video coding standard, the time allocated to the video coding process, etc. For example, a specific video coding standard can set a valid range of values for the quantization parameter, and the upper and lower thresholds can define the boundaries of the valid range according to the specific video coding standard. As another example, the second quantization parameter can be determined during iteration, and the upper and lower thresholds can represent the maximum and minimum number of iterations that can be performed within the time allocated to the video coding process, respectively. As yet another example, the upper and lower thresholds can represent the image quality standard associated with the intermediate region. In other words, the upper threshold can represent the maximum image quality of the intermediate region at a specific bit rate, ensuring that the bit rate is not slowed down by additional details in the intermediate region, and the lower threshold can represent the minimum image quality of the intermediate region at a specific bit rate, ensuring that the minimum image quality of the intermediate region remains at that specific bit rate. In some specific implementations, the maximum and minimum image quality of the intermediate region at a specific bit rate can be lower than the maximum and minimum image quality of the complex virtual region at a specific bit rate, to ensure that more bits are allocated to the complex virtual region than to the intermediate region. In step 560, the electronic device encodes each of the virtual regions that do not satisfy the first complexity criterion based on the corresponding second quantization parameter.
[0041] Returning from step 230 to at least one real region, in step 540, the electronic device determines whether the at least one real region satisfies a second complexity criterion. The second complexity criterion can represent a threshold amount of image complexity, contrast, etc., such that a real region satisfying the second complexity criterion is more complex than a real region not satisfying the second complexity criterion and is considered a complex real region or an intermediate region. A complex real region satisfying the second complexity criterion may include highly detailed portions of the background image 140, such as a portion of the background image 140 showing the legs of a table 130 resting on a carpet 120, including multiple edges and contrast in texture and color between the table 130 and the carpet 120. A real region not satisfying the second complexity criterion may include relatively simple portions of the background image 140, such as a dressing table 110 and a consistent portion of the wall and carpet 120. In response to the at least one real region satisfying the second complexity criterion, the electronic device proceeds to... Figure 5B The above-described step 550 is shown in process 500B. In step 550, the electronic device determines a corresponding second quantization parameter for each of the real regions (i.e., complex real regions) that satisfy the second complexity criterion, based on the initial quantization parameter associated with the intermediate region. The second quantization parameter of at least one real region that satisfies the second complexity criterion may be the same as or different from the second quantization parameter of at least one virtual region that does not satisfy the first complexity criterion. Then, in step 560, the electronic device encodes each of the real regions that satisfy the second complexity criterion based on the corresponding second quantization parameter.
[0042] Returning to step 540, in response to determining that at least one real region does not satisfy the second complexity criterion, the electronic device proceeds to... Figure 5C The process is shown in step 570 of process 500C. In step 570, the electronic device determines a corresponding third quantization parameter for each of the real regions (i.e., simple real regions) that do not satisfy the second complexity criterion, based on the initial quantization parameter associated with the simple real region. The initial quantization parameter associated with the simple real region may be greater than the initial quantization parameter associated with the intermediate region and the initial quantization parameter associated with the complex virtual region, such that the simple real region is encoded using fewer bits and at a lower resolution than the number of bits used to encode the intermediate region and the complex virtual region.
[0043] A third quantization parameter can be further determined based on an upper and lower threshold associated with the simple real-world region. In response to the third quantization parameter reaching the upper or lower threshold associated with the simple real-world region, the electronic device stops determining the corresponding third quantization parameter. The upper and lower thresholds associated with the simple real-world region can be selected based on the complexity of the background image 140, the image quality requirements associated with a given video coding standard, the time allocated to the video coding process, etc. For example, a specific video coding standard can set a valid range of values for the quantization parameter, and the upper and lower thresholds can define the boundaries of the valid range according to the specific video coding standard. As another example, the third quantization parameter can be determined during iteration, and the upper and lower thresholds can represent the maximum and minimum number of iterations that can be performed within the time allocated to the video coding process, respectively. As yet another example, the upper and lower thresholds can represent the image quality standard associated with the simple real-world region. In other words, the upper threshold can represent the maximum image quality of a simple real region at a specific bit rate, such that the bit rate is not slowed down by the additional details in the simple real region, and the lower threshold can represent the minimum image quality of a simple real region at a specific bit rate, such that the minimum image quality of the simple real region is maintained at that specific bit rate. In some specific implementations, the maximum and minimum image quality of the simple real region at a specific bit rate can be lower than the maximum and minimum image quality of the complex virtual region at a specific bit rate and the maximum and minimum image quality of the intermediate region at a specific ratio, to ensure that more bits are allocated to the complex virtual region and the intermediate region than to the simple real region. In step 580, the electronic device encodes each of the real regions that do not satisfy the second complexity criterion based on the corresponding third quantization parameter. Although process 500 shows three types of regions: complex virtual region, intermediate region, and simple real region, any number of region types and corresponding complexity criteria, initial quantization parameters associated with the region type, and upper and lower thresholds associated with the region type can be used alternatively.
[0044] Figure 6 It shows Figure 1The example diagram of XR video frame 100 shown is divided into regions based on the first and second complexity criteria discussed herein with respect to process 500. Virtual region 610 includes a portion of virtual object 150 and a portion of background image 140 surrounding virtual object 150, showing the surface of table 130 and a portion of carpet 120. Virtual region 610 satisfies the first complexity criterion and is therefore encoded using a first quantization parameter. Simple real region 620 includes a portion of background image 140 that does not satisfy the second complexity criterion and shows a portion of dressing table 110, carpet 120, and table 130. Simple real region 620 is encoded using a third quantization parameter. Intermediate region 630 includes a portion of background image 140 that satisfies the second complexity criterion and shows the leg of table 130 resting on a portion of carpet 120. Intermediate region 630 is encoded using a second quantization parameter. Negative space in simple real region 620 indicates where virtual region 610 and intermediate region 630 are located. The virtual region 610, the simple real region 620, and the intermediate region 630 can be divided into one or more additional smaller regions to allow for further refinement of quantization parameters based on complexity, contrast, etc., in different parts of each region.
[0045] Figures 7A to 7C An exemplary process 700 for encoding XR video frames based on an adaptive quantization matrix, a first complexity criterion, a second complexity criterion, and an adjusted region size is illustrated in flowchart form. For illustrative purposes, the following steps are described as being performed by a specific component. However, it should be understood that various actions can be performed by alternative components. Furthermore, various actions can be performed in different orders. Additionally, some actions can be performed simultaneously, and some actions may be unnecessary, or additional actions may be added. For ease of explanation, refer to the references herein. Figure 2 The process described in 200 and the references in this article Figures 5A to 5C The process described in process 500 is used to describe process 700.
[0046] Flowchart 700 Figure 7A The above is a reference for China and Israel. Figure 2 Steps 210, 220, and 230 begin. After dividing the XR video frame into at least one virtual region and at least one real region, the electronic device proceeds to step 510 and determines whether the at least one virtual region satisfies the first complexity criterion, as referenced above. Figure 5AThe process 500A shown is described. In response to determining that at least one virtual region satisfies a first complexity criterion, the electronic device may optionally proceed to step 710, and determine a corresponding region size for each of the virtual regions satisfying the first complexity criterion based on an initial region size associated with the complex virtual region. The region sizes may be selected such that the complex portions of the XR video frame have smaller region sizes, while the simple portions of the XR video frame have larger region sizes.
[0047] Then, in step 720, the electronic device may optionally divide a particular virtual region into one or more additional virtual regions for each of the virtual regions satisfying the first complexity criterion and based on the corresponding region size. The electronic device proceeds to step 520 and determines a corresponding first quantization parameter for each of the virtual regions satisfying the first complexity criterion and the additional virtual regions based on the initial quantization parameters associated with the complex virtual region, as referenced above. Figure 5A The process 500A shown is described above. In step 530, the electronic device encodes each of the virtual region and the additional virtual region that satisfy the first complexity criterion based on the corresponding first quantization parameter, as referenced above. Figure 5A The process 500A shown is described.
[0048] Returning to step 510, in response to determining that at least one virtual region does not satisfy the first complexity criterion, in Figure 7B In step 730 of process 700B, the electronic device may optionally determine a corresponding region size for each of the virtual regions that does not satisfy the first complexity criterion based on an initial region size associated with the intermediate region. The initial region size associated with the intermediate region may be larger than the initial region size associated with the complex virtual region. Then, in step 740, the electronic device may optionally divide the specific region into one or more additional regions for each of the virtual regions that does not satisfy the first complexity criterion and based on the corresponding region size.
[0049] In step 550, the electronic device determines a corresponding second quantization parameter for each of the virtual regions and the additional virtual regions that do not satisfy the first complexity criterion, based on the initial quantization parameters associated with the intermediate region, as referenced above. Figure 5B The process described in step 500B is as follows. In step 560, the electronic device encodes each of the virtual regions and the additional virtual regions that do not satisfy the first complexity criterion based on the corresponding second quantization parameters, as referenced above. Figure 5B The process shown in 500B is described.
[0050] Returning from step 230 to at least one real region, in step 540, the electronic device determines whether the at least one real region satisfies the second complexity criterion, as referenced above. Figure 5A The process 500A shown is described. In response to the at least one real region satisfying the second complexity criterion, the electronic device may optionally proceed to... Figure 7B The above-described step 730 is shown in process 700B. In step 730, the electronic device may optionally determine a corresponding region size for each of the real regions that satisfies the second complexity criterion based on the initial region size associated with the intermediate region. The electronic device may optionally proceed to step 740 and divide a particular real region into one or more additional real regions for each of the real regions that satisfies the second complexity criterion based on the corresponding region size.
[0051] In step 550, the electronic device determines the corresponding second quantization parameter for each of the real region and the additional real region that satisfy the second complexity criterion, based on the initial quantization parameter associated with the intermediate region, as referenced above. Figure 5B The process 500B is illustrated. The second quantization parameters of the real region and the additional real region that satisfy the second complexity criterion can be the same as or different from the second quantization parameters of the virtual region and the additional virtual region that do not satisfy the first complexity criterion. In step 560, the electronic device then encodes each of the real region and the additional real region that satisfy the second complexity criterion based on the corresponding second quantization parameters, as referenced above. Figure 5B The process shown in 500B is described.
[0052] Returning to step 540, in response to determining that at least one real region does not satisfy the second complexity criterion, the electronic device may optionally proceed to... Figure 7C The process 700C shows step 750. In step 750, the electronic device may optionally determine a corresponding region size for each real region that does not satisfy the second complexity criterion based on the initial region size associated with a simple real region. The initial region size associated with a simple real region may be larger than the initial region size associated with an intermediate region and the initial region size associated with a complex virtual region. The electronic device may optionally proceed to step 760 and divide the specific real region into one or more additional real regions based on the corresponding region size for each real region that does not satisfy the second complexity criterion.
[0053] In step 570, the electronic device determines a corresponding third quantization parameter for each of the real region and the additional real region that do not satisfy the second complexity criterion, based on the initial quantization parameters associated with the simple real region, as referenced above. Figure 5C The process described in step 500C is as follows. In step 580, the electronic device encodes each of the real region and the additional real region that do not satisfy the second complexity criterion based on the corresponding third quantization parameter, as referenced above. Figure 5CThe process 500C shown is described. Although process 700 illustrates three types of regions: complex virtual regions, intermediate regions, and simple real regions, any number of region types and corresponding complexity criteria, initial region sizes associated with the region type, initial quantization parameters associated with the region type, and upper and lower thresholds associated with the region type can be used instead.
[0054] Figure 8 An example diagram of the intermediate region 630 of an XR video frame 100 is shown, which is divided into multiple regions based on a first complexity criterion, a second complexity criterion, and an adjusted region size discussed herein with respect to process 700. The intermediate region 630 includes a portion of the background image 140 that satisfies the second complexity criterion and shows the legs of a table 130 resting on a portion of carpet 120. The intermediate region 630 is further divided into additional intermediate regions 810, 820, and 830. Additional intermediate region 810 includes the two legs of the table 130 resting on a portion of carpet 120. Additional intermediate region 820 includes a portion of carpet 120. Additional intermediate region 830 includes the legs of the table 130 resting on a portion of carpet 120. An initial region size associated with the intermediate region allows the electronic device to determine a smaller region size for the intermediate region 630 and to divide the intermediate region 630 into additional smaller intermediate regions 810, 820, and 830. Although Figure 8 The diagram shows an intermediate region 630 divided into three additional smaller intermediate regions 810, 820, and 830, but the intermediate region can be divided into any number of additional intermediate regions. Furthermore, the additional intermediate regions 810, 820, and 830 can have the same or different dimensions.
[0055] The corresponding second quantization parameters for intermediate regions 810 and 830 can be smaller than the corresponding second quantization parameter for intermediate region 820, to account for increased edge complexity, contrast, etc., of the legs of the table 130 resting on a portion of the carpet 120 in intermediate regions 810 and 830 compared to intermediate region 820, which only shows a portion of the carpet 120. In other words, during video encoding, intermediate regions 810 and 830 can be allocated more bits and have a higher image resolution than intermediate region 820. Figure 8 Example diagrams of additional intermediate regions 810, 820 and 830 of intermediate region 630 are shown, but the complex virtual region 610 and the simple real region 620 can be similarly divided into additional regions.
[0056] refer to Figure 9This document depicts a simplified block diagram of an electronic device 900, communicatively connected via a network 905 to an additional electronic device 980 and a network device 990, according to one or more embodiments of this disclosure. The electronic device 900 may be part of a multi-functional device, such as a mobile phone, tablet computer, personal digital assistant, portable music / video player, wearable device, head-mounted system, projection-based system, base station, laptop computer, desktop computer, network device, or any other electronic system as described herein. The electronic device 900, the additional electronic device 980, and / or the network device 990 may additionally or alternatively include one or more additional devices, which may contain various functions or may distribute various functions across these devices, such as server devices, base stations, auxiliary devices, etc. Exemplary networks such as network 905 include, but are not limited to, local area networks (such as Universal Serial Bus (USB) networks), organizational LANs, and wide area networks (such as the Internet). According to one or more embodiments, the electronic device 900 is used to enable a multi-view video codec. It should be understood that the various components and functions within the electronic device 900, the additional electronic device 980, and the network device 990 may be distributed differently on the device or on the additional device.
[0057] Electronic device 900 may include one or more processors 910, such as a central processing unit (CPU). Processor 910 may include systems-on-a-chip (SoCs) such as those present in mobile devices, and may include one or more dedicated graphics processing units (GPUs). Additionally, processor 910 may include multiple processors of the same or different types. Electronic device 900 may also include memory 930. Memory 930 may include one or more different types of memory that can be used in conjunction with processor 910 to perform device functions. For example, memory 930 may include cache, ROM, RAM, or any kind of transient or non-transitory computer-readable storage medium capable of storing computer-readable code. Memory 930 may store various programming modules executed by processor 910, including a video encoding module 935, a renderer 940, a gaze tracking module 945, and various other applications 950. Electronic device 900 may also include storage device 920. Storage device 920 may include one or more non-transitory computer-readable media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and magnetic tapes, optical media (such as CD-ROMs and digital video optical discs (DVDs)), and semiconductor storage devices (such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM)). According to one or more embodiments, storage device 920 may be configured to store virtual object data 925. The electronic device may additionally include a network interface 970 from which electronic device 900 can communicate via network 905.
[0058] Electronic device 900 may also include one or more cameras 960 or other sensors 965, such as depth sensors that can determine the depth of a scene. In one or more embodiments, each of the one or more cameras 960 may be a conventional RGB camera or a depth camera. Furthermore, the cameras 960 may include stereo cameras or other multi-camera systems, time-of-flight camera systems, etc. Electronic device 900 may also include a display 975. The display device 975 may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection technology, which projects graphic images onto the human retina. Projection systems may also be configured to project virtual objects onto a physical environment, such as as holograms or on a physical surface.
[0059] Storage device 920 can be used to store various data and structures that can be used to divide XR video frames into virtual and real regions and encode the virtual regions based on a first quantization parameter and the real regions based on a second quantization parameter. According to one or more embodiments, memory 930 may include one or more modules comprising computer-readable code executable by processor 910 to perform functions. Memory 930 may include, for example, a video encoding module 935 for encoding XR video frames, a renderer 940 for generating XR video frames, a gaze tracking module 945 for determining the user's gaze position and regions of interest in the image stream, and other application programs 950.
[0060] Although the electronic device 900 is depicted as including the numerous components described above, in one or more embodiments, the various components may be distributed across multiple devices. Therefore, although certain calls and transmissions are described herein with respect to the specific system depicted, in one or more embodiments, various calls and transmissions may be directed differently based on the functions of different distributions. Additionally, additional components may be used, and certain combinations of the functions of any components may be combined.
[0061] Now for reference Figure 10A simplified functional block diagram of an exemplary programmable electronic device 1000 for providing access to an app store, according to one embodiment, is shown. The electronic device 1000 may be, for example, a mobile phone, personal media device, portable camera or tablet computer, laptop or desktop computer system, network device, wearable device, etc. As shown, the electronic device 1000 may include a processor 1005, a display 1010, a user interface 1015, graphics hardware 1020, device sensors 1025 (e.g., proximity sensor / ambient light sensor, accelerometer and / or gyroscope), a microphone 1030, an audio codec 1035, a speaker 1040, communication circuitry 1045, image capture circuitry or unit 1050 (e.g., which may include multiple camera units / optical sensors with different characteristics (and a camera unit housed externally to the device 1000 but in electronic communication with the device)), a video codec 1055, a memory 1060, a storage device 1065, and a communication bus 1070.
[0062] Processor 1005 can execute instructions necessary for implementing or controlling various functions performed by device 1000 (e.g., the generation and / or processing of app store metrics according to the various embodiments described herein). Processor 1005 can, for example, drive display 1010 and can receive user input from user interface 1015. User interface 1015 can take various forms, such as buttons, keypad, dial pad, click wheel, keyboard, display screen, and / or touchscreen. User interface 1015 can, for example, be a conduit through which a user can view a captured video stream and / or indicate a specific image that the user wants to capture or share (e.g., by clicking a physical or virtual button at the moment the desired image is displayed on the device's display screen).
[0063] In one embodiment, display 1010 may display a video stream captured while processor 1005 and / or graphics hardware 1020 and / or image capture circuitry simultaneously store a video stream (or a single image frame from the video stream) in memory 1060 and / or storage device 1065. Processor 1005 may be a system-on-a-chip, such as those present in mobile devices, and may include one or more dedicated graphics processing units (GPUs). Processor 1005 may be based on a Reduced Instruction Set Computer (RISC) or Complex Instruction Set Computer (CISC) architecture or any other suitable architecture, and may include one or more processing cores. Graphics hardware 1020 may be dedicated computing hardware for processing graphics and / or assisting processor 1005 in performing computational tasks. In one embodiment, graphics hardware 1020 may include one or more programmable graphics processing units (GPUs).
[0064] For example, according to this disclosure, image capture circuitry 1050 may include one or more camera units configured to capture images. The output from image capture circuitry 1050 may be processed at least in part by video codec 1055 and / or processor 1005 and / or graphics hardware 1020 and / or a dedicated image processing unit incorporated within circuitry 1050. The captured images may thus be stored in memory 1060 and / or storage device 1065. Memory 1060 may include one or more different types of media used by processor 1005, graphics hardware 1020, and image capture circuitry 1050 to perform device functions. For example, memory 1060 may include memory cache, read-only memory (ROM), and / or random access memory (RAM). Storage device 1065 may store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device configuration file information, and any other suitable data. Storage device 1065 may include one or more non-transitory storage media, including, for example, magnetic disks (fixed hard disks, floppy disks, and removable disks) and magnetic tapes, optical media such as CD-ROMs and digital video discs (DVDs), and semiconductor memory devices such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM). Memory 1060 and storage device 1065 can be used to hold computer program instructions or code organized into one or more modules and written in any desired computer programming language. When executed, for example, by processor 1005, such computer program code can implement one or more of the methods described herein. Power supply 1075 may include a rechargeable battery (e.g., a lithium-ion battery, etc.) for managing electronic components and associated circuitry of electronic device 1000 and / or providing power to said electronic components and associated circuitry, or other electrical connections to a power source (e.g., mains power).
[0065] It should be understood that the above description is intended to be exemplary and not restrictive. The material has been presented to enable any person skilled in the art to make and use the disclosed matters protected by the claims, and is provided in the context of a particular embodiment, variations of which will be readily apparent to a person skilled in the art (e.g., some embodiments of the disclosed embodiments may be used in combination with each other). Therefore, Figure 2 , Figure 4 , Figures 5A to 5C as well as Figures 7A to 7C The specific arrangement of the steps or actions shown or Figure 9 and Figure 10The arrangement of the elements shown should not be construed as limiting the scope of the disclosed subject matter. Therefore, the scope of the invention should be determined by reference to the appended claims and the full scope of their equivalents. In the appended claims, the terms “comprising” and “therein” are used as common English equivalents of the corresponding terms “including” and “wherein”.
Claims
1. A method for encoding extended reality (XR) video frames, comprising: Obtain an XR video frame, the XR video frame including a background image and a virtual object covering at least a portion of the background image; The XR video frame is divided into a first region including virtual content, a second region including real content within the physical environment, and a third region including both virtual content and real content. The first region includes at least a portion of the virtual object, the second region includes a region of the background image separated from the first region, and the third region includes a portion of the virtual object and a portion of the background image. The first quantization parameter is determined for the first region based on the initial quantization parameter associated with the virtual region; The second quantization parameters are determined for the second region based on the initial quantization parameters associated with the real region. The third quantization parameters are determined for the third region based on the initial quantization parameters associated with the intermediate region. as well as The first region is encoded based on the corresponding first quantization parameter, the second region is encoded based on the corresponding second quantization parameter, and the third region is encoded based on the corresponding third quantization parameter.
2. The method of claim 1, wherein determining the corresponding first quantization parameter for the first region is further based on an upper threshold and a lower threshold associated with the virtual region.
3. The method of claim 1, wherein determining the corresponding second quantization parameter for the second region is further based on an upper threshold and a lower threshold associated with the real region.
4. The method according to claim 1, wherein: The first region includes at least a portion of the virtual object that satisfies the first complexity criterion; The second region includes a first portion of the region of the background image that is separated from the first region, wherein the first portion of the region of the background image that is separated from the first region does not satisfy the second complexity criterion; The third region includes at least one of the following: (i) a portion of the at least one virtual object that does not satisfy the first complexity criterion, and (ii) a second portion of the region of the background image separated from the first region, wherein the second portion of the region of the background image separated from the first region satisfies the second complexity criterion.
5. The method of claim 1, wherein determining the corresponding third quantization parameter for the third region is further based on an upper threshold and a lower threshold associated with the intermediate region.
6. The method of claim 1, further comprising obtaining input indicating a focus area via a gaze-tracking user interface, wherein the division of the XR video frame is at least partially based on the focus area.
7. A non-transitory computer-readable medium comprising computer code, said computer code being executable by at least one processor to: Obtain an extended reality (XR) video frame, the XR video frame including a background image and a virtual object covering at least a portion of the background image; The XR video frame is divided into a first region including virtual content, a second region including real content within the physical environment, and a third region including both virtual content and real content. The first region includes at least a portion of the virtual object, the second region includes a region of the background image separated from the first region, and the third region includes a portion of the virtual object and a portion of the background image. The first quantization parameter is determined for the first region based on the initial quantization parameter associated with the virtual region; The second quantization parameters are determined for the second region based on the initial quantization parameters associated with the real region. The third quantization parameters are determined for the third region based on the initial quantization parameters associated with the intermediate region. as well as The first region is encoded based on the corresponding first quantization parameter, the second region is encoded based on the corresponding second quantization parameter, and the third region is encoded based on the corresponding third quantization parameter.
8. The non-transitory computer-readable medium of claim 7, wherein the computer-readable code for determining the corresponding first quantization parameter for the first region further includes computer-readable code for further determining the corresponding first quantization parameter based on an upper threshold and a lower threshold associated with the virtual region.
9. The non-transitory computer-readable medium of claim 7, wherein the computer-readable code for determining the corresponding second quantization parameter for the second region further includes computer-readable code for further determining the corresponding second quantization parameter based on an upper threshold and a lower threshold associated with the real region.
10. The non-transitory computer-readable medium of claim 7, wherein the initial quantization parameter associated with the virtual region is smaller than the initial quantization parameter associated with the real region.
11. The non-transitory computer-readable medium of claim 7, wherein the computer-readable medium further comprises computer-readable code executable by the at least one processor to perform the following operations: For the first region: The size of the first region is determined based on the initial region size associated with the virtual region; and Based on the corresponding first region size, the first region is divided into one or more additional virtual regions; For the second region: The size of the second region is determined based on the initial region size associated with the real region; as well as Based on the corresponding second region size, the second region is divided into one or more additional real regions.
12. The non-transitory computer-readable medium of claim 11, wherein the initial region size associated with the virtual region is smaller than the initial region size associated with the real region.
13. The non-transitory computer-readable medium according to claim 7, wherein: The first region includes at least a portion of the virtual object that satisfies the first complexity criterion; The second region includes a first portion of the region of the background image that is separated from the first region, wherein the first portion of the region of the background image that is separated from the first region does not satisfy the second complexity criterion; and The third region includes at least one of the following: (i) a portion of the virtual object that does not satisfy the first complexity criterion, and (ii) a second portion of the region of the background image that is separated from the first region, wherein the second portion of the region satisfies the second complexity criterion.
14. The non-transitory computer-readable medium of claim 7, wherein the computer-readable code for determining the corresponding third quantization parameter for the third region further comprises computer-readable code for further determining the corresponding third quantization parameter based on an upper threshold associated with an intermediate region and a lower threshold associated with an intermediate region.
15. The non-transitory computer-readable medium of claim 13, wherein the initial quantization parameter associated with the intermediate region is less than the initial quantization parameter associated with the real region and greater than the initial quantization parameter associated with the virtual region.
16. The non-transitory computer-readable medium of claim 13, wherein the computer-readable medium further comprises computer-readable code executable by the at least one processor to perform the following operations: For the first region: The size of the first region is determined based on the initial region size associated with the virtual region; and Based on the corresponding first region size, the first region is divided into one or more additional virtual regions; For the second region: The size of the second region is determined based on the initial region size associated with the real region; and Based on the corresponding second region size, the second region is divided into one or more additional real regions; as well as For the third region: The size of the corresponding third region is determined based on the initial region size associated with the intermediate region; and The third region is divided into one or more additional intermediate regions based on the corresponding third region size.
17. The non-transitory computer-readable medium of claim 16, wherein the initial region size associated with the intermediate region is greater than the initial region size associated with the virtual region and smaller than the initial region size associated with the real region.
18. The non-transitory computer-readable medium of claim 7, wherein the computer-readable medium further comprises computer-readable code executable by the at least one processor to obtain input indicating a focus region via a gaze-tracking user interface, wherein the computer-readable code for segmenting the XR video frame further comprises computer-readable code for segmenting the XR video frame at least in part based on the focus region.
19. An apparatus for video encoding, comprising: An image capturing device configured to capture a background image; At least one processor; as well as At least one computer-readable medium, the at least one computer-readable medium comprising computer-readable code, the computer-readable code being executable by the at least one processor to: Obtain an extended reality (XR) video frame, the XR video frame including the background image and a virtual object covering at least a portion of the background image; The XR video frame is divided into a first region including virtual content, a second region including real content within the physical environment, and a third region including both virtual content and real content. The first region includes at least a portion of the virtual object, the second region includes a region of the background image separated from the first region, and the third region includes a portion of the virtual object and a portion of the background image. The first quantization parameter is determined for the first region based on the initial quantization parameter associated with the virtual region; The second quantization parameters are determined for the second region based on the initial quantization parameters associated with the real region. The third quantization parameters are determined for the third region based on the initial quantization parameters associated with the intermediate region. as well as The first region is encoded based on the corresponding first quantization parameter, the second region is encoded based on the corresponding second quantization parameter, and the third region is encoded based on the corresponding third quantization parameter.
20. The device of claim 19, wherein the computer-readable code for determining the corresponding first quantization parameter for the first region further includes computer-readable code for further determining the corresponding first quantization parameter based on a threshold upper limit associated with the virtual region and a threshold lower limit associated with the virtual region.