An adaptive quantization matrix for extended reality video encoding

IN598799BActive Publication Date: 2026-08-12APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
IN202214047952
Authority / Receiving Office
IN · IN
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-27
Filing Date
2022-08-23
Publication Date
2026-08-12
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

Existing video encoding systems face challenges in efficiently allocating bits to regions of interest in video frames, often requiring computationally expensive and time-consuming image analysis to ensure proper quality, especially in extended reality (XR) environments where viewer focus is on virtual objects.

Method used

The system divides an XR video frame into virtual and real regions based on the position of virtual objects and background images, using adaptive quantization matrices to allocate more bits to virtual regions of interest, allowing for improved encoding efficiency without extensive computational analysis.

Benefits of technology

This approach enhances bit allocation efficiency, prioritizing regions of interest while maintaining image quality, thereby reducing computational load and improving encoding speed in XR environments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Encoding an extended-reality (XR) video frame may include obtaining an XR video 5 frame comprising a background image and a virtual object; obtaining, from an image renderer, a first region of the background image over which the virtual object is overlaid; dividing the XR video frame into a virtual region and a real region, wherein the virtual region comprises the first region of the background image and the virtual object and the real region comprises a second region of the background 10 image; determining, for the virtual region, a corresponding first quantization parameter based on an initial quantization parameter associated with virtual regions; determining, for the real region, a corresponding second quantization parameter based on an initial quantization parameter associated with real regions; and encoding the virtual region based on the corresponding first quantization 15 parameter and the real region based on the corresponding second quantization parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Background

[0001] This disclosure relates generally to image processing. Moreparticularly, but not by way of limitation, this disclosure relates to techniques andsystems of video encoding.

[0002] Some video encoding systems use bit-rate control algorithms todetermine how many bits to allocate to a particular region of a video frame to ensurea uniform picture quality for a given video-encoding standard and reduce thebandwidth needed to transmit the encoded video frame. Some bit-rate controlalgorithms use frame-level and macroblock-level content statistics such ascomplexity and contrast to determine quantization parameters and correspondingbit allocations. A quantization parameter is an integer mapped to a quantization stepsize and controls an amount of compression for each region of a video frame. Forexample, an eight by eight region of pixels is multiplied by the quantizationparameter and divided by a quantization matrix. The resulting values are thenrounded to the nearest integer. A large quantization parameter corresponds to higherquantization, more compression, and lower image quality than a small quantizationparameter that corresponds to lower quantization, less compression, and higherimage quality. Bit-rate control algorithms may use a constant quantizationparameter or varying quantization parameters to accommodate a target averagebitrate, a constant bitrate, a constant image quality, or the like. However, many bitrate control algorithms are objective and cannot guarantee that more bits areallocated to a region of interest than to the background. Some bit-rate controlalgorithms are able determine a region of interest and allocate more bits to theregion of interest than to the background, but they are often computationally-expensive and time-consuming to operate. What is needed is an improved techniqueto encode video frames.Brief Description of the Drawings

[0003] FIG. 1 shows an example diagram of an extended reality (XR) videoframe.

[0004] FIG. 2 shows, in flow chart form, an example process for encodingan extended reality video frame based on an adaptive quantization matrix.

[0005] FIG. 3 shows an example diagram of an extended reality video framedivided into a virtual region and a real region.

[0006] FIG. 4 shows, in flowchart form, an example process for encodingan extended reality video frame based on an adaptive quantization matrix and inputfrom a gaze-tracking user interface.

[0007] FIGs. 5A-C show, in flowchart form, an example process forencoding an extended reality video frame based on an adaptive quantization matrixand first and second complexity criteria.

[0008] FIG. 6 shows an example diagram of an extended reality video frame divided into regions based on first and second complexity criteria.

[0009] FIGs. 7A-C show, in flowchart form, an example process forencoding an extended reality video frame based on an adaptive quantization matrix,first and second complexity criteria, and adjusted region sizes.

[0010] FIG. 8 shows an example diagram of a medial region of an extendedreality video frame divided into regions based on first and second complexitycriteria and adjusted region sizes.

[0011] FIG. 9 shows, in block diagram form, exemplary systems forencoding extended reality video streams.

[0012] FIG. 10 shows an exemplary system for use in various videoencoding systems, including for encoding extended reality video streams.Detailed Description

[0013] This disclosure pertains to systems, methods, and computer readablemedia for a video-encoding extended reality (XR) video streams. In particular, anXR video frame comprising a background image and at least one virtual object maybe obtained. A first region of the background image over which the at least onevirtual object is to be overlaid may be obtained from an image renderer. The XRvideo frame may be divided into at least one virtual region and at least one realregion. The at least one virtual region comprises the first region of the backgroundimage and the at least one virtual object. The at least one real region comprises asecond region of the background image. For each of the at least one virtual regions,a corresponding first quantization parameter may be determined based on an initialquantization parameter associated with virtual regions. For each of the at least onereal regions, a corresponding second quantization parameter may be determinedbased on an initial quantization parameter associated with real regions. Each of theat least one virtual regions may be encoded based on the corresponding firstquantization parameter, and each of the at least one real regions may be encodedbased on the corresponding second quantization parameter.

[0014] Various examples of electronic systems and techniques for usingsuch systems in relation to encoding extended reality video streams are described.

[0015] A physical environment refers to a physical world that people cansense and / or interact with without aid of electronic systems. Physicalenvironments, such as a physical park, include physical articles, such as physicaltrees, physical buildings, and physical people. People can directly sense and / orinteract with the physical environment, such as through sight, touch, hearing, taste,and smell.

[0016] In contrast, an extended reality (XR) environment refers to a whollyor partially simulated environment that people sense and / or interact with via anelectronic system. In XR, a subset of a person's physical motions, orrepresentations thereof, are tracked, and, in response, one or more characteristics ofone or more virtual objects simulated in the XR environment are adjusted in amanner that comports with at least one law of physics. For example, a XR systemmay detect a person's head turning and, in response, adjust graphical content andan acoustic field presented to the person in a manner similar to how such views andsounds would change in a physical environment. In some situations (e.g., foraccessibility reasons), adjustments to characteristic(s) of virtual object(s) in a XRenvironment may be made in response to representations of physical motions (e.g.,vocal commands).

[0017] A person may sense and / or interact with a XR object using any oneof their senses, including sight, sound, touch, taste, and smell. For example, aperson may sense and / or interact with audio objects that create 3D or spatial audioenvironment that provides the perception of point audio sources in 3D space. Inanother example, audio objects may enable audio transparency, which selectivelyincorporates ambient sounds from the physical environment with or withoutcomputer-generated audio. In some XR environments, a person may sense and / orinteract only with audio objects.

[0018] A virtual reality (VR) environment refers to a simulatedenvironment that is designed to be based entirely on computer-generated sensoryinputs for one or more senses. A VR environment comprises a plurality of virtualobjects with which a person may sense and / or interact. For example, computer-generated imagery of trees, buildings, and avatars representing people are examplesof virtual objects. A person may sense and / or interact with virtual objects in theVR environment through a simulation of the person's presence within thecomputer-generated environment, and / or through a simulation of a subset of theperson's physical movements within the computer-generated environment.

[0019] In contrast to a VR environment, which is designed to be basedentirely on computer-generated sensory inputs, a mixed reality (MR) environmentrefers to a simulated environment that is designed to incorporate sensory inputsfrom the physical environment, or a representation thereof, in addition to includingcomputer-generated sensory inputs (e.g., virtual objects). On a virtualitycontinuum, a mixed reality environment is anywhere between, but not including, awholly physical environment at one end and virtual reality environment at the otherend.

[0020] In some MR environments, computer-generated sensory inputs mayrespond to changes in sensory inputs from the physical environment. Also, someelectronic systems for presenting an MR environment may track location and / ororientation with respect to the physical environment to enable virtual objects tointeract with real objects (that is, physical articles from the physical environment orrepresentations thereof). For example, a system may account for movements sothat a virtual tree appears stationery with respect to the physical ground.

[0021] An augmented reality (AR) environment refers to a simulatedenvironment in which one or more virtual objects are superimposed over a physicalenvironment, or a representation thereof. For example, an electronic system forpresenting an AR environment may have a transparent or translucent displaythrough which a person may directly view the physical environment. The systemmay be configured to present virtual objects on the transparent or translucentdisplay, so that a person, using the system, perceives the virtual objectssuperimposed over the physical environment. Alternatively, a system may have anopaque display and one or more imaging sensors that capture images or video ofthe physical environment, which are representations of the physical environment.The system composites the images or video with virtual objects, and presents thecomposition on the opaque display. A person, using the system, indirectly viewsthe physical environment by way of the images or video of the physicalenvironment, and perceives the virtual objects superimposed over the physicalenvironment. As used herein, a video of the physical environment shown on anopaque display is called "pass-through video," meaning a system uses one or moreimage sensor(s) to capture images of the physical environment, and uses thoseimages in presenting the AR environment on the opaque display. Furtheralternatively, a system may have a projection system that projects virtual objectsinto the physical environment, for example, as a hologram or on a physical surface,so that a person, using the system, perceives the virtual objects superimposed overthe physical environment.

[0022] An augmented reality environment also refers to a simulatedenvironment in which a representation of a physical environment is transformed bycomputer-generated sensory information. For example, in providing pass-throughvideo, a system may transform one or more sensor images to impose a selectperspective (e.g., viewpoint) different than the perspective captured by the imagingsensors. As another example, a representation of a physical environment may betransformed by graphically modifying (e.g., enlarging) portions thereof, such thatthe modified portion may be representative but not photorealistic versions of theoriginally captured images. As a further example, a representation of a physicalenvironment may be transformed by graphically eliminating or obfuscating portionsthereof.

[0023] An augmented virtuality (AV) environment refers to a simulatedenvironment in which a virtual or computer generated environment incorporatesone or more sensory inputs from the physical environment. The sensory inputs maybe representations of one or more characteristics of the physical environment. Forexample, an AV park may have virtual trees and virtual buildings, but people withfaces photorealistically reproduced from images taken of physical people. Asanother example, a virtual object may adopt a shape or color of a physical articleimaged by one or more imaging sensors. As a further example, a virtual object mayadopt shadows consistent with the position of the sun in the physical environment.

[0024] There are many different types of electronic systems that enable aperson to sense and / or interact with various XR environments. Examples includehead mounted systems, projection-based systems, heads-up displays (HUDs),vehicle windshields having integrated display capability, windows havingintegrated display capability, displays formed as lenses designed to be placed on aperson's eyes (e.g., similar to contact lenses), headphones / earphones, speakerarrays, input systems (e.g., wearable or handheld controllers with or without hapticfeedback), smartphones, tablets, and desktop / laptop computers. A head mountedsystem may have one or more speaker(s) and an integrated opaque display.Alternatively, a head mounted system may be configured to accept an externalopaque display (e.g., a smartphone). The head mounted system may incorporateone or more imaging sensors to capture images or video of the physicalenvironment, and / or one or more microphones to capture audio of the physicalenvironment. Rather than an opaque display, a head mounted system may have atransparent or translucent display. The transparent or translucent display may havea medium through which light representative of images is directed to a person'seyes. The display may utilize digital light projection, OLEDs, LEDs, uLEDs, liquidcrystal on silicon, laser scanning light source, or any combination of thesetechnologies. The medium may be an optical waveguide, a hologram medium, anoptical combiner, an optical reflector, or any combination thereof. In oneembodiment, the transparent or translucent display may be configured to becomeopaque selectively. Projection-based systems may employ retinal projectiontechnology that projects graphical images onto a person's retina. Projectionsystems also may be configured to project virtual objects into the physicalenvironment, for example, as a hologram or on a physical surface.

[0025] In the following description, for purposes of explanation, numerousspecific details are set forth in order to provide a thorough understanding of thedisclosed concepts. As part of this description, some of this disclosure's drawingsrepresent structures and devices in block diagram form in order to avoid obscuringthe novel aspects of the disclosed concepts. In the interest of clarity, not all featuresof an actual implementation may be described. Further, as part of this description,some of this disclosure's drawings may be provided in the form of flowcharts. Theboxes in any particular flowchart may be presented in a particular order. It shouldbe understood however that the particular sequence of any given flowchart is usedonly to exemplify one embodiment. In other embodiments, any of the variouselements depicted in the flowchart may be deleted, or the illustrated sequence ofoperations may be performed in a different order, or even concurrently. In addition,other embodiments may include additional steps not depicted as part of theflowchart. Moreover, the language used in this disclosure has been principallyselected for readability and instructional purposes, and may not have been selectedto delineate or circumscribe the inventive subject matter, resort to the claims beingnecessary to determine such inventive subject matter. Reference in this disclosureto "one embodiment" or to "an embodiment" means that a particular feature,structure, or characteristic described in connection with the embodiment is includedin at least one embodiment of the disclosed subject matter, and multiple referencesto "one embodiment" or "an embodiment" should not be understood as necessarilyall referring to the same embodiment.

[0026] It will be appreciated that in the development of any actualimplementation (as in any software and / or hardware development project),numerous decisions must be made to achieve a developers' specific goals (e.g.,compliance with system- and business-related constraints), and that these goals mayvary from one implementation to another. It will also be appreciated that suchdevelopment efforts might be complex and time-consuming, but would nevertheless be a routine undertaking for those of ordinary skill in the design andimplementation of video encoding systems having the benefit of this disclosure.

[0027] FIG. 1 shows an example diagram of an XR video frame 100. TheXR video frame 100 includes a background image 140 showing real objects, suchas the dresser 110, the rug 120, and the table 130, and a virtual object 150 that isoverlaid with the background image 140 such that the virtual object 150 appearsatop the table 130. The background image 140 is described as a "backgroundimage" to indicate the image is behind the virtual object 150 and may have aforeground region and a background region. With XR video, viewers often focuson virtual objects and the areas immediately surrounding the virtual objects, ratherthan the background environment. For example, a viewer looking at the XR videoframe 100 may focus on the virtual object 150 and the portion of the table 130 andrug 120 immediately surrounding the virtual object 150, rather than the dresser 110.Instead of performing computationally expensive and time consuming imageanalysis of each frame in an XR video to determine a region of interest based onthe image content of each frame, a video-encoding system may use the virtual object150 and the known region of the background image 140 over which the virtualobject 150 is placed to determine a region of interest for the viewer. Based on thevirtual object 150 and its position over the background image 140, the video-encoding system may allocate more bits to the region of interest for the viewer thanto the remainder of background image 140.

[0028] FIG. 2 shows, in flow chart form, an example process 200 forencoding an XR video frame 100 based on an adaptive quantization matrix. Forpurposes of explanation, the following steps are described as being performed byparticular components, However, it should be understood that the various actionsmay be performed by alternate components. In addition, the various actions may beperformed in a different order. Further, some actions may be performedsimultaneously, and some may not be required, or others may be added. For ease ofexplanation, the process 200 is described with reference to the XR video frame 100shown in FIG. 1.

[0029] The flowchart begins at step 210, where an electronic device obtainsan XR video frame 100 comprising a background image 140 and at least one virtualobject 150. At step 220, the electronic device obtains, from an image renderer, afirst region of the background image 140 over which the virtual object 150 isoverlaid. For example, the first region of the background image 140 may indicatethe portion of the rug 120 and table 130 over which the virtual object 150 ispositioned. The electronic device divides the XR video frame 100 into at least onevirtual region and at least one real region based on the first region of the backgroundimage 140 at step 230. The virtual region includes at least a portion of the virtualobject. The virtual region may further include the entire virtual object, and includenone of the background image or a portion of the background image. For example,a virtual region may include the virtual object 150 and a portion of the rug 120 andtable 130, and a real region may include the remainder of the background image140, such as the dresser 110 and the other portions of the rug 120 and the table 130.

[0030] At step 240, the electronic device determines, for each of the at leastone virtual regions, a corresponding first quantization parameter based on an initialquantization parameter associated with virtual regions. For example, the electronicdevice may determine an image complexity of a particular virtual region is greaterthan an image complexity of a reference virtual region associated with the initialquantization parameter for virtual regions and decrease the initial quantizationparameter by a proportional amount. At step 250, the electronic device determines,for each of the at least one real regions, a corresponding second quantizationparameter based on an initial quantization parameter associated with real regions.For example, the electronic device may determine an image complexity of aparticular real region is less than an image complexity of a reference real regionassociated with the initial quantization parameter for real regions and increase theinitial quantization parameter by a proportional amount. The initial quantizationparameter associated with virtual regions may be smaller than the initialquantization parameter associated with real regions to indicate a larger amount ofdetail and complexity in the virtual regions than in the real regions. That is, theinitial quantization parameters associated with the virtual and real regions may bechosen such that the virtual regions corresponding to the viewer's region of interestare allocated more bits than real regions outside the region of interest during videoencoding of the XR video frame 100. At step 260, the electronic device encodes theat least one virtual region based on the first quantization parameter and the at leastone real region based on the second quantization parameter. The resulting encodedXR video frame allocates more bits to the at least one virtual region based on thefirst quantization parameter than to the at least one real region based on the secondquantization parameter.

[0031] FIG. 3 shows an example diagram of the XR video frame 100 shownin FIG. 1 divided into a virtual region 310 and a real region 320. In step 230 ofprocess 200, the electronic device divides the XR video frame 100 into a virtualregion 310 and a real region 320. The virtual region 310 includes the virtual object150 and a portion of the background image 140 around the virtual object 150,showing the surface of the table 130 and a portion of the rug 120. In this example,the virtual region 310 includes the entire virtual object 150 and a portion of thebackground image 140, but in other implementations, the virtual region 310 mayinclude the entire virtual object 150 but omit the portion of the background image140, or include a portion of the virtual object 150 and a portion of the backgroundimage 140, or include a portion of the virtual object but omit the portion of thebackground image 140. The negative space in the real region 320 indicates wherethe virtual region 310 is located. The virtual region 310 and the real region 320 maybe divided into one or more additional, smaller regions to allow further refinementof the quantization parameters based on the complexity, contrast, etc. in differentportions of the regions 310 and 320.

[0032] FIG. 4 shows, in flowchart form, an example process 400 forencoding an XR video frame based on an adaptive quantization matrix and inputfrom a gaze-tracking user interface. For purposes of explanation, the followingsteps are described as being performed by particular components, However, itshould be understood that the various actions may be performed by alternatecomponents. In addition, the various actions may be performed in a different order.Further, some actions may be performed simultaneously, and some may not berequired, or others may be added. For ease of explanation, the process 400 isdescribed with reference to the process 200 described herein with reference to FIG.2.

[0033] The flowchart 400 begins with steps 210 and 220, as describedabove with reference to FIG. 2. Dividing the XR video frame into at least one virtualregion and at least one real region in step 230 may optionally include steps 410 and420. At step 410, the electronic device obtains input indicative of an area of focus,for example via a gaze-tracking user interface, a cursor-based user interface, andthe like. For example, where the XR video frame includes a plurality of virtual objects, the input indicative of an area of focus via a gaze-tracking user interfacemay indicate which particular virtual object the user is looking at out of the pluralityof virtual objects.

[0034] At step 420, the electronic device divides the XR video frame intothe at least one virtual region and the at least one real region based on the area offocus. The electronic device may divide the particular virtual object and thecorresponding portion of the background image over which the particular virtualobject is overlaid into a unique virtual region and the remaining virtual objects outof the plurality of virtual objects into one or more additional virtual regions.Similarly, the electronic device may divide the remaining portions of thebackground image not included in the real regions into one or more additional,smaller regions to further refine the quantization parameters based on thecomplexity, contrast, etc. in different regions of the remaining portion of thebackground image.

[0035] Determining, for each of the virtual regions, a corresponding firstquantization parameter based on an initial quantization parameter associated withvirtual regions at step 240 may optionally include step 430. At step 430, theelectronic device determines a corresponding first quantization parameter based onthe area of focus indicated by the input from the gaze-tracking user interface. Forexample, the first quantization parameter for the virtual region that includes the areaof focus may be smaller than the first quantization parameter for other virtualregions. That is, the virtual region that includes the area of focus may be allocatedmore bits and encoded with a higher resolution than the other virtual regions. Theelectronic device proceeds to steps 250 and 260, as described above with referenceto FIG. 2 and based on the regions of the XR video frame as divided in step 420and the corresponding first quantization parameters determined at step 430.

[0036] FIGs. 5A-C show, in flowchart form, an example process 500 forencoding an XR video frame based on an adaptive quantization matrix and first andsecond complexity criteria. For purposes of explanation, the following steps aredescribed as being performed by particular components, However, it should beunderstood that the various actions may be performed by alternate components. Inaddition, the various actions may be performed in a different order. Further, someactions may be performed simultaneously, and some may not be required, or othersmay be added. For ease of explanation, the process 500 is described with referenceto the process 200 described herein with reference to FIG. 2 and the XR video frame100 described herein with reference to FIG. 1.

[0037] The flowchart 500 begins in FIG. 5A with steps 210, 220, and 230as described above with reference to FIG. 2. After dividing the XR video frame intoat least one virtual region and at least one real region, the electronic device proceedsto step 510 and determines whether at least one virtual region satisfies a firstcomplexity criterion. The first complexity criterion may be representative of athreshold amount of image complexity, contrast, and the like, such that virtualregions that satisfy the first complexity criterion are more complex than virtualregions that do not satisfy the first complexity criterion and are considered complexvirtual regions. For example, a complex virtual region that satisfies the firstcomplexity criterion may include a highly-detailed virtual object, such as a useravatar's face, while a virtual region that does not satisfy the first complexitycriterion includes a comparatively simple virtual object, such as a ball. In responseto determining at least one of the virtual regions satisfies the first complexitycriterion, the electronic device proceeds to step 520 and determines, for each of thevirtual regions that satisfy the first complexity criterion (that is, the complex virtualregions), a corresponding first quantization parameter based on an initialquantization parameter associated with complex virtual regions.

[0038] The corresponding first quantization parameter may further bedetermined based on a threshold upper limit and a threshold lower limit associatedwith complex virtual regions. In response to the first quantization parameterreaching the threshold upper or lower limit associated with complex virtual regions,the electronic device stops determining the corresponding first quantizationparameter. The threshold upper and lower limits associated with complex virtualregions may be chosen based on the complexity of the virtual object 150 and thebackground image 140, the image quality requirements associated with a givenvideo-encoding standard, the time allotted to the video-encoding process, and thelike. For example, a particular video-encoding standard may set a range of validvalues for the quantization parameter, and the threshold upper and lower limits maydefine the boundaries of the range of valid values according to the particular video encoding standard. As another example, the first quantization parameter may bedetermined in an iterative process, and the threshold upper and lower limits mayrepresent a maximum and a minimum number of iterations, respectively, that maybe performed in the time allotted to the video encoding process. As a furtherexample, the threshold upper and lower limits may represent image quality criterionassociated with complex virtual regions. That is, the threshold upper limit mayrepresent a maximum image quality for complex virtual regions at a particular bitrate, such that the bit rate is not slowed by the additional detail included in thecomplex virtual regions, and the threshold lower limit may represent a minimumimage quality for complex virtual regions at the particular bit rate, such that aminimum image quality for complex virtual regions is maintained at the particularbit rate. At step 530, the electronic device encodes each of the virtual regions thatsatisfy the first complexity criterion based on the corresponding first quantizationparameter.

[0039] Returning to step 510, in response to determining at least one virtualregion does not satisfy the first complexity criterion, the electronic device proceedsto step 550 shown in process 500B of FIG. 5B. At step 550, the electronic devicedetermines, for each of the virtual regions that do not satisfy the first complexitycriterion (that is, the simple virtual regions), a corresponding second quantizationparameter based on an initial quantization parameter associated with medialregions. Medial regions may include comparatively simple virtual regions that do not satisfy the first complexity criterion and comparatively complex real regionsthat satisfy the second complexity criterion. The initial quantization parameterassociated with medial regions may be greater than the initial quantizationparameter associated with complex virtual regions, such that medial regions areencoded using fewer bits and in a lower resolution than the number of bits andresolution with which complex virtual regions are encoded.

[0040] The corresponding second quantization parameter may further bedetermined based on a threshold upper limit and a threshold lower limit associatedwith medial regions. In response to the second quantization parameter reaching thethreshold upper or lower limit associated with medial regions, the electronic devicestops determining the corresponding second quantization parameter. The thresholdupper and lower limits associated with medial regions may be chosen based on thecomplexity of the virtual object 150 and the background image 140, the imagequality requirements associated with a given video-encoding standard, the timeallotted to the video-encoding process, and the like. For example, a particular video encoding standard may set a range of valid values for the quantization parameter,and the threshold upper and lower limits may define the boundaries of the range ofvalid values according to the particular video-encoding standard. As anotherexample, the second quantization parameter may be determined in an iterativeprocess, and the threshold upper and lower limits may represent a maximum and aminimum number of iterations, respectively, that may be performed in the timeallotted to the video encoding process. As a further example, the threshold upperand lower limits may represent image quality criterion associated with medialregions. That is, the threshold upper limit may represent a maximum image qualityfor medial regions at a particular bit rate, such that the bit rate is not slowed by theadditional detail included in the medial regions, and the threshold lower limit mayrepresent a minimum image quality for medial regions at the particular bit rate, suchthat a minimum image quality for medial regions is maintained at the particular bitrate. In some implementations, the maximum and minimum image qualities formedial regions at a particular bit rate may be lower than the maximum andminimum image qualities for complex virtual regions at the particular bitrate, toensure that more bits are allocated to the complex virtual regions than to the medialregions. The electronic device encodes each of the virtual regions that do not satisfythe first complexity criterion based on the corresponding second quantizationparameter at step 560.

[0041] Returning to the at least one real region from step 230, the electronicdevice determines whether the at least one real region satisfies a second complexitycriterion at step 540. The second complexity criterion may be representative of athreshold amount of image complexity, contrast, and the like, such that real regionsthat satisfy the second complexity criterion are more complex than real regions thatdo not satisfy the second complexity criterion and are considered complex realregions or medial regions. A complex real region that satisfies the secondcomplexity criterion may include a highly-detailed portion of the background image140 such as the portion of the background image 140 showing the legs of table 130against the portion of the rug 120, which includes multiple edges and contrasts intexture and color between the table 130 and the rug 120. A real region that does notsatisfy the second complexity criterion may include a comparatively simple portionof the background image 140, such as the dresser 110 and uniform portions of thewalls and rug 120. In response to the at least one real region satisfying the secondcomplexity criterion, the electronic device proceeds to step 550 shown in process500B of FIG. 5B and described above. At step 550, the electronic devicedetermines, for each of the real regions that satisfy the second complexity criterion(that is, the complex real regions), a corresponding second quantization parameterbased on an initial quantization parameter associated with medial regions. Thesecond quantization parameter for the at least one real region satisfying the secondcomplexity criterion may be the same or different than the second quantizationparameter for the at least one virtual region not satisfying the first complexitycriterion. The electronic device then encodes each of the real regions that satisfy thesecond complexity criterion based on the corresponding second quantizationparameter at step 560.

[0042] Returning to step 540, in response to determining the at least onereal region does not satisfy the second complexity criterion, the electronic deviceproceeds to step 570 shown in process 500C of FIG. 5C. At step 570, the electronicdevice determines, for each of the real regions that do not satisfy the secondcomplexity criterion (that is, the simple real regions), a corresponding thirdquantization parameter based on an initial quantization parameter associated withsimple real regions. The initial quantization parameter associated with simple realregions may be greater than the initial quantization parameter associated withmedial regions and the initial quantization parameter associated with complexvirtual regions, such that simple real regions are encoded using fewer bits and in alower resolution than the number of bits and resolution with which medial regionsand complex virtual regions are encoded.

[0043] The corresponding third quantization parameter may further bedetermined based on a threshold upper limit and a threshold lower limit associatedwith simple real regions. In response to the third quantization parameter reachingthe threshold upper or lower limit associated with simple real regions, the electronicdevice stops determining the corresponding third quantization parameter. Thethreshold upper and lower limits associated with simple real regions may be chosenbased on the complexity of the background image 140, the image qualityrequirements associated with a given video-encoding standard, the time allotted tothe video-encoding process, and the like. For example, a particular video-encodingstandard may set a range of valid values for the quantization parameter, and thethreshold upper and lower limits may define the boundaries of the range of validvalues according to the particular video-encoding standard. As another example,the third quantization parameter may be determined in an iterative process, and thethreshold upper and lower limits may represent a maximum and a minimum numberof iterations, respectively, that may be performed in the time allotted to the videoencoding process. As a further example, the threshold upper and lower limits mayrepresent image quality criterion associated with simple real regions. That is, thethreshold upper limit may represent a maximum image quality for simple realregions at a particular bit rate, such that the bit rate is not slowed by the additionaldetail included in the simple real regions, and the threshold lower limit mayrepresent a minimum image quality for simple real regions at the particular bit rate,such that a minimum image quality for simple real regions is maintained at theparticular bit rate. In some implementations, the maximum and minimum imagequalities for simple real regions at a particular bit rate may be lower than themaximum and minimum image qualities for complex virtual regions and themaximum and minimum image qualities for medial regions at the particular bitrate,to ensure that more bits are allocated to the complex virtual regions and medialregions than to the simple real regions. The electronic device encodes each of thereal regions that do not satisfy the second complexity criterion based on thecorresponding third quantization parameter at step 580. While the process 500illustrates three types of regions-complex virtual regions, medial regions, andsimple real regions-any number of types of regions and corresponding complexitycriterion, initial quantization parameters associated with the types of regions, andupper and lower threshold limits associated with the types of regions may be usedinstead.

[0044] FIG. 6 shows an example diagram of the XR video frame 100 shownin FIG. 1 divided into regions based on the first and second complexity criteriadiscussed herein with respect to process 500. The virtual region 610 includes thevirtual object 150 and a portion of the background image 140 around the virtualobject 150, showing the surface of the table 130 and a portion of the rug 120. Thevirtual region 610 satisfies the first complexity criterion and so is encoded using thefirst quantization parameter. The simple real region 620 includes portions of thebackground image 140 that do not satisfy the second complexity criterion andshows the dresser 110, a portion of the rug 120, and a portion of the table 130. Thesimple real region 620 is encoded using the third quantization parameter. Themedial region 630 includes portions of the background image 140 that satisfy thesecond complexity criterion and shows the legs of the table 130 against a portion ofthe rug 120. The medial region 630 is encoded using the second quantizationparameter. The negative space in the simple real region 620 indicates where thevirtual region 610 and the medial region 630 are located. The virtual region 610,the simple real region 620, and the medial region 630 may be divided into one ormore additional, smaller regions to allow further refinement of the quantizationparameters based on the complexity, contrast, etc. in different portions of eachregion.

[0045] FIGs. 7A-C show, in flowchart form, an example process 700 forencoding an XR video frame based on an adaptive quantization matrix, first andsecond complexity criteria, and adjusted region sizes. For purposes of explanation,the following steps are described as being performed by particular components,However, it should be understood that the various actions may be performed byalternate components. In addition, the various actions may be performed in adifferent order. Further, some actions may be performed simultaneously, and somemay not be required, or others may be added. For ease of explanation, the process700 is described with reference to the process 200 described herein with referenceto FIG. 2 and the process 500 described herein with reference to FIGs. 5A-C.

[0046] The flowchart 700 begins in FIG. 7A with steps 210, 220, and 230as described above with reference to FIG. 2. After dividing the XR video frame intoat least one virtual region and at least one real region, the electronic device proceedsto step 510 and determines whether at least one virtual region satisfies a firstcomplexity criterion as described above with reference to process 500A shown inFIG. 5A. In response to determining at least one of the virtual regions satisfies thefirst complexity criterion, the electronic device may optionally proceed to step 710and determines, for each of the virtual regions that satisfy the first complexitycriterion, a corresponding region size based on an initial region size associated withcomplex virtual regions. The region size may be chosen such that complex portionsof the XR video frame have smaller region sizes and simple portions of the XR video frame have larger region sizes.

[0047] The electronic device may then optionally, for each of the virtualregions that satisfy the first complexity criterion and based on the correspondingregion size, divide the particular virtual region into one or more additional virtualregions at step 720. The electronic device proceeds to step 520 and determines, foreach of the virtual regions and additional virtual regions that satisfy the firstcomplexity criterion, a corresponding first quantization parameter based on aninitial quantization parameter associated with complex virtual regions as describedabove with reference to process 500A shown in FIG. 5A. At step 530, the electronicdevice encodes each of the virtual regions and additional virtual regions that satisfythe first complexity criterion based on the corresponding first quantizationparameter as described above with reference to process 500A shown in FIG. 5A.

[0048] Returning to step 510, in response to determining at least one virtualregion does not satisfy the first complexity criterion, the electronic device mayoptionally determine, for each of the virtual regions that do not satisfy the first complexity criterion, a corresponding region size based on an initial region sizeassociated with medial regions at step 730 shown in process 700B of FIG. 7B. Theinitial region size associated with medial regions may be larger than the initialregion size associated with complex virtual regions. The electronic device may thenoptionally, for each of the virtual regions that do not satisfy the first complexitycriterion and based on the corresponding region size, divide the particular regioninto one or more additional regions at step 740.

[0049] At step 550, the electronic device determines, for each of the virtualregions and additional virtual regions that do not satisfy the first complexitycriterion, a corresponding second quantization parameter based on an initialquantization parameter associated with medial regions as described above withreference to process 500B shown in FIG. 5B. The electronic device encodes eachof the virtual regions and additional virtual regions that do not satisfy the firstcomplexity criterion based on the corresponding second quantization parameter atstep 560 as described above with reference to process 500B shown in FIG. 5B.

[0050] Returning to the at least one real region from step 230, the electronicdevice determines whether the at least one real region satisfies a second complexitycriterion at step 540 as described above with reference to process 500A shown inFIG. 5A. In response to the at least one real region satisfying the second complexitycriterion, the electronic device may optionally proceed to step 730 shown in process700B of FIG. 7B and described above. At step 730, the electronic device mayoptionally determine, for each of the real regions that satisfy the second complexitycriterion, a corresponding region size based on the initial region size associated withmedial regions. The electronic device may optionally proceed to step 740 and foreach of the real regions that satisfy the second complexity criterion and based onthe corresponding region size, divide the particular real region into one or moreadditional real regions.

[0051] At step 550, the electronic device determines, for each of the realregions and additional real regions that satisfy the second complexity criterion, acorresponding second quantization parameter based on an initial quantizationparameter associated with medial regions as described above with reference toprocess 500B shown in FIG. 5B. The second quantization parameters for the realregions and additional real regions satisfying the second complexity criterion maybe the same or different than the second quantization parameters for the virtualregions and additional virtual regions not satisfying the first complexity criterion.The electronic device then encodes each of the real regions and additional realregions that satisfy the second complexity criterion based on the correspondingsecond quantization parameter at step 560 as described above with reference toprocess 500B shown in FIG. 5B.

[0052] Returning to step 540, in response to determining the at least onereal region does not satisfy the second complexity criterion, the electronic devicemay optionally proceed to step 750 shown in process 700C of FIG. 7C. At step 750,the electronic device may optionally determine, for each of the real regions that donot satisfy the second complexity criterion, a corresponding region size based onan initial region size associated with simple real regions. The initial region sizeassociated with simple real regions may be larger than the initial region sizeassociated with medial regions and the initial region size associated with complexvirtual regions. The electronic device may optionally proceed to step 760 and foreach of the real regions that do not satisfy the second complexity criterion and basedon the corresponding region size, divide the particular real region into one or moreadditional real regions.

[0053] At step 570, the electronic device determines, for each of the realregions and additional real regions that do not satisfy the second complexitycriterion, a corresponding third quantization parameter based on an initialquantization parameter associated with simple real regions as described above with reference to process 500C shown in FIG. 5C. The electronic device encodes eachof the real regions and additional real regions that do not satisfy the secondcomplexity criterion based on the corresponding third quantization parameter atstep 580 as described above with reference to process 500C shown in FIG. 5C.While the process 700 illustrates three types of regions-complex virtual regions,medial regions, and simple real regions-any number of types of regions andcorresponding complexity criterion, initial region sizes associated with the types ofregions, initial quantization parameters associated with the types of regions, andupper and lower threshold limits associated with the types of regions may be usedinstead.

[0054] FIG. 8 shows an example diagram of a medial region 630 of the XRvideo frame 100 divided into regions based on the first and second complexitycriteria and adjusted region sizes discussed herein with respect to process 700. Themedial region 630 includes portions of the background image 140 that satisfy thesecond complexity criterion and shows the legs of the table 130 against a portion ofthe rug 120. The medial region 630 is divided into additional medial regions 810,820, and 830. The additional medial region 810 includes two legs of the table 130against a portion of the rug 120. The additional medial region 820 includes a portionof the rug 120. The additional medial region 830 includes two legs of the table 130against a portion of the rug 120. The initial region size associated with medialregions may cause the electronic device to determine a smaller region size formedial region 630, and divide medial region 630 into the additional, smaller medialregions 810, 820, and 830. While FIG. 8 shows the medial region 630 divided intothree additional, smaller medial regions 810, 820, and 830, the medial regions maybe divided into any number of additional medial regions. In addition, the additionalmedial regions 810, 820, and 830 may be the same or different sizes.

[0055] The corresponding second quantization parameters for the medialregions 810 and 830 may be smaller than the corresponding second quantizationparameter for the medial region 820 to account for the added edge complexity,contrast, and the like of the legs of the table 130 against a portion of rug 120 inmedial regions 810 and 830 compared to the medial region 820 showing only aportion of the rug 120. That is, the medial regions 810 and 830 may be allocatedmore bits and a higher image resolution than the medial region 820 during video-encoding. FIG. 8 shows an example diagram of additional medial regions 810, 820,and 830 for the medial region 630, but complex virtual region 610 and simple realregion 620 may be similarly divided into additional regions.

[0056] Referring to FIG. 9, a simplified block diagram of an electronicdevice 900 is depicted, communicably connected to additional electronic devices980 and a network device 990 over a network 905, in accordance with one or moreembodiments of the disclosure. Electronic device 900 may be part of amultifunctional device, such as a mobile phone, tablet computer, personal digitalassistant, portable music / video player, wearable device, head-mounted systems,projection-based systems, base station, laptop computer, desktop computer,network device, or any other electronic systems such as those described herein.Electronic device 900, additional electronic device 980, and / or network device 990may additionally, or alternatively, include one or more additional devices withinwhich the various functionality may be contained, or across which the variousfunctionality may be distributed, such as server devices, base stations, accessorydevices, and the like. Illustrative networks, such as network 905 include, but arenot limited to, a local network such as a universal serial bus (USB) network, anorganization's local area network, and a wide area network such as the Internet.According to one or more embodiments, electronic device 900 is utilized to enablea multi-view video codec. It should be understood that the various components andfunctionality within electronic device 900, additional electronic device 980 andnetwork device 990 may be differently distributed across the devices, or may bedistributed across additional devices.

[0057] Electronic device 900 may include one or more processors 910, suchas a central processing unit (CPU). Processor(s) 910 may include a system-on-chipsuch as those found in mobile devices and include one or more dedicated graphicsprocessing units (GPUs). Further, processor(s) 910 may include multipleprocessors of the same or different type. Electronic device 900 may also include amemory 930. Memory 930 may include one or more different types of memory,which may be used for performing device functions in conjunction withprocessor(s) 910. For example, memory 930 may include cache, ROM, RAM, orany kind of transitory or non-transitory computer readable storage medium capableof storing computer readable code. Memory 930 may store various programmingmodules for execution by processor(s) 910, including video encoding module 935,renderer 940, a gaze-tracking module 945, and other various applications 950.Electronic device 900 may also include storage 920. Storage 920 may include onemore non-transitory computer-readable mediums including, for example, magneticdisks (fixed, floppy, and removable) and tape, optical media such as CD-ROMs and digital video disks (DVDs), and semiconductor memory devices such asElectrically Programmable Read-Only Memory (EPROM), and ElectricallyErasable Programmable Read-Only Memory (EEPROM). Storage 920 may beconfigured to store virtual object data 925, according to one or more embodiments.Electronic device may additionally include a network interface 970 from which theelectronic device 900 can communicate across network 905.

[0058] Electronic device 900 may also include one or more cameras 960 orother sensors 965, such as a depth sensor, from which depth of a scene may bedetermined. In one or more embodiments, each of the one or more cameras 960may be a traditional RGB camera, or a depth camera. Further, cameras 960 mayinclude a stereo- or other multi-camera system, a time-of-flight camera system, orthe like. Electronic device 900 may also include a display 975. The display device975 may utilize digital light projection, OLEDs, LEDs, uLEDs, liquid crystal onsilicon, laser scanning light source, or any combination of these technologies. Themedium may be an optical waveguide, a hologram medium, an optical combiner,an optical reflector, or any combination thereof. In one embodiment, thetransparent or translucent display may be configured to become opaque selectively.Projection-based systems may employ retinal projection technology that projectsgraphical images onto a person's retina. Projection systems also may be configuredto project virtual objects into the physical environment, for example, as a hologramor on a physical surface.

[0059] Storage 920 may be utilized to store various data and structureswhich may be utilized for dividing an XR video frame into virtual and real regionsand encoding the virtual regions based on a first quantization parameter and the realregions based on a second quantization parameter. According to one or moreembodiments, memory 930 may include one or more modules that comprisecomputer readable code executable by the processor(s) 910 to perform functions.The memory 930 may include, for example a video encoding module 935 whichmay be used to encode an XR video frame, a renderer 940 which may be used togenerate an XR video frame, a gaze-tracking module 945 which may be used todetermine a user's gaze position and an area of interest in the image stream, as wellas other applications 950.

[0060] Although electronic device 900 is depicted as comprising thenumerous components described above, in one or more embodiments, the variouscomponents may be distributed across multiple devices. Accordingly, althoughcertain calls and transmissions are described herein with respect to the particularsystems as depicted, in one or more embodiments, the various calls andtransmissions may be made differently directed based on the differently distributedfunctionality. Further, additional components may be used, some combination ofthe functionality of any of the components may be combined.

[0061] Referring now to FIG. 10, a simplified functional block diagram ofan illustrative programmable electronic device 1000 for providing access to anapp store is shown, according to one embodiment. Electronic device 1000 couldbe, for example, a mobile telephone, personal media device, portable camera, or atablet, notebook or desktop computer system, network device, wearable device, orthe like. As shown, electronic device 1000 may include processor 1005, display1010, user interface 1015, graphics hardware 1020, device sensors 1025 (e.g.,proximity sensor / ambient light sensor, accelerometer and / or gyroscope),microphone 1030, audio codec(s) 1035, speaker(s) 1040, communicationscircuitry 1045, image capture circuit or unit 1050, which may, e.g., comprisemultiple camera units / optical sensors having different characteristics (as well ascamera units that are housed outside of, but in electronic communication with,device 1000), video codec(s) 1055, memory 1060, storage 1065, andcommunications bus 1070.

[0062] Processor 1005 may execute instructions necessary to carry out orcontrol the operation of many functions performed by device 1000 (e.g., such asthe generation and / or processing of app store metrics accordance with the variousembodiments described herein). Processor 1005 may, for instance, drive display1010 and receive user input from user interface 1015. User interface 1015 cantake a variety of forms, such as a button, keypad, dial, a click wheel, keyboard,display screen and / or a touch screen. User interface 1015 could, for example, bethe conduit through which a user may view a captured video stream and / orindicate particular images(s) that the user would like to capture or share (e.g., byclicking on a physical or virtual button at the moment the desired image is beingdisplayed on the device's display screen).

[0063] In one embodiment, display 1010 may display a video stream as itis captured while processor 1005 and / or graphics hardware 1020 and / or imagecapture circuitry contemporaneously store the video stream (or individual imageframes from the video stream) in memory 1060 and / or storage 1065. Processor1005 may be a system-on-chip such as those found in mobile devices and includeone or more dedicated graphics processing units (GPUs). Processor 1005 may bebased on reduced instruction-set computer (RISC) or complex instruction-setcomputer (CISC) architectures or any other suitable architecture and may includeone or more processing cores. Graphics hardware 1020 may be special purposecomputational hardware for processing graphics and / or assisting processor 1005perform computational tasks. In one embodiment, graphics hardware 1020 mayinclude one or more programmable graphics processing units (GPUs).

[0064] Image capture circuitry 1050 may comprise one or more cameraunits configured to capture images, e.g., in accordance with this disclosure.Output from image capture circuitry 1050 may be processed, at least in part, byvideo codec(s) 1055 and / or processor 1005 and / or graphics hardware 1020, and / ora dedicated image processing unit incorporated within circuitry 1050. Images socaptured may be stored in memory 1060 and / or storage 1065. Memory 1060 mayinclude one or more different types of media used by processor 1005, graphicshardware 1020, and image capture circuitry 1050 to perform device functions.For example, memory 1060 may include memory cache, read-only memory(ROM), and / or random access memory (RAM). Storage 1065 may store media(e.g., audio, image and video files), computer program instructions or software,preference information, device profile information, and any other suitable data.Storage 1065 may include one more non-transitory storage mediums including,for example, magnetic disks (fixed, floppy, and removable) and tape, opticalmedia such as CD-ROMs and digital video disks (DVDs), and semiconductormemory devices such as Electrically Programmable Read-Only Memory(EPROM), and Electrically Erasable Programmable Read-Only Memory(EEPROM). Memory 1060 and storage 1065 may be used to retain computerprogram instructions or code organized into one or more modules and written inany desired computer programming language. When executed by, for example,processor 1005, such computer program code may implement one or more of themethods described herein. Power source 1075 may comprise a rechargeablebattery (e.g., a lithium-ion battery, or the like) or other electrical connection to apower supply, e.g., to a mains power source, that is used to manage and / or provideelectrical power to the electronic components and associated circuitry ofelectronic device 1000.

[0065] It is to be understood that the above description is intended to beillustrative, and not restrictive. The material has been presented to enable anyperson skilled in the art to make and use the disclosed subject matter as claimed andis provided in the context of particular embodiments, variations of which will bereadily apparent to those skilled in the art (e.g., some of the disclosed embodimentsmay be used in combination with each other). Accordingly, the specificarrangement of steps or actions shown in FIGS. 2, 4, 5A-C, and 7A-C or thearrangement of elements shown in FIGS. 9 and 10 should not be construed aslimiting the scope of the disclosed subject matter. The scope of the inventiontherefore should be determined with reference to the appended claims, along withthe full scope of equivalents to which such claims are entitled. In the appendedclaims, the terms "including" and "in which" are used as the plain-Englishequivalents of the respective terms "comprising" and "wherein."

Claims

1. A method for encoding an extended-reality (XR) video frame, comprising: obtaining an XR video frame comprising a background image and a virtual object overlaying at least a portion of the background image; dividing the XR video frame into a virtual region and a real region, wherein the virtual region comprises at least a portion of the virtual object, and wherein the real region comprises a region of the background image separate from the virtual region; determining, for the virtual region, a corresponding first quantization parameter based on an initial quantization parameter associated with virtual regions; determining, for the real region, a corresponding second quantization parameter based on an initial quantization parameter associated with real regions; and encoding the virtual region based on the corresponding first quantization parameter and the real region based on the corresponding second quantization parameter.

2. The method of claim 1, wherein determining, for the virtual region, the corresponding first quantization parameter is further based on a threshold upper limit associated with virtual regions and a threshold lower limit associated with virtual regions.

3. The method of claim 1, wherein determining, the real region, the corresponding second quantization parameter is further based on a threshold upper limit associated with real regions and a threshold lower limit associated with real regions.

4. The method of claim 1, wherein: dividing the XR video frame further comprises dividing the XR video frame into a medial region; the virtual region comprises at least a portion of the virtual object that satisfies a first complexity criterion; the real region comprises a first portion of the region of the background image separate from the virtual region, wherein the first portion of the region fails to satisfy a second complexity criterion; the at least one medial region comprises at least one of: (i) a portion of the at least one virtual object that fails to satisfy the first complexity criterion, and (ii) a second portion of the region of the background image separate from the virtual region, wherein the second portion of the region satisfies the second complexity criterion; and the method further comprising: determining, for the medial region, a corresponding third quantization parameter based on an initial quantization parameter associated with medial regions; and encoding the medial region based on the corresponding third quantization parameter.

5. The method of claim 4, wherein determining, for the medial region, the corresponding third quantization parameter is further based on a threshold upper limit associated with medial regions and a threshold lower limit associated with medial regions.

6. The method of claim 1, further comprising obtaining an input indicative of an area of focus via a gaze-tracking user interface, wherein dividing the XR video frame is based at least in part on the area of focus.

7. The method of claim 1, wherein the initial quantization parameter associated with virtual regions is smaller than the initial quantization parameter associated with real regions.

8. The method of claim 1, wherein the method comprises: for the virtual region: determining a corresponding first region size based on an initial region size associated with virtual regions; and dividing the virtual region into one or more additional virtual regions based on the corresponding first region size; for the real regions: determining a corresponding second region size based on an initial region size associated with real regions; and dividing the real region into one or more additional real regions based on the corresponding second region size.

9. The method of claim 8, wherein the initial region size associated with virtual regions is smaller than the initial region size associated with real regions.

10. The method of claim 4, wherein the initial quantization parameter associated with medial regions is smaller than the initial quantization parameter associated with real regions and larger than the initial quantization parameter associated with virtual regions.

11. The method of claim 4, wherein the method comprises: for the virtual region: determining a corresponding first region size based on an initial region size associated with virtual regions; and dividing the virtual region into one or more additional virtual regions based on the corresponding first region size; for the real region: determining a corresponding second region size based on an initial region size associated with real regions; and dividing the real region into one or more additional real regions based on the corresponding second region size; and for the medial regions: determining a corresponding third region size based on an initial region size associated with medial regions; and dividing the medial region into one or more additional medial regions based on the corresponding third region size.

12. The method of claim 11, wherein the initial region size associated with medial regions is larger than the initial region size associated with virtual regions and smaller than the initial region size associated with real regions.

13. A device, comprising: an image capturing device configured to capture a background image; and at least one processor, wherein the at least one processor is to: obtain an extended reality (XR) video frame comprising the background image and a virtual object overlaying at least a portion of the background image; divide the XR video frame into a virtual region and a real region, wherein the virtual region comprises at least a portion of the virtual object, and wherein the real region comprises a region of the background image separate from the virtual region; determine, for the virtual region, a corresponding first quantization parameter based on an initial quantization parameter associated with virtual regions; determine, for the real region, a corresponding second quantization parameter based on an initial quantization parameter associated with real regions; and encode the virtual region based on the corresponding first quantization parameter and the real region based on the corresponding second quantization parameter.

14. The device of claim 13, wherein to determine, for the virtual region, the corresponding first quantization parameter, the at least one processor is to determine the corresponding first quantization parameter further based on a threshold upper limit associated with virtual regions and a threshold lower limit associated with virtual regions.

15. The device of claim 13, wherein to determine, for the real region, the corresponding second quantization parameter, the at least one processor is to determine the corresponding second quantization parameter further based on a threshold upper limit associated with real regions and a threshold lower limit associated with real regions.

16. The device of claim 13, wherein: to divide the XR video frame, the at least one processor is to divide the XR video frame into a medial region; the virtual region comprises at least a portion of the virtual object that satisfies a first complexity criterion; the real region comprises a first portion of the region of the background image separate from the virtual region, wherein the first portion of the region fails to satisfy a second complexity criterion; and the medial region comprises at least one of: (i) a portion of the virtual object that fails to satisfy the first complexity criterion, and (ii) a second portion of the region of the background image separate from the virtual region, wherein the second portion of the region satisfies the second complexity criterion; and the at least one processor is to: determine, for the medial region, a corresponding third quantization parameter based on an initial quantization parameter associated with medial regions; and encode the medial region based on the corresponding third quantization parameters.

17. The device of claim 16, wherein to determine, for the medial region, the corresponding third quantization parameter, the at least one processor is to determine the corresponding third quantization parameter further based on a threshold upper limit associated with medial regions and a threshold lower limit associated with medial regions.

18. The device of claim 13, wherein the at least one processor is to obtain an input indicative of an area of focus via a gaze-tracking user interface, wherein to divide the XR video frame, the at least one processor is to divide the XR video frame based at least in part on the area of focus.