Object based depth image encoding

WO2026195155A1PCT designated stage Publication Date: 2026-09-24TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/057409
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-24

Smart Images

  • Figure EP2025057409_24092026_PF_FP_ABST
    Figure EP2025057409_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a technique for achieving higher precision for depth image data associated with remote rendered eXtended Reality, XR, scenes in situations where conventional 2D video encoders are used to encode depth information. Accordingly, the present embodiments provide an encoder node (200), such as a network node (16) in a cloud network (14), a decoder node (300), such as an XR device (e.g., a Head Mounted Display, HMD, (20)), and corresponding methods (30, 90) that use conventional 2-Dimensional, 2D, video compression technology to encode and decode depth image data. More particularly, the embodiments provided herein employ object-based depth scaling, thereby enabling higher quality depth precision for remote rendered eXtended Reality, XR, scenes even in situations where depth image data is encoded according to the limited bit-per-pixel, BPP, value of 2D video compression. Additionally, the present embodiments provided herein utilize the same encoded image to hold both the 2D and the depth information, thereby eliminating at least some of the overhead associated with synchronizing the 2D information and the depth information. Moreover, the embodiments presented herein calculate and add the depth information and the scale information for the 3D objects in a scene to the composite frame as supplemental information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] OBJECT BASED DEPTH IMAGE ENCODING

[0002] TECHNICAL FIELD

[0003] This application relates generally to depth encoding techniques, and more particularly to a method and apparatus for encoding depth into a video frame in extended Reality (XR) applications.

[0004] BACKGROUND

[0005] Various techniques are available for creating depth images for Extended Reality (XR) applications. Some techniques, for example, use different sensors to scan the actual environment for output to a display. Other techniques, however, render 3-dimensional (3D) virtual objects to the display. With distributed XR applications, the actual location where depth images are used can be different from where these images were created. In these cases, therefore, the depth images must be transmitted on a communication channel between these locations. However, some form of compression is necessary to transmit the images prior to transmission. Currently, there are two main approaches to this problem. The first approach is known as Geometry-based Point Cloud Compression (G-PCC), which is used for compressing depth or point cloud information. The other approach is the well-known 2-dimensional (2D) Video-based Point Cloud Compression (V-PCC), which is used to compress 3D point clouds into 2D video sequences.

[0006] Remote rendering is a technology that first renders 3D content on a server (e.g., a network server) and then streams that rendered content in real time to a receiving device (e.g., an XR headset). To accomplish this, XR applications process and render 3D scenes in a remote cloud processing environment and then send the 3D scenes to the receiving device. Remote rendering processes also exist for processing and rendering 2D images, along with techniques for improving that process. For example, some techniques use split rendering to improve the quality of the transport channel and reduce warp-related issues. In such cases, a game engine, for example, generates separate 2D graphics regions from 3D objects. The graphics regions are then augmented with information from a 3D simulation and encoded into a media stream as a Composite Video Frame (CVF).

[0007] In remote rendering XR use cases, transmitting depth images along with 2D images can improve the display quality of a presented scene. For example, transmitting the depth images can add depth to a rendered virtual object, add more realistic occlusion effects, and improve the efficiency and quality of warp processing.

[0008] SUMMARY

[0009] The present disclosure provides methods and corresponding devices (i.e., an encoder and a decoder) for encoding and decoding depth into a video frame, respectively. More specifically,the methods and techniques provided herein employ currently available 2D video compression technology to encode depth image data, thereby enabling hardware acceleration at both the encoder and the decoder. These methods and techniques also employ object-based depth scaling. This enables higher quality depth precision for remote rendered XR scenes even in situations where the depth information is encoded according to a limited bit-per-pixel (BPP) value (e.g., the 8-10 bits / pixel of 2D video compression). Further, the methods and techniques presented herein utilize the same encoded image to hold both the 2D and the depth information, thereby eliminating synchronization overhead. Moreover, for each graphics region, the methods and techniques presented herein calculate and add both the depth information and the scale information for the 3D objects in a scene to the composite frame as supplemental information.

[0010] Accordingly, in a first aspect, the present disclosure provides a method for encoding depth into a video frame. In this aspect, for each of a plurality of 3D objects in an image, the method comprises rendering the 3D object to a 2D graphics region and a depth graphics region, determining distance information and scale information for the 3D object in the depth graphics region, and adding the 2D graphics region and the depth graphics region to one or more composite frames. The method then further comprises generating a composite video frame comprising the one or more composite frames. Once generated, the method calls for encoding the composite video frame and the distance information and the scale information for the 3D object to generate an encoded composite video frame and sending the encoded composite video frame to a decoder in a bitstream.

[0011] In a second aspect, the present disclosure provides a method for decoding depth in a video frame. In this aspect, the method comprises, for each of a plurality of encoded composite video frames received in a bitstream, decoding the encoded composite video frame to generate a composite video frame comprising a 2D graphics region of an image, a depth graphics region of the image that is independent of the 2D graphics region, and distance information and scale information for a 3D object in the image. So decoded, the method calls for extracting the 2D graphics region and the depth graphics region from the composite video frame, adjusting the depth graphics region based on the distance information and the scale information for the 3D object, and rendering the 2D graphics region and the adjusted depth graphics region for output to a display.

[0012] In a third aspect, the present disclosure provides an encoder node for encoding depth into a video frame. In this aspect, the encoder node is configured to, for each of a plurality of 3D objects in an image, render the 3D object to a 2D graphics region and a depth graphics region, determine distance information and scale information for the 3D object in the depth graphics region, and add the 2D graphics region and the depth graphics region to one or more compositeframes. The encoder node is then further configured to generate a composite video frame comprising the one or more composite frames. Once generated, the encoder node is further configured to encode the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame and send the encoded composite video frame to a decoder in a bitstream.

[0013] In a fourth aspect, the present disclosure provides an encoder node for encoding depth into a video frame. In this aspect, the encoder node comprises communication circuitry and processing circuitry. The processing circuitry is operatively connected to the communication circuitry and, for each of a plurality of 3D objects in an image, is configured to render the 3D object to a 2D graphics region and a depth graphics region, determine distance information and scale information for the 3D object in the depth graphics region, and add the 2D graphics region and the depth graphics region to one or more composite frames. The processing circuitry is then further configured to generate a composite video frame comprising the one or more composite frames. Once generated, the processing circuitry is further configured to encode the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame and send the encoded composite video frame to a decoder in a bitstream.

[0014] In a fifth aspect, the present disclosure provides a computer program comprising instructions that, when executed by processing circuitry of an encoder node, causes the processing circuitry to perform a method according to the first aspect.

[0015] In a sixth aspect, the present disclosure provides a non-transitory computer-readable medium comprising a computer program that, when executed by processing circuitry of an encoder node, causes the processing circuitry to perform a method according to the first aspect.

[0016] In a seventh aspect, the present disclosure provides a decoder node for decoding depth in a video frame. In this aspect, the decoder node is configured to, for each of a plurality of encoded composite video frames received in a bitstream, decode the encoded composite video frame to generate a composite video frame comprising a 2D graphics region of an image, a depth graphics region of the image that is independent of the 2D graphics region, distance information for a 3D object in the image, and scale information for the 3D object. So processed, the decoder node is further configured to extract the 2D graphics region and the depth graphics region from the composite video frame, adjust the depth graphics region based on the distance information and the scale information for the 3D object, and render the 2D graphics region and the adjusted depth graphics region for output to a display.

[0017] In an eighth aspect, the present disclosure provides a decoder node for decoding depth in a video frame. In this aspect, the decoder node comprises communication circuitry andprocessing circuitry operatively connected to the communication circuitry. The processing circuitry is configured to, for each of a plurality of encoded composite video frames received in a bitstream, decode the encoded composite video frame to generate a composite video frame comprising a 2D graphics region of an image, a depth graphics region of the image that is independent of the 2D graphics region, distance information for a 3D object in the image, and scale information for the 3D object. The decoder node is further configured to extract the 2D graphics region and the depth graphics region from the composite video frame, adjust the depth graphics region based on the distance information and the scale information for the 3D object, and render the 2D graphics region and the adjusted depth graphics region for output to a display.

[0018] In a ninth aspect, the present disclosure provides a computer program comprising instructions that, when executed by processing circuitry of a decoder node, causes the processing circuitry to perform a method according to the second aspect.

[0019] In a tenth aspect, the present disclosure provides a non-transitory computer-readable medium comprising a computer program that, when executed by processing circuitry of a decoder node, causes the processing circuitry to perform a method according to the second aspect.

[0020] In an eleventh aspect, the present disclosure provides a system in a communication network for encoding and decoding depth in a video frame. In this aspect, the system comprises an encoder node and a decoder node. The encoder node is configured to, for each of a plurality of 3D objects in an image, render the 3D object to a 2D graphics region and a depth graphics region, determine distance information and scale information for the 3D object in the depth graphics region, and add the 2D graphics region and the depth graphics region to one or more composite frames. The processing circuitry is then further configured to generate a composite video frame comprising the one or more composite frames. The encoder node is also configured to encode the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame and send the encoded composite video frame to a decoder in a bitstream.

[0021] The decoder node of the system is configured to, for each of a plurality of encoded composite video frames received in the bitstream, decode the encoded composite video frame to generate the composite video frame comprising the 2D graphics region of an image, the depth graphics region of the image that is independent of the 2D graphics region, the distance information for the 3D object in the image, and the scale information for the 3D object. Once decoded, the decoder node is further configured to extract the 2D graphics region and the depth graphics region from the composite video frame, adjust the depth graphics region based on thedistance information and the scale information for the 3D object, and render the 2D graphics region and the adjusted depth graphics region for output to a display.

[0022] BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 illustrates a communication system configured according to an embodiment of the present disclosure.

[0024] Figure 2 is a flow diagram illustrating a method of depth encoding according to an embodiment of the present disclosure.

[0025] Figure 3 illustrates a composite video frame (CVF) having a plurality of 2D graphics regions and a plurality of depth graphics regions according to an embodiment of the present disclosure.

[0026] Figure 4 is a flow diagram illustrating a method of depth decoding according to an embodiment of the present disclosure.

[0027] Figure 5 is a schematic block diagram illustrating some exemplary functional components of a network node configured according to an embodiment of the present disclosure.

[0028] Figure 6 is a schematic block diagram illustrating some exemplary functional components of a client device, such as a XR headset, configured according to an embodiment of the present disclosure.

[0029] DETAILED DESCRIPTION

[0030] Using only 2D images to render 3D XR scenes, even in cases where stereo images are available, has limitations that negatively affect the quality of a displayed image. For example, handling occlusion artifacts in such cases is an especially difficult process absent 3D information. Another concern is the performance of compression techniques. Future depth or point cloud information specific (G-PCC) compression techniques, for example, may achieve better compression performance than do conventional compression techniques. However, none of these techniques are currently capable of handling real-time and low latency use cases such as those handled by remote rendering XR applications.

[0031] Traditional 2D video compression-based solutions, especially in situations where currently existing algorithms are used, also present complications. For example, conventional 2D video compression techniques attempt to enable real-time and low latency use cases by using hardware accelerated encoders and decoders. However, despite such hardware acceleration, limitations such as depth image quality still need to be considered. Particularly, current videobased PCC applications can only be used in limited depth scenarios. For example, full XR scenes with high dynamic range - i.e., where the objects in a scene are very close and very far from a viewpoint of the camera - are encoded to the limited bit resolution (e.g., 8-bit) of the 2D video compression. Such processes degrades depth image quality.The present disclosure, however, addresses these and other issues by providing a system, devices, and corresponding methods for achieving a higher precision for depth image data of remote rendered XR scenes in situations where conventional 2D video encoders are used to encode depth information. To accomplish this, the present embodiments use, inter alia, depth image information determined from a depth image associated with a scene. Particularly, the objects of a depth image comprise a plurality of pixels. Each depth image pixel has a respective pixel value represented by 32-bit floating-point number. According to the present embodiments, these depth image pixel values are used to measure the distance between the camera that captured the scene having the 3D object and a part of the 3D object currently represented by the depth image pixel.

[0032] In more detail, a graphics engine (e.g., a game engine) renders separate 2D graphics regions and depth graphics regions from the 3D objects in a 3D scene and allocates unique ObjectID for each of those 3D objects. The width and distance of every 3D object in the depth graphics region is then calculated separately, while the scale of those 3D objects is calculated using the width of the depth graphics region and the available bit depth of the encoder. So calculated, the pixels of depth graphics regions are then scaled to the available bit depth based on the scale calculated from the width of the depth graphics region.

[0033] Next, a composite image is created from both the 2D graphics regions and the scaled depth graphics regions. To avoid synchronization latency, a single composite image is created by adding two images side-by-side and then encoding the composite image using a state-of-the-art 2D video encoder. In at least one embodiment, for every 2D graphics region, the ObjectID, distance information, and scale information are added to the composite image (i.e., a composite frame) as supplemental information, and sent to a decoder (e.g., an XR device, such as an XR headset) in a bitstream.

[0034] Upon receipt of the bitstream, the decoder decodes the images using hardware accelerated video decoding. Specifically, the color information of the objects is reconstructed using the separated 2D side of the composite image (i.e., the 2D graphics region), while the depth information of the objects is reconstructed from the separated depth side of the composite image (i.e., the depth graphics region) using the distance and scale information sent by the encoder. So processed, the resultant images are rendered for display to a user of a XR device.

[0035] As detailed throughout this disclosure, the present embodiments provide benefits and advantages that conventional methods and techniques of depth encoding / decoding cannot or do not provide. For example, object-based depth scaling enables higher quality depth precision for XR scenes even in cases where the depth information is encoded with a limited BPP value (e.g., the 8-10 bits / pixel of 2D video compression). Additionally, in some XR applications (e.g.,Augmented Reality (AR) and Mixed Reality (MR) applications), the depth of a scene can be on the order of a few meters. For outdoor applications, the depth of a scene can be even higher. By using object-based scaling with high depth dynamic range (i.e., where the objects in a scene are both very close and very far away), the present embodiments achieve better depth image quality results even in situations where the encoding adheres to the limited bit resolution of the 2D video compression.

[0036] In another advantage, the present embodiments use currently available 2D video compression technology to encode the depth information, thereby enabling the use of conventional hardware acceleration circuitry and functionality at both the encoder and the decoder. The present embodiments also eliminate, or greatly reduce, synchronization overhead by using the same encoded image to hold both the 2D graphics data and the depth data. This is an important advantage for low latency XR applications. Particularly, the depth and scale information for an object are encoded and transmitted to the decoder in an encoded composite video frame together with the image that comprises the object. As such, the decoder has everything it needs to process the image upon receipt of the encoded composite video frame. There is no need for the decoder to delay processing the image until it receives a separate transmission containing the depth and scale information.

[0037] Turning now to the drawings, Figure 1 illustrates a communications system 10 configured according to one embodiment of the present disclosure. As seen in Figure 1, system 10 comprises an access network 12 communicatively connecting a client device (e.g., a headmounted display, HMD, 20) with a network node 16 disposed in a cloud network 14. As described in more detail later, network node 16 may comprise encoding circuitry that encodes data and images for display to a user in XR scenarios. Similarly, HMD 20 may comprise decoder circuitry for decoding the encoded images and data sent by network node 16; however, the present embodiments are not so limited. In some situations, a computing device 18 may be disposed between HMD 20 and network node 16. In these cases, computing device 18 can be configured to perform at least some of the processing functions of HMD 20 (e.g., decoding a bitstream sent by network node 16). Therefore, in such embodiments, computing device 18 may comprise the decoder circuitry that performs the decoding functions.

[0038] The access network 12 may be any type of communications network (e.g., WiFi, ETHERNET, Wireless LAN (WLAN), 3G, LTE, etc.), and functions to connect subscriber devices, such as HMD 20, to one or more service provider nodes, such as network node 16. The cloud network 14 provides such subscriber devices with “on-demand” availability of computer resources (e.g., memory, data storage, processing power, etc.) without requiring the user to directly, actively manage those resources. According to embodiments of the present disclosure,such resources include, but are not limited to, one or more XR applications being executed on network node 16. The XR applications may comprise, for example, gaming applications and / or simulation applications used for training.

[0039] In general, one or more sensors (not shown) on HMD 20 measure the translational and / or rotational movement of the user’s head as the user views images rendered by network node 16 on HMD 20. Signals representing the detected and measured movement are then sent to network node 16. Upon receipt, network node 16 utilizes those signals to compensate the video images for the user’s movement and sends the compensated video to HMD 20 for display to a user. In some situations, to help reduce and / or eliminate latency associated with the communications between HMD 20 and network node 16, the present embodiments place network node 16 in an Edge Data Network (EDN).

[0040] As stated above, hardware accelerated encoders and decoders are necessary for both 2D images and depth images to enable real-time and low latency use cases. Where traditional 2D video compression-based solutions are used, especially in situations where traditional encoding / decoding algorithms are used, such hardware accelerated encoders and decoders are already available on both the network side and the client (e.g., XR headset) side.

[0041] Additionally, XR scenes, in general, can have wide depth range. With Virtual Reality (VR) applications, the depth range is limited to the actual virtual view; however, with Augmented Reality (AR) or Mixed Reality (MR) applications, a scene can have a depth range anywhere from a few meters to much higher in outdoor applications. Further, in remote rendering scenarios, the full depth XR scenes need to be streamed to the client device. In AR and MR applications, only the virtual objects need to be streamed to the client device. Nevertheless, the virtual objects are in the same depth range. Thus, the depth image quality will be degraded in situations where full XR scenes having high depth dynamic range (i.e., objects in the same scene that are very close to and very far from the camera) are encoded to the limited bit resolution of the 2D video compression.

[0042] Figure 2 is a flow diagram illustrating a method 30, implemented at a 2D encoder, of depth encoding according to an embodiment of the present disclosure. As seen in Figure 2, one or more 3D objects in a scene captured by a camera are input into a simulation engine (e.g., a game engine). According to the present embodiments, a 3D object is selected for processing (box 32). A unique ObjectID is then allocated to the selected 3D object and is maintained throughout the lifetime of the 3D object (box 34). The simulation engine then renders separate 2D graphics regions and depth graphics regions from the selected 3D object and associates the unique ObjectID for the selected 3D object with both the 2D graphics region and depth graphics region (box 36).The 2D graphics region and the depth graphics region are then processed on two separate, independent paths. For example, in embodiments of the present disclosure, the 2D graphics regions are processed along path 38. In this embodiment, a variety of techniques may be used to process the 2D graphics region. Such techniques include, but are not limited to, split-rendering processes for reducing undesirable visual artifacts, such as judder, processes supporting the transmission of the video data of a 3D scene and object information (e.g., pose information comprising a position and orientation of one or more virtual objects within the 3D scene) over various transport channels having different performance characteristics (e.g., low latency channels), and processes for avoiding disocclusion artifacts. Regardless of the particular processing techniques used along path 38, however, the 2D graphics region is subsequently added (along with previously processed 2D graphics regions) to a composite frame (box 46).

[0043] The depth graphics regions, however, are separately processed using different techniques and along an entirely independent processing path 40. Particularly, the 2D encoder configured to implement method 30 first calculates the width and distance of the selected 3D object for the depth graphics region (box 42). In one embodiment, the width Wnof the 3D object is determined

[0044]

[0045] where:

[0046] maxis the highest value depth pixel and represents the point of the nth3D object that is n

[0047] furthest from the camera that captured the image; and

[0048] is the lowest value depth pixel and represents the point of the nth3D object that is closest to the camera that captured the image.

[0049] The distance Dnof the nth3D object is determined as being equal to the distance between the camera that captured the image having the nth3D object and the point of the nth3D object that is closest to the camera. In this embodiment, the point at which the nth3D object is closest to the camera (i.e., the minimum distance between the nth3D object and the camera) is equal to

[0050]

[0051] Next, the 2D encoder implementing method 30 determines the scaling information for the region (box 44). In this embodiment, the scale of the nth3D object is calculated as the ratio between the previously calculated width of the nth3D object and the bit depth of the 2D image encoder.

[0052]

[0053] where:

[0054] Snis a scale of the nth3D object;

[0055]

[0056] the width of the 3D object; and

[0057] 2Nis a bit depth of the encoder.

[0058] Once the scaling information has been determined, the 2D encoder scales the pixels of the current depth graphics region (i.e., the depth graphics region associated with the ObjectID of the nth3D object) (box 46) and then adds both the 2D graphics region and the depth graphics region to one or more composite frames (box 48). In one embodiment, for example, both the 2D graphics region and the depth graphics region are added to a single composite frame. In other embodiments, however, the 2D graphics region is added to a composite frame for 2D graphics regions while the depth graphics regions are added to a composite frame for depth graphics regions.

[0059] Additionally, in this embodiment, the higher resolution depth pixels (e.g., measured in meters and represented using 32-bit floating point numbers) are mapped to lower bit resolution (typically 8-bit) 2D image pixel values according to:

[0060]

[0061] where:

[0062] °n is a scaled depth pixel value;

[0063] is an original depth pixel value;

[0064] Dnis the distance of the depth pixel from the camera that captured the image; and

[0065] Snis a scale of the nth3D object.

[0066] According to one or more embodiments, every scaled depth pixel value d*Yof the nth3D object that is encoded into the 2D image is calculated using this equation.

[0067] This process (i.e., illustrated in box 32 - box 48) continues in a loop until the there are no more 3D objects to process (box 50). Then, once all 3D objects have been processed, the 2D encoder assembles (i.e., combines) all composite frames having the 2D graphics regions and depth graphics regions created by the simulation engine, along with the width W, distance D, and scaling S information for those frames, into a single composite video frame 60 (box 52). An example composite video frame 60 showing its constituent 2D graphics regions 70 and depth graphics regions 80 is illustrated in Figure 3. Assembling regions 70, 80 into a single composite video frame, as described herein, ensures that each 2D graphics region 70 and depth graphics region 80 can be decoded in a stand-alone manner (i.e., independently of each other). Further, as seen in Figure 3, the meta-information, including but not limited to the depth and the scalevalues Dn, Sn, respectively. Once assembled, the single composite video frame is encoded using a 2D encoder (box 54) and then added to a bitstream for transmission to a decoder, such as HMD 20, for example (box 56).

[0068] It should be noted that most composite video frame features are supported by modem video encoders. Nevertheless, extensions may still be needed in some situations. For example, coding tools, such as those that support the H.264 Baseline Profile features Flexible Macroblock Ordering, FMO, and Arbitrary Slice Ordering, ASO, already exist. These tools enable parts of a composite video frame to be grouped together, as well as set different coding parameters and transmission ordering. Additionally, some existing video codecs also support encoding non-rectangular areas. Any of these tools may be utilized with any of the embodiments disclosed herein.

[0069] Notwithstanding the above, in one or more embodiments, the method 30 further comprises allocating an ObjectID to the 3D object uniquely identifying the 3D object and associating the ObjectID with both the 2D graphics region 70 and the depth graphics region 80.

[0070] In at least one embodiment, the 2D graphics region 70 is processed independently of the depth graphics region 80.

[0071] Additionally, in some embodiments, the 3D object in the depth graphics region 80 comprises a plurality of depth pixels d. Each depth pixel dnhas a respective depth value representing a distance Dnof the depth pixel dnfrom a camera that captured the image.

[0072] As stated above, embodiments of the present disclosure determine the width Wnof the 3D object as:

[0073]

[0074] where:

[0075] is the highest value depth pixel of the nth3D object and represents the point of the nth3D object that is furthest from the camera that captured the image; and

[0076] is the lowest value depth pixel of the nth3D object and represents the point of the nth3D object that is closest to the camera that captured the image.

[0077] In at least one embodiment, the distance information for the 3D object in the depth graphics region 80 comprises ^}min.

[0078] n

[0079] In one embodiment, the scale information Snfor the 3D object is determined as:

[0080]

[0081] where:

[0082] Snis a scale of the nth3D object;

[0083]

[0084] the width of the nth3D object; and

[0085] 2Nis a bit depth of the encoder.

[0086] Additionally, in one embodiment, method 30 further comprises scaling (46) one or more pixels in the depth graphics region 80 according to the scale information.

[0087] In one embodiment, scaling the one or more pixels in the depth graphics region comprises mapping high-resolution depth pixels to low-resolution depth pixels according to:

[0088] "

[0089]

[0090] where:

[0091] rfxy

[0092] n is a scaled depth pixel value;

[0093] is an original depth pixel value;

[0094] Dnis the distance of the depth pixel from the camera that captured the image; and

[0095] Snis a scale of the nth3D object.

[0096] In at least some embodiments, each high-resolution depth pixel is represented by a first resolution value, and each low-resolution depth pixel is represented by a second resolution value that is less than the first resolution value.

[0097] In one embodiment, the first resolution value is a 32-bit value, and the second resolution value is an 8-bit value.

[0098] In one embodiment, generating a composite video frame comprising the one or more composite frames comprises assembling the one or more composite frames into the composite video frame 60 responsive to determining 50 that the distance information and the scale information have been determined for each of the 3D objects in the depth graphics region.

[0099] In one embodiment, encoding the composite video frame further comprises encoding the scale information 5nand the distance information Dnfor the 3D object to generate the encoded composite video frame.

[0100] In one embodiment, the distance information Dnand the scale information Sncomprise metadata.

[0101] In one embodiment, each of the 2D graphics region 70 and the depth graphics region 80 are independently decodable.

[0102] In one embodiment, method 30 is implemented by a user device, such as HMD 20, for example.

[0103] In another embodiment, however, method 30 is implemented by a network node disposed in the cloud, such as network node 16, for example.Figure 4 is a flow diagram illustrating a method 90 of depth decoding according to an embodiment of the present disclosure. In this embodiment, method 90 is implemented by HMD 20. However, this is merely for illustrative purposes. Those of ordinary skill in the art should readily appreciate that the logic of this embodiment may be implemented at some other computing device, such as computing device 18, for example, and the results provided to HMD 20 for display to a user.

[0104] As seen in Figure 4, a decoder at HMD 20 decodes an encoded composite video frame received from an encoder, such as the encoder implemented at network node 16, for example (box 92). So decoded, the decoder recreates the graphics regions structure based on the meta-data included in the composite frame. In this embodiment, for example, the decoder separates the depth information (contained in the depth graphics region 70) and the RGB frame (contained in the 2D graphics region 80) from the decoded composite video frame (box 94), selects a graphics region (box 96), and then extracts the selected graphics region from the decoded composite video frame (box 98). The 2D graphics region 70 and the depth graphics region 80 are then processed on two separate, independent paths 100, 102, respectively.

[0105] In this embodiment, the 2D graphics region 70 is processed along path 100 before and provided to the HMD 20 rendering logic (box 108). As previously described, various techniques may be utilized to process the 2D graphics region 70 including, but not limited to, split-rendering processes for reducing undesirable visual artifacts, such as judder, processes supporting the reception of video data of a 3D scene and object information (e.g., pose information comprising a position and orientation of one or more virtual objects within the 3D scene) over various transport channels having different performance characteristics (e.g., low latency channels), and processes for avoiding disocclusion artifacts. The corresponding depth graphics region 80, however, is further processed by the decoder logic at HMD 20 along path 102. In this embodiment, for example, decoder circuitry at HMD 20 uses the scaling information Snreceived with the composite video frame to unscale the pixels associated with the currently selected depth graphics region 80 (box 104). Similarly, the decoder uses the distance information Dnreceived with the composite video frame to add the distance back into the value of the depth pixels (box 106). By way of example only, this embodiment performs these two functions according to the following equation:

[0106] d nxv= d^ n '-S nn+ D nn

[0107] where:

[0108] is the original nthdepth pixel value;

[0109] is the scaled nthdepth pixel value;Snis a scale of the nth3D object; and

[0110] Dnis the distance of the nthdepth pixel from the camera that captured the image.

[0111] Thus, according to the present disclosure, this single equation recovers the original depth pixel valuexvof the nth3D object using the distance and scaling information received with the composite video frame and according to the techniques described in U.S. Patent No. 11,798,196, which is incorporated herein by reference in its entirety. Particularly, the decoder multiplies the scaled depth pixel value d*yby the scale Snto unscale the depth pixel and then adds the distance Dnback to the depth pixel value to restore the distance.

[0112] Next, the 2D graphics region 70 and the unsealed depth graphics region 80, together with the corresponding ObjectID, are provided to the rendering logic at HMD 20. The rendering logic creates a 2D image using the decoded 2D information and depth information of the “remote” 3D objects (i.e., 3D objects captured by the camera and encoded at network node 16), as well as the 3D information associated with any local objects (i.e., objects at or near HMD 20). The result is then placed to the current framebuffer (box 108) and the process repeated so long as the current depth graphics region 80 still has 3D objects to process (box 110). However, once rendering is complete, the frame is provided for display on HMD 20 (box 112).

[0113] In one embodiment, adjusting the depth graphics region 70 based on the distance information Dnand the scale information Snfor the 3D object comprises adding the distance information for the 3D object to one or more depth pixels in the depth graphics region 80.

[0114] In one embodiment, adjusting the depth graphics region 80 based on the distance information Dnand the scale Sninformation for the 3D object further comprises unsealing the one or more depth pixels in the depth graphics region.

[0115] In one embodiment, adding the distance information Dnand the scale Sninformation for the 3D object to one or more depth pixels and unsealing the one or more depth pixels in the depth graphics region 80 are performed according to:

[0116]

[0117] "

[0118] where:

[0119] "dxyis the original nthdepth pixel value;

[0120] d*yis the scaled nthdepth pixel value;

[0121] Snis a scale of the nth3D object; and

[0122] Dnis the distance of the nthdepth pixel from the camera that captured the image.

[0123] In one embodiment, rendering the 2D graphics region 70 and the adjusted depth graphics region 80 for output to a display comprises rendering 108 the 2D graphics region 70 and the adjusted depth graphics region 80 along with an ObjectID uniquely identifying the 3D object.In one embodiment, method 80 is implemented by a user device, such as HMD 20.

[0124] In at least some embodiments, the user device comprises an extended Reality, XR, device.

[0125] In one embodiment, the encoder comprises a 2D encoder and the decoder comprises a 2D decoder.

[0126] An apparatus can perform any of the methods herein described by implementing any functional means, modules, units, or circuitry. In one embodiment, for example, the apparatuses comprise respective circuits or circuitry configured to perform the steps shown in the method figures. The circuits or circuitry in this regard may comprise circuits dedicated to performing certain functional processing and / or one or more microprocessors in conjunction with memory. For instance, the circuitry may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include Digital Signal Processors (DSPs), special-purpose digital logic, and the like. The processing circuitry may be configured to execute program code stored in memory, which may include one or several types of memory such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, etc. Program code stored in memory may include program instructions for executing one or more telecommunications and / or data communications protocols as well as instructions for carrying out one or more of the techniques described herein, in several embodiments. In embodiments that employ memory, the memory stores program code that, when executed by the one or more processors, carries out the techniques described herein.

[0127] Figure 5 is a schematic block diagram illustrating some exemplary functional components of an encoder node 200, such as network node 16, for example, configured for depth encoding according to embodiments of the present disclosure. As seen in Figure 5, the encoder node 200 in this embodiment comprises, inter alia, processing circuitry 202 communicatively connected to communication circuitry 204 and memory 206.

[0128] In some embodiments, communication circuitry 204 comprises a network interface circuit 204a that facilitates communication with other network nodes, such as OA&M nodes, core network nodes, and / or other nodes in external systems. The network interface circuitry 204a in this regard may, for example, comprise an ETHERNET interface, an optical network interface, or in some situations, a wireless interface.

[0129] The processing circuitry 202 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of encoder node 200. The processing circuitry 202 in this regard can be configured by software to perform the functionality herein described, including the functionality described above with respect to method 30 shown in Figure 2.Memory 206 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 202 for operation. Memory 206 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. Memory 206 stores one or more computer programs 208 comprising executable instructions that configure the processing circuitry 202 in encoder node 200 to perform the functionality herein described, including method 30 shown in Figure 2. A computer program 208 in this regard may comprise one or more code modules corresponding to the means or units described above.

[0130] In general, computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM). In some embodiments, computer program 208 for configuring the processing circuitry 202 as herein described may be stored in a removable memory, such as a portable compact disc, portable digital video disc, or other removable media. The computer program 208 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0131] It should be noted here that the previously described embodiments illustrate the encoding functionality described herein on an encoder node 200 (e.g., network node 16). However, as previously stated, this is merely illustrative. In other embodiments, the encoding functionality described herein may be implemented on computing device 18, or on any computing device, such as the user’s laptop computer, notebook computer, desktop computer, local server computer, mobile communications device (e.g., a SMARTPHONE), tablet computer, and the like. Additionally, as with the network node 16, a computing device in these embodiments may, for example, comprise dedicated hardware circuitry, such as one or more graphics processing units (GPU), and be capable of executing games and other software programs associated with an XR environment.

[0132] Figure 6 is a schematic block diagram illustrating some exemplary functional components of a decoder node 300, such as HMD 20, configured for depth decoding according to embodiments of the present disclosure. As seen in Figure 6, decoder node 300 in this embodiment comprises, inter alia, processing circuitry 302, communication circuitry 304, and memory 306. In some embodiments, the communication circuitry 304 comprises both radio frequency (RF) circuitry 304a and network interface circuitry (NIC) 304b. In other embodiments, however, the network node may comprise only NIC 304b. More particularly, the RF circuitry 204a can be located at one or more Transmission Reception Points, TRPs, and comprises the RF components necessary for communicating with various network nodes andcomputer devices (e.g., Radio Access Network, RAN, nodes, network node 16, and computing device 18), directly or indirectly, via a wireless communication link. According to the present embodiments, the RF circuitry 304a may comprise, for example, a transmitter and receiver configured to operate according to the 5G standards or other wireless communication standard.

[0133] The communication circuitry 304 also comprises network interface circuitry (e.g., NIC 304b) for communication with RAN nodes, OA&M nodes, core network nodes, such as network node 16, and / or other nodes and computers, such as computing device 18, in external systems. The network interface circuitry 304b in this regard may, for example, comprise an ETHERNET interface, an optical network interface, or a wireless interface.

[0134] The processing circuitry 302 comprises one or more microprocessors, hardware, firmware, or a combination thereof that controls the overall operation of decoder node 300. The processing circuitry 302 in this regard can be configured by software to perform the functionality described with respect to one or more of the methods herein described, including method 90 as shown in Figure 4.

[0135] Memory 306 comprises both volatile and non-volatile memory for storing computer program code and data needed by the processing circuitry 302 for operation. Memory 306 may comprise any tangible, non-transitory computer-readable storage medium for storing data including electronic, magnetic, optical, electromagnetic, or semiconductor data storage. Memory 306 also stores one or more computer programs 308 comprising executable instructions that configure the processing circuitry 302 to perform the functionality described with respect to one or more of the methods herein described, including method 90 as shown in Figure 4. A computer program 308 in this regard may comprise one or more code modules corresponding to the means or units described above.

[0136] In general, computer program instructions and configuration information are stored in a non-volatile memory, such as a ROM, erasable programmable read only memory (EPROM) or flash memory. Temporary data generated during operation may be stored in a volatile memory, such as a random access memory (RAM). In some embodiments, computer program 308 for configuring the processing circuitry 302 as herein described may be stored in a removable memory, such as a portable compact disc or other removable media. The computer program 308 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0137] It should be noted here that the previously described embodiment illustrates the functionality described herein on HMD 20. However, as previously described, this is illustrative only. In some embodiments, the functionality described herein may be implemented on computing device 18, such as a user’s laptop computer, notebook computer, desktop computer,local server computer, mobile communications device (e.g., a SMARTPHONE), tablet computer, and the like. Regardless of the particular computer or type of computer, however, a computing device 18 in these embodiments may, for example, comprise dedicated hardware circuitry, such as one or more graphics processing units (GPU), and be capable of executing games and other software programs associated with an XR environment.

[0138] Those of ordinary skill in the art will appreciate that the embodiments herein further include corresponding computer programs. A computer program comprises instructions which, when executed on at least one processor of an apparatus, causes the apparatus to carry out any of the respective processing described above, including the functionality of methods 30 and 90 described in Figures 2 and 4, respectively. A computer program in this regard may comprise one or more code modules corresponding to the means or units described above.

[0139] Embodiments may also include a carrier containing such a computer program. This carrier may comprise one of an electronic signal, optical signal, radio signal, or computer readable storage medium.

[0140] In this regard, embodiments herein also include a computer program product stored on a non-transitory computer readable (storage or recording) medium and comprising instructions that, when executed by a processor of an apparatus, cause the apparatus to perform the functionality of methods 30 and 80 described in Figures 2 and 4, respectively, as described above.

[0141] Embodiments further include a computer program product comprising program code portions for performing the method of any of the embodiments herein when the computer program product is executed by a computing device. This computer program product may be stored on a computer readable recording medium.

[0142] The present embodiments may, of course, be carried out in other ways than those specifically set forth herein without departing from characteristics described herein. The present embodiments are therefore to be considered in all respects as illustrative and not restrictive, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced therein.

Claims

CLAIMSWhat is claimed is:

1. A method (30) for encoding depth into a video frame, the method comprising:for each of a plurality of 3D objects in an image:rendering (36) the 3D object to a 2D graphics region (70) and a depth graphics region (80);determining (42, 44) distance information and scale information for the 3D object in the depth graphics region; andadding the 2D graphics region and the depth graphics region to one or more composite frames;generating (52) a composite video frame comprising the one or more composite frames; encoding (54) the composite video frame and the distance information and the scale information for the 3D object to generate an encoded composite video frame; and sending (56) the encoded composite video frame to a decoder in a bitstream.

2. The method of claim 1, further comprising:allocating (34) an ObjectID to the 3D object uniquely identifying the 3D object; and associating the ObjectID with both the 2D graphics region and the depth graphics region.

3. The method of any of claims 1-2, wherein the 2D graphics region is processed independently of the depth graphics region.

4. The method of any of claims 1-3, wherein the 3D object in the depth graphics region comprises a plurality of depth pixels d, and wherein each depth pixel dnhas a respective depth value representing a distance Dnof the depth pixel dnfrom a camera that captured the image.

5. The method of claim 4, further comprising determining (42) a width of the 3D object as:Wnax _ minununwherein:is the highest depth pixel value and represents the point of the nth3D object that is furthest from the camera that captured the image; andis the lowest depth pixel value and represents the point of the nth3D object that is closest to the camera that captured the image.

6. The method of any of claims 1-4, wherein the distance information for the 3D object in the depth graphics region comprises ™°-7. The method of any of claims 1-6, wherein the scale information for the 3D object is determined as:wherein:Snis a scale of the nth3D object;the width of the 3D object; and2Nis a bit depth of the encoder; andfurther comprising scaling (46) one or more pixels in the depth graphics region according to the scale information.

8. The method of claim 7, wherein scaling the one or more pixels in the depth graphics region comprises mapping high-resolution depth pixels to low-resolution depth pixels according to:where:duxyn is a scaled nthdepth pixel value;duxnyis an original nthdepth pixel value;is the distance of the nthdepth pixel from the camera that captured the image; and scale of the nth 3D object.

9. The method of claim 8, wherein each high-resolution depth pixel is represented by a first resolution value, and wherein each low-resolution depth pixel is represented by a second resolution value that is less than the first resolution value.

10. The method of claim 9, wherein the first resolution value is a 32-bit value, and the second resolution value is an 8-bit value.

11. The method of any of claims 1-10, wherein generating a composite video frame comprising the one or more composite frames comprises assembling one or more composite frames into thecomposite video frame responsive to determining (50) that the distance information and the scale information have been determined for each of the 3D objects in the depth graphics region.

12. The method of any of claims 1-11, wherein encoding the composite video frame further comprises encoding the scale information and the distance information for the 3D object to generate the encoded composite video frame.

13. The method of any of the preceding claims, wherein the distance information and the scale information comprise metadata.

14. The method of any of the preceding claims, wherein each of the 2D graphics region and the depth graphics region are independently decodable.

15. The method of any of the preceding claims implemented by a user device (20).

16. The method of any of the preceding claims implemented by a network node (16).

17. A method (90) for decoding depth in a video frame, the method comprising:for each of a plurality of encoded composite video frames received in a bitstream:decoding (92) the encoded composite video frame to generate a composite video frame comprising:a 2D graphics region (70) of an image;a depth graphics region (80) of the image that is independent of the 2D graphics region; anddistance information and scale information for a 3D object in the image; and extracting (94) the 2D graphics region and the depth graphics region from the composite video frame; andadjusting (104, 106) the depth graphics region based on the distance information and the scale information for the 3D object; andrendering (98) the 2D graphics region and the adjusted depth graphics region for output to a display.

18. The method of claim 17, wherein adjusting the depth graphics region based on the distance information and the scale information for the 3D object comprises adding (106) the distance information for the 3D object to one or more depth pixels in the depth graphics region.

19. The method of claims 17-18, wherein adjusting the depth graphics region based on the distance information and the scale information for the 3D object further comprises unsealing (104) the one or more depth pixels in the depth graphics region.

20. The method of claims 18-19, wherein adding the distance information and the scale information for the 3D object to one or more depth pixels and unsealing the one or more depth pixels in the depth graphics region are performed according to:d nxv= d^ n '-S nn+ D nnwherein:is the original nthdepth pixel value;is the scaled nthdepth pixel value;Snis a scale of the nth3D object; andDnis the distance of the nthdepth pixel from the camera that captured the image.

21. The method of any of claims 17-20, wherein rendering the 2D graphics region and the adjusted depth graphics region for output to a display comprises rendering the 2D graphics region and the adjusted depth graphics region along with an ObjectID uniquely identifying the 3D object.

22. The method of any of claims 17-21 implemented by a user device (20).

23. The method of claim 22, wherein the user device comprises an extended Reality, XR, device.

24. The method of any of the preceding claims, wherein the encoder node comprises a 2D encoder and the decoder node comprises a 2D decoder.

25. An encoder node (200) for encoding depth into a video frame, the encoder node configured to:for each of a plurality of 3D objects in an image:render (36) the 3D object to a 2D graphics region (70) and a depth graphics region (80); determine (42, 44) distance information and scale information for the 3D object in the depth graphics region; andadd the 2D graphics region and the depth graphics region to one or more composite frames;generate (52) a composite video frame comprising the one or more composite frames; encode (54) the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame; and send (56) the encoded composite video frame to a decoder in a bitstream.

26. An encoder node (200) for encoding depth into a video frame, the encoder node comprising:communication circuitry (204); andprocessing circuitry (202) operatively connected to the communication circuitry and configured to:for each of a plurality of 3D objects in an image:render (36) the 3D object to a 2D graphics region (70) and a depth graphics region (80);determine (42, 44) distance information and scale information for the 3D object in the depth graphics region; andadd the 2D graphics region and the depth graphics region to one or more composite frames;generate (52) a composite video frame comprising the one or more composite frames; encode (54) the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame; and send (56) the encoded composite video frame to a decoder in a bitstream.

27. The encoder node of claim 26, wherein the processing circuitry is further configured to perform the method of any of claims 1-16 and 24.

28. A computer program (208), comprising instructions that, when executed by processing circuitry (202) of an encoder node, causes the processing circuitry to perform a method according to any of claims 1-16 and 24.

29. A non-transitory computer-readable medium (206) comprising a computer program (208) that, when executed by processing circuitry (202) of an encoder node, causes the processing circuitry to perform a method according to any of claims 1-16 and 24.

30. A carrier containing the computer program of claim 29, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer-readable medium.

31. A decoder node (300) for decoding depth in a video frame, the decoder node configured to:for each of a plurality of encoded composite video frames received in a bitstream:decode (92) the encoded composite video frame to generate a composite video frame comprising:a 2D graphics region of an image (70);a depth graphics region (80) of the image that is independent of the 2D graphics region;distance information for a 3D object in the image; andscale information for the 3D object;extract (94) the 2D graphics region and the depth graphics region from the composite video frame;adjust (104, 106) the depth graphics region based on the distance information and the scale information for the 3D object; andrender (108) the 2D graphics region and the adjusted depth graphics region for output to a display.

32. A decoder node (300) for decoding depth in a video frame, the decoder comprising:communication circuitry (304); andprocessing circuitry (302) operatively connected to the communication circuitry and configured to:for each of a plurality of encoded composite video frames received in a bitstream:decode (92) the encoded composite video frame to generate a composite video frame comprising:a 2D graphics region (60) of an image;a depth graphics region (80) of the image that is independent of the 2D graphics region;distance information for a 3D object in the image; andscale information for the 3D object;extract (94) the 2D graphics region and the depth graphics region from the composite video frame;adjust (104, 106) the depth graphics region based on the distance information and the scale information for the 3D object; andrender (108) the 2D graphics region and the adjusted depth graphics region for output to a display.

33. The decoder node of claim 32, wherein the processing circuitry is further configured to perform the method of any of claims 16-23.

34. A computer program (308), comprising instructions that, when executed by processing circuitry (302) of a decoder node (300), causes the processing circuitry to perform a method according to any of claims 16-23.

35. A non-transitory computer-readable medium (306) comprising a computer program (308) that, when executed by processing circuitry (302) of a decoder node (300), causes the processing circuitry to perform a method according to any of claims 17-24.

36. A carrier containing the computer program of claim 35, wherein the carrier is one of an electronic signal, optical signal, radio signal, or computer-readable medium.

37. A system (10) in a communication network for encoding and decoding depth in a video frame, the system comprising:an encoder node (200) and a decoder node (300), wherein:the encoder node is configured to:for each of a plurality of 3D objects in an image:render (36) the 3D object to a 2D graphics region (70) and a depth graphics region (80);determine (42, 44) distance information and scale information for the 3D object in the depth graphics region; andadd the 2D graphics region and the depth graphics region to one or more composite frames;generate (52) a composite video frame comprising the one or more composite frames; encode (54) the composite video frame, along with the distance information and the scale information for the 3D object, to generate an encoded composite video frame; and send (56) the encoded composite video frame to a decoder in a bitstream; andthe decoder node is configured to:for each of a plurality of encoded composite video frames received in the bitstream: decode (92) the encoded composite video frame to generate the composite video frame comprising:the 2D graphics region of an image;the depth graphics region of the image that is independent of the 2D graphics region;the distance information for the 3D object in the image; andthe scale information for the 3D object;extract (94) the 2D graphics region and the depth graphics region from the composite video frame;adjust (104, 106) the depth graphics region based on the distance information and the scale information for the 3D object; andrender (108) the 2D graphics region and the adjusted depth graphics region for output to a display.26