Deep single layer encoding / decoding

By encoding depth image sequences and representation information in a single-layer bitstream, the problem of the inability to separate and process depth information from video image texture information in existing technologies is solved. This enables independent encoding and decoding of depth information, supports multiple codecs and transmission methods, and improves the independence and flexibility of depth information.

CN122228658APending Publication Date: 2026-06-16BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-01-29
Publication Date
2026-06-16

Smart Images

  • Figure CN122228658A_ABST
    Figure CN122228658A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method (200) of decoding a sequence of depth images from a single-layer bitstream, wherein the method comprises decoding (210) a sequence of depth maps from the single-layer bitstream; and decoding (220) depth representation information from the single-layer bitstream, the depth representation information defining information for obtaining the sequence of depth images based on the decoded sequence of depth maps.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application is based on and claims priority to European Patent Application No. 23307069.7 filed on 28 November 2023 and European Patent Application No. 23307419.4 filed on 29 December 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure generally relates to the encoding and decoding of depth information in video images. Background Technology

[0003] Depth information in video images may follow different conventions, but it typically represents the distance from a specific point (often associated with certain types of sensor devices) along the direction given by the principal z-axis to point P (a point on the depth image plane IP) in the scene, such as... Figure 1 Shown (L. Xia, C. -C. Chen and JK Aggarwal, "Human detection using depth information by Kinect," CVPR 2011 WORKSHOPS, Colorado Springs, CO, USA, 2011, pp. 15-22, doi: 10.1109 / CVPRW.2011.5981811, https: / / cvrc.ece.utexas.edu / Publications / HAU3D11_Xia.pdf).

[0004] The process of sensing depth in a scene is called distance imaging. Distance imaging is a collective term for a series of techniques used to generate depth images (two-dimensional images) (or a sequence of depth images containing at least one depth image) that show the distances from a specific point in the scene to other points, typically associated with some type of sensor device. For example, this specific point might be the center of the sensor device's focal plane. The pixel values ​​of the generated depth image correspond to these distances.

[0005] If the sensor used to generate the distance image is properly calibrated, the depth image value can be given directly in physical units (such as meters).

[0006] Different sensors and technologies exist for generating depth image sequences by sensing the environment, such as stereo triangulation, light sheet triangulation, structured light, time-of-flight, interferometry, or coded aperture. Some imaging systems simultaneously possess depth and color sensors, which are different and spatially distant from each other, such as... Figure 2See (Azure Kinect DKcoordinate systems | Microsoft Docs, https: / / docs.microsoft.com / en-us / azure / kinect-dk / coordinate-systems).

[0007] Figure 3 Examples of content captured by a combination of depth and color sensors are given (Sun, Wenxiu, Lingfeng Xu, Oscar C. Au, Sung Him Chui, and Chun Wing Kwok. "An overview of free view-point depth-image-based rendering (DIBR)." In APSIPA Annual Summit and Conference (2010, pp. 1023-1030). 3D data representation of captured content can be achieved using a video plus depth format. Figure 3 The upper part represents the video content, and the lower part represents the minimum depth value Z. near and maximum depth value Z far The depth image values ​​between. Summary of the Invention

[0008] The following section provides a brief overview of at least one embodiment to provide a basic understanding of some aspects of this disclosure. This summary is not an exhaustive overview of exemplary embodiments. Its purpose is not to identify key or essential elements of the exemplary embodiments. The following summary presents only some aspects of at least one exemplary embodiment in a simplified form, serving as a prelude to a more detailed description provided elsewhere in this document.

[0009] According to a first aspect of this disclosure, a method for encoding a depth image sequence in a single-layer bitstream is provided, wherein the method includes: - Obtain depth representation information, which defines information used to obtain a depth map sequence based on a depth image sequence; - Obtain the depth map sequence by encoding the depth image sequence based on the depth representation information; - Encode the depth map sequence and the depth representation information in a single-layer bitstream.

[0010] In some embodiments, the method further includes encoding single-layer video data information in the bitstream, the single-layer video data information indicating that the bitstream is a single-layer bitstream including a depth map sequence.

[0011] According to a second aspect of this disclosure, a method for decoding a depth image sequence from a single-layer bitstream is provided, wherein the method includes: - Decode the depth map sequence from the single-layer bitstream; and - Decode depth representation information from the single-layer bitstream, the depth representation information defining information for obtaining a depth image sequence based on the decoded depth map sequence.

[0012] In some embodiments, the method further includes obtaining the depth image sequence from the decoded depth map sequence based on the decoded depth representation information.

[0013] In some embodiments, the method further includes decoding single-layer video data information from the single-layer bitstream, the single-layer video data information indicating that the bitstream is a single-layer bitstream including a depth map sequence.

[0014] In some embodiments, the depth representation information includes the minimum and maximum values ​​of the depth image sequence.

[0015] In some embodiments, single-layer video data information is a depth marker.

[0016] In some embodiments, the depth flag is signaled in a sequence-level parameter set.

[0017] In some embodiments, the depth flag is notified via a signal as a supplementary enhancement information message to the depth map sequence.

[0018] In some embodiments, the single-layer video data information is a specific value of the content type parameter.

[0019] In some embodiments, the content type parameter is signaled within a sequence-level parameter set.

[0020] In some embodiments, the content type parameter is notified via a signal as a supplementary information message to the content type information.

[0021] In some embodiments, the content type information supplementary enhancement message further includes a syntax element that specifies the mapping type between the depth map format and the video image format.

[0022] In some embodiments, the depth representation information is notified via a signal as a supplementary enhancement information message to the main depth representation information.

[0023] In some embodiments, single-layer video data information is signaled in a scalability dimension information supplementation and enhancement message.

[0024] In some embodiments, single-layer video data information is notified via signaling as a content type parameter in a scalability dimension scalability information supplementary enhancement message.

[0025] According to a third aspect of this disclosure, a bitstream formatted according to a method as described in a first or second aspect of this disclosure is provided.

[0026] According to a fourth aspect of this disclosure, a system is provided, the system including means for performing one of the methods described in the first or second aspect of this disclosure.

[0027] According to a fifth aspect of this disclosure, a computer program product is provided, the computer program product including instructions that, when the program is executed by one or more processors, cause the one or more processors to perform a method according to a first or second aspect of this disclosure.

[0028] The specific nature of at least one exemplary embodiment, as well as other objects, advantages, features, and uses of the at least one exemplary embodiment, will become more apparent from the following description of the examples in conjunction with the accompanying drawings.

[0029] Brief description of the attached figures Reference will now be made to the accompanying drawings, which illustrate exemplary embodiments of the present disclosure, wherein: Figure 1 An example of depth values ​​based on z-axis conventions according to existing technology is shown; Figure 2 Examples of depth and color cameras based on existing technology are shown; Figure 3 The RGBD format in a real capture sequence according to existing technology is shown; Figure 4 An extension of a 2D video codec for processing depth encoding and decoding according to the prior art is shown; Figure 5 This illustrates the definitions of different types of MV-HEVC layers according to existing technologies; Figure 6 Table G.3 of the HEVC standard according to the prior art is shown; Figure 7 Table F.1 of the 3D-HEVC standard according to the prior art is shown; Figure 8 Table 15 is shown according to the VVC standard in the prior art; Figure 9 A schematic block diagram of the steps of a method 100 for encoding a depth image sequence in a single-layer bitstream according to at least one exemplary embodiment of the present disclosure is shown. Figure 10A schematic block diagram of the steps of a method 200 for decoding a depth image sequence from a single-layer bitstream according to at least one exemplary embodiment of the present disclosure is shown. Figure 11 An example of a depth flag notified by signal in an SPS according to at least one exemplary embodiment of the present disclosure is shown; Figure 12 An example of a depth map sequence SEI message according to at least one exemplary embodiment of the present disclosure is shown; Figure 13 An example of a content type parameter notified by signal in an SPS according to at least one exemplary embodiment of the present disclosure is shown; Figure 14 An example of a Content Type Information (SEI) message according to at least one exemplary embodiment of this disclosure is shown; Figure 15 and Figure 16 An example of a Master Depth Representation Information (SEI) message according to at least one exemplary embodiment of this disclosure is shown; Figure 17 An example of a Content Type Information (SEI) message according to at least one exemplary embodiment of this disclosure is shown; Figure 18 An example of a Scalability Dimension Information (SDI) SEI message according to at least one exemplary embodiment of this disclosure is shown; Figure 19 An example of a Scalability Dimension Information (SDI) SEI message according to at least one exemplary embodiment of this disclosure is shown; Figure 20 A schematic block diagram of the steps of a method 300 for encoding video data associated with video images having a specific content type in a bitstream according to at least one exemplary embodiment of the present disclosure is shown. Figure 21 A schematic block diagram of the steps of a method 400 for decoding video data associated with video images having a specific content type from a bitstream according to at least one exemplary embodiment of the present disclosure is shown. Figure 22 A schematic block diagram is shown, illustrating an example of a system in which various aspects and exemplary embodiments are implemented.

[0030] Similar or identical elements are indicated by the same reference numerals.

[0031] Description of exemplary embodiments At least one exemplary embodiment will be described more fully below with reference to the accompanying drawings, which depict examples of at least one exemplary embodiment. However, exemplary embodiments may be embodied in various alternative forms and should not be construed as limited to the examples described herein. Therefore, it should be understood that this disclosure is not intended to limit exemplary embodiments to the specific forms disclosed. Rather, this disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of this disclosure.

[0032] An image can be a video frame belonging to a video, that is, a time sequence of video frames. There are temporal relationships between the video frames of a video.

[0033] An image can also be a still image.

[0034] An image includes at least one component (also called a channel) determined by a specific image / video format, which specifies all information related to the sample values ​​as well as all information used by the display unit and / or any other device to generate pixel values ​​for displaying and / or decoding image data associated with the image.

[0035] An image consists of at least one component and is typically represented as a two-dimensional array of samples.

[0036] A monochrome image consists of a single component, while a color image (also known as a texture image) can consist of three components.

[0037] For example, when the image / video format is the well-known (Y,Cb,Cr) format, a color image can include a luminance (or brightness) component and two chrominance components; while when the image / video format is the well-known (R,G,B) format, it can include three color components (one for red, one for green, and one for blue). The image / video format can also be the well-known (R,G,B,D) format (D represents depth information).

[0038] The image can also be an infrared image.

[0039] Each component of an image may include a number of samples related to the number of pixels on the display screen on which the image will be displayed. For example, the number of samples included in a component may be the same as, or a multiple (or fraction) of, the number of pixels on the display surface on which the image will be displayed.

[0040] The number of samples included in a component can also be a multiple (or fraction) of the number of samples included in another component of the same image.

[0041] For example, in image / video formats that include a luminance component and two chrominance components (such as the (Y,Cb,Cr) format), the chrominance component may contain half the number of samples relative to the luminance component in width and / or height, depending on the color format under consideration.

[0042] The texture information of an image can be defined as what is represented by a single component of a monochrome image, or what is represented by the luminance and chrominance components of a color image to be displayed to the user.

[0043] A sample is the smallest unit of visual information that makes up the components of an image. Sample values ​​can be, for example, luminance or chrominance values, or color values ​​of the red, green, or blue components in (R, G, B) format.

[0044] For monochrome images, the pixel value of the display surface can be represented by a single sample; while for color images, it can be represented by multiple co-location samples. A co-location sample associated with a pixel refers to the sample corresponding to the pixel's position on the display screen.

[0045] An image is typically viewed as a set of pixel values, with each pixel represented by at least one sample.

[0046] In addition, among other things, this disclosure also relates to systems and methods for encoding, decoding, acquiring, generating, processing and / or editing depth information of video images.

[0047] Depth representation information can define information used to obtain depth map values ​​based on depth image values, which represent the distance from the sensing device to a point in the scene, or conversely, to obtain depth image values ​​based on depth map values ​​that represent the distance from the sensing device to a point.

[0048] In some embodiments, depth representation information may include the minimum depth value Z of a depth image or a sequence of depth images. near and maximum depth value Z far .

[0049] In some embodiments, the depth representation information may also indicate whether the value of the depth map represents a depth image value (Z value) or the reciprocal of a depth image value.

[0050] A depth map sequence may include at least one depth map. The depth maps in a depth map sequence are consecutive depth maps in chronological order, typically in ascending order of time.

[0051] Figure 4This demonstrates the Universal Video Coding (VVC, ISO / IEC 23090-3 Versatile Video Coding (VVC) / ITU-T Rec. H.266) and the High Efficiency Video Coding (HEVC) standard (ISO / IEC 23008-2 High Efficiency Video Coding (HEVC) / ITU-T Rec. H.265, Chen, Ying, and Anthony Vetro. "Next-generation 3D formats with depth map support.") IEEE MultiMedia21 Various extensions of ( , no. 2 (2014): 90-94). These extensions handle depth encoding and decoding, and have been developed or are under development.

[0052] For example, one of these extensions is Multi-View Video Coding (MVC, ISO / IEC 14496-10, Annex H), a multi-view extension of Advanced Video Coding (AVC, ISO / IEC 14496-10 Advanced Video Coding (AVC) / ITU-T Rec. H.264), which has been developed for encoding two or more views. A key feature of this standard is maintaining AVC compatibility, meaning that one view in a compressed multi-view bitstream can be decoded by a conventional AVC decoder. Compression gain is achieved by allowing simultaneous instances of images from different views to predict other views, compared to independent encoding and decoding of all views. This process, called inter-view prediction, determines the block-based disparity offset between the reference view and the current view and uses it to perform disparity compensation prediction. This is similar to motion compensation prediction used in conventional video coding, but it is based on images with different viewpoints, rather than images from different temporal instances. In simple terms, the MVC approach is defined by two features: an extended high-level syntax to support appropriate signaling for view identifiers and their references, and a definition process through which the current image in another view can be predicted using the decoded image of another view.

[0053] For example, an extension of AVC called MVC+D (ISO / IEC 14496-10, Annex I) specifies a separate second stream for the representation of depth information and high-level syntactic signaling for the information needed to express the interpretation of depth information and its association with video data. The MVC+D method adds support for depth information using Network Abstraction Layer (NAL) units specifically designed for depth information and defines dedicated profiles for decoding multi-view plus depth video. This method does not involve macroblock-level changes to the AVC or MVC syntax, semantics, or decoding process. The corresponding 3D video codec is called MVC+D.

[0054] For example, an extension of the HEVC framework called Multi-View High-Efficiency Video Coding and Decoding (MV-HEVC, ISO / IEC 23008-2 Annex G) supports depth maps with auxiliary picture syntax. This auxiliary picture decoding process is the same for both video and multi-view video, but it doesn't necessarily have a prescriptive decoding requirement as part of a configuration file. This allows applications to optionally decode the depth map. If the industry desires an additional level of interoperability, a configuration file requiring the ability to decode depth information can be added at a later stage.

[0055] More precisely, MV-HEVC leverages the ability to encode different layers into the same bitstream. A layer is a set of VCL NAL units, all of which have a specific nuh_layer_id value and associated non-VCL NAL units, or one of a set of hierarchical syntax structures. Depending on the context, either a first-layer concept or a second-layer concept applies. The first-layer concept is also called a scalability layer, where a layer can be a spatially scalable layer, a quality-scalable layer, a view, etc. A temporally proper subset of a scalability layer is not called a layer, but rather a sublayer or temporal sublayer. The second-layer concept is also called a codec layer, where higher layers include lower layers. Codec layers are Encoded Video Sequences (CVS), pictures, slices, segments, and Codec Tree Unit (CTU) layers.

[0056] MV-HEVC design defines different types of layers: texture or depth information, multiple views, spatial / quality scalability, and auxiliary layers. The scalability mask index indicates the different types of layers, such as... Figure 5 As shown, scalability_mask_flag[i] equal to 1 indicates... Figure 5The dimension_id syntax element corresponding to the i-th scalability dimension exists. A scalability_mask_flag[i] equal to 0 indicates that the dimension_id syntax element corresponding to the i-th scalability dimension does not exist. In MV-HEVC, the dimension related to depth information is the auxiliary dimension of the scalability mask, which is index 3, i.e., auxiliary. By checking the AuxId value, it can be determined whether the current layer (lId) is an auxiliary image (value 0), an auxiliary α-plane (value 1), or an auxiliary depth image (value 2), as shown below. Then, if the layer has an auxiliary type of AUX_DEPTH, the layer also includes a Depth Representation Information (DRI) SEI (Supplemental Enhancement Information) message, the semantics of which are defined in specification G.14.3.3 "Semantics of Depth Representation Information SEI Messages". In the context of the depth map, this message allows signaling to Z. near and Z far Depth value (sometimes also called Z) near and Z far (Plane), and encodes and decodes depth values ​​(Z-values) or the reciprocal of the Z-value. These representation types are in Table G.3 of the HEVC standard ( Figure 6 Defined in (), and notified via a signal in the depth_representation_type parameter of the depth representation information SEI message.

[0057] Note that values ​​1 and 3 in Table G.3 define disparity measurements, i.e., the relative displacement of an object in two stereo images, which implies at least two views of the scene. This is irrelevant to this disclosure, as this disclosure focuses on single-view, single-layer, depth map sequences.

[0058] In MV-HEVC, auxiliary IDs can be used to notify the signaling layer that carries depth information.

[0059] For example, in an extension of AVC called 3D-AVC, the encoding and decoding of the depth view is similar to MVC+D, and no block-level changes for depth encoding and decoding are introduced. In 3D-AVC, the texture information of a single view is encoded and decoded in a manner compatible with AVC.

[0060] For example, in an extension of HEVC called 3D-HEVC, the multi-layer design of HEVC is utilized to encode both texture layers and depth layers into the same bitstream. A depth layer is a layer whose nuh_layer_id value equals i, thus DepthLayerFlag[i] equals 1, and DependencyId[i] and AuxId[i] equal 0. A texture layer is a layer whose nuh_layer_id value equals i, thus DepthLayerFlag[i], DependencyId[i], and AuxId[i] all equal 0. As seen in the definition, in 3D-HEVC, whether a layer is a texture layer or a depth layer depends on the value of its DepthLayerFlag. This variable is not present in the bitstream but can be derived from it from the bitstream using the following pseudocode sequence already seen in the MV-HEVC context:

[0061] In other words, if the value of dimension_id index 0 of the i-th layer is 1, then the layer is a depth layer. If the value of dimension_id index 0 of the i-th layer is 0, then the layer is a texture layer.

[0062] According to the definition of 3D-HEVC, depth layers and texture layers are not auxiliary layers because AuxId must be 0.

[0063] Furthermore, 3D-HEVC defines specific codecs for both texture and depth layers, making them incompatible with regular HEVC single-layer models. For example, an overview of these specific codecs can be found in Tech, Gerhard, YingChen, Karsten Müller, Jens-Rainer Ohm, Anthony Vetro, and Ye-Kui Wang's "Overview of the multiview and 3D extensions of high-efficiency video coding." IEEE Transactions on Circuits and Systems for Video Technology 26, no. 1(2015): 35-49.

[0064] To support these tools, the bitstream syntax is also consistent with that of Tech, Gerhard, Ying Chen, Karsten Müller, Jens-Rainer Ohm, Anthony Vetro, and Ye-Kui Wang. "Overview of the multiview and 3D extensions of high-efficiency video coding." IEEE Transactions on Circuits and Systems for Video Technology This differs from the conventional HEVC single-layer described in 26, no. 1 (2015):35-49.

[0065] In 3D-HEVC, a new scalable ID element called depth flags has been introduced. Unlike layers that indicate depth information via secondary IDs, layers with depth flags enabled can use the new 3D-HEVC codec tools.

[0066] VVC supports using depth information as an auxiliary layer (similar to MV-HEVC) and signaling it as a specific SEI in the VSEI (General Supplemental Enhancement Information Message for Encoding and Decoding Video Bitstreams / ITU-T Rec H.274). This specific SEI is called the Scalable Dimension Information (SDI) SEI. The Scalable Dimension Information (SDI) SEI message provides the SDI for each layer in the current CVS, i.e., the CVS including the SDI SEI message, for example: 1) the view ID of each layer when multiple views may exist; 2) the auxiliary ID of each layer when auxiliary information (such as depth information or α) may exist carried by one or more layers. When an SDI SEI message exists in any AU of the CVS, an SDI SEI message should also exist in the first AU of the CVS. All SDI SEI messages in the CVS should have the same content. Therefore, the (primary) video layer must exist, while depth information can only be an auxiliary layer, such as... Figure 8 Table 15 shows that sdi_aux_id[i] equal to 0 indicates that the i-th layer in the current CVS does not include auxiliary images; sdi_aux_id[i] greater than 0 indicates the type of auxiliary images in the i-th layer of the current CVS, such as... Figure 8 As shown in Table 15. When sdi_auxiliary_info_flag equals 0, the value of sdi_aux_id[i] is inferred to be equal to 0. Note that... Figure 8The interpretation of auxiliary pictures in Table 15 with sdi_aux_id[i] in the range of 128 to 159 (inclusive) is specified in a manner other than the value of sdi_aux_id[i]. For bitstreams conforming to this version of the specification, sdi_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive). Although for bitstreams conforming to this version of the specification, the value of sdi_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive), the decoder should also allow other values ​​of sdi_aux_id[i] in the range of 0 to 255 (inclusive).

[0067] If sdi_aux_id[i] equals 0, then the i-th layer is called the main layer. Otherwise, the i-th layer is called the auxiliary layer. When sdi_aux_id[i] equals 1, the i-th layer is also called the α-auxiliary layer. When sdi_aux_id[i] equals 2, the i-th layer is also called the deep auxiliary layer.

[0068] In existing technologies, the method of signaling depth information in video images exhibits some defects and shortcomings.

[0069] For example, using MV-HEVC, the depth information of the video images is carried as an auxiliary image to the main image (base layer). Therefore, without a base layer, the auxiliary layer with depth information cannot exist, and thus a single-layer bitstream independent of the base layer cannot be formed. Secondly, the depth representation information carried by the depth representation information (SEI) message must be part of the depth information type auxiliary layer (AUX_DEPTH). Therefore, this SEI message cannot be part of the base layer, because the base layer, by definition, is not an auxiliary layer.

[0070] For example, when using 3D-HEV, the depth information of video images is carried in a layer type called a depth layer, which has specific syntax elements and specific encoding / decoding tools, making the depth layer incompatible with the regular HEVC base layer.

[0071] The shortcomings of MV-HEVC described in this article also exist in the multi-layer support of VVC / H.266. The construction design principles of VVC / H.266 are similar to those of HEVC / H.265 and MV-HEVC.

[0072] There are multiple use cases where the depth information of video images must be processed, encoded, decoded, and / or transmitted together with or separately from the texture information of video images.

[0073] For example, Xia et al. (L. Xia, C.-C. Chen and JK Aggarwal, "Human detection using depth information by Kinect," CVPR 2011 WORKSHOPS, Colorado Springs,CO, USA, 2011, pp. 15-22, doi: 10.1109 / CVPRW.2011.5981811, https: / / cvrc.ece.utexas.edu / Publications / HAU3D11_Xia.pdf) proposed a method for human detection using depth information acquired by a sensing device. Their model-based method uses a 2D head contour model and a 3D head surface model to detect the human body. A segmentation scheme is also proposed to separate the human body from its surrounding environment and extract the overall contour of the person based on the detection points. The detection method described by Xia et al. only requires depth information from video images.

[0074] Choi et al. (B. Choi, Ç. Meriçli, J. Biswas and M. Veloso, "Fast human detection for indoor mobile robots using depth images," 2013 IEEE International Conference on Robotics and Automation, Karlsruhe, Germany, 2013, pp. 1108-1113, doi: 10.1109 / ICRA.2013.6630711., https: / / sci-hub.se / 10.1109 / icra.2013.6630711) proposed a "fast human detection algorithm for mobile robots equipped with depth cameras." This method uses a graph-based segmentation algorithm to segment depth images and applies a parameterized set of heuristics to filter and merge segmented regions to obtain a candidate region set. Finally, the method computes an oriented depth histogram (HOD) descriptor for each candidate region and uses a linear SVM to test for human presence. The detection method described by Choi et al. only requires depth information from video images.

[0075] Yan et al. (Yan, S., Yang, J., Leonardis, A. and Kamarainen, JK, 2021.Depth-only object tracking. arXiv preprint arXiv:2110.11679, https: / / arxiv.org / pdf / 2110.11679.pdf) proposed a deep learning approach for object tracking. The architecture is derived from a previous model that worked with an RGB-only tracker, namely DiMP. Yan et al. introduced a variant using a depth-only tracker, called Depth-DiMP, and its extension RGBD-DiMP (which additionally considers RGB data), and evaluated the performance of all variants (DiMP, Depth-DiMP, and RGBD-DiMP), which share the same architecture: a pre-trained feature extraction module, an offline-trained object estimation module, and an online-trained classifier module. The deep learning approach described by Yan et al. can require both depth and texture information from video images, but can also require only depth information, as their results show that the depth-only tracker outperforms state-of-the-art RGB trackers.

[0076] Pulli et al. (Pulli, K. and Pietikäinen, M., 1993, May. Range imagesegmentation based on decomposition of surface normals. In Proceedings of the Scandinavian conference on image analysis (Vol. 2, pp. 893-893). Proceedings published by various publishers, https: / / people.csail.mit.edu / kapu / papers / SCIApap.pdf) proposed a depth map-based segmentation method. Specifically, this segmentation method is based on a local approximation of normal vectors, which segments the distance image into uniform surface regions. The segmentation method described by Pulli et al. can be implemented using only depth information from video images.

[0077] Shotton et al. (J. Shotton et al., "Real-time human pose recognition inparts from single depth images," CVPR 2011, Colorado Springs, CO, USA, 2011, pp. 1297-1304, doi: 10.1109 / CVPR.2011.5995316, https: / / www.microsoft.com / en-us / research / wp content / uploads / 2016 / 02 / BodyPartRecognition.pdf) proposed a method for fast and accurate prediction of the 3D position of human joints from a single depth image without using temporal information. The method described by Shotton et al. can be implemented using only depth information from video images.

[0078] Li et al. (Li, Z. and Stamos, I., 2023. Depth-based 6DoF Object PoseEstimation using Swin Transformer. arXiv preprint arXiv:2303.02133, https: / / arxiv.org / pdf / 2303.02133.pdf) proposed a deep learning-based method for inferring the 6DoF pose of an object from a depth image. The method computes the angles between surface normal vectors and the three coordinate axes. These angles are then normalized to the RGB color range to form an RGB-like image where each pixel has three normalized angle values. This image is fed into an image representation learning network. In addition to the normal vector angle image, the method also upscales the depth image to a point cloud using given camera parameters and extracts point cloud features through a point cloud representation learning network. Combining these two embeddings, our SwinDePose architecture can leverage both depth image and point cloud information for more accurate 6D pose estimation. The method described by Li et al. may require both depth and texture information from video images.

[0079] Newcombe et al. (RA Newcombe et al., "KinectFusion: Real-time dense surface mapping and tracking," 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland, 2011, pp. 127-136, doi:10.1109 / ISMAR.2011.6092378, https: / / www.microsoft.com / en-us / research / wp-content / uploads / 2016 / 02 / ismar2011.pdf) proposed a method for real-time reconstruction and mapping of indoor targets under variable lighting conditions, using only a mobile, low-cost depth camera and general-purpose graphics hardware. All depth information flowing from the Kinect sensor is fused in real-time into a single global implicit surface model of the observed scene. The pose of the current sensor is simultaneously acquired by tracking the real-time depth frame relative to the global model using a coarse-to-fine Iterative Closest Point (ICP) algorithm, which utilizes all available observation depth information. The method described by Newcombe et al. can require only depth information from video images.

[0080] There are other use cases where the depth information of video images must be processed, encoded, decoded, and / or transmitted together with or separately from the texture information of video images, such as in fields like object recognition (https: / / link.springer.com / content / pdf / 10.1007 / 978-3-642-37444-9_41.pdf). See https: / / sci-hub.se / 10.1109 / icarsc.2016.35 or https: / / link.springer.com / content / pdf / 10.1007 / 978-3-642-37444-9_41.pdf https: / / sci-hub.se / 10.1109 / icarsc.2016.35), Robot / unmanned vehiclenavigation (https: / / ieeexplore.ieee.org / document / 6224766 or https: / / ieeexplore.ieee.org / document / 6224766), scene completion from single depthimage (https: / / www.ijcai.org / proceedings / 2018 / 0101.pdf or (https: / / www.ijcai.org / proceedings / 2018 / 0101.pdf).

[0081] In existing technologies, encoding / decoding and transmitting depth map sequences of video images requires encoding / decoding and transmitting the texture information of the video images, as the depth map sequence is considered auxiliary information for the video images. Therefore, it is impossible to implement the encoding / decoding and transmission of video image depth map sequences within any single-layer video codec, such as AVC, EVC, or even AV1 (AOMedia Video 1) from the AOM (Open Media Consortium), because this would result in the loss of depth-related information and prevent the receiver from interpreting and retrieving the depth image from the video bitstream.

[0082] Furthermore, some use cases, as described below, only require the encoding / decoding / transmission of depth images / graph sequences, while existing technologies do not provide a method for encoding / decoding / transmitting only the depth images / graph sequences of video images without encoding / decoding / transmitting the texture information of the video images.

[0083] In view of the above, at least one exemplary embodiment of this disclosure has been designed.

[0084] Embodiments of this disclosure enable the encoding of depth images / image sequences and their associated depth representation information into a single-layer bitstream, referred to as a depth bitstream. This depth bitstream does not include texture information of the video images and can be encoded / decoded / transmitted independently of the bitstream carrying the texture information of the video images (referred to as a texture bitstream). The depth bitstream can be decoded by conventional existing 2D video decoders (without requiring specific encoding / decoding tools).

[0085] Embodiments of this disclosure allow encoding depth image sequences within a single-layer bitstream and embedding sufficient metadata (depth representation information). This enables a receiver to decode the single-layer bitstream and process the depth representation information using a conventional 2D decoder, thereby correctly interpreting decoded samples within the decoded depth map sequence. Both the encoding and transmission of the depth map sequence can be optimized, regardless of the presence of associated texture information sequences to be encoded. Compared to MV-HEVC, 3D-HEVC, and multi-layer VVC, embodiments of this disclosure require less processing because only the depth bitstream needs to be encoded and transmitted.

[0086] For use cases where the sender only needs to transmit the depth information of the video (rather than the texture information of the video), this allows the transmission over the network of depth images / graph sequences and / or their storage in binary file form, such as the ISOBMFF format as defined in the ISO / IEC 14496-12 standard, just like any other regular 2D video bitstream. However, other file instances may also be used without being limited to the scope of this disclosure.

[0087] For use cases where the sender needs to transmit both texture and depth information of a video image, embodiments of this disclosure allow the texture information to be encoded and transmitted in parallel using a multi-stream method (texture bitstream and depth bitstream). This multi-stream method allows: - Choose two different codecs for texture and depth information (e.g., AVC and HEVC), or choose the same codec but with different profiles / levels; - Set different encoding qualities between texture information and depth information; - Set different picture group encoding levels in each bitstream, i.e., different positions of reference frames (I-frames and P-frames); - Set different quality of service for network streams of texture bitstream and depth bitstream, for example, texture bitstream can be considered to have a higher priority than depth bitstream, or vice versa; - Set different retransmission strategies for each texture bitstream and depth bitstream. For example, if one texture bitstream or the other is deemed more important than the other, the more important bitstream can be sent when retransmitting lost network packets, while the other is not sent. Lost packets will cause decoding artifacts, but for less important bitstreams, this may be considered acceptable depending on the application. - Use different transport protocols for each texture and depth stream, that is, use completely different protocols. For example, one bit stream will be sent via a TCP-based network protocol, while another bit stream will be sent via a UDP-based network protocol. - Easily reuse texture bitstreams and depth bitstreams using multiplexed network protocols, for example, each bitstream can be mapped to a different QUIC stream (RFC 9000: QUIC: A UDP-Based Multiplexed and SecureTransport (rfc-editor.org), https: / / www.rfc-editor.org / rfc / rfc9000.html).

[0088] Embodiments of this disclosure relate to a method for encoding a depth image sequence in a single-layer bitstream. The method obtains depth representation information, which defines information for obtaining a depth map sequence based on the depth image sequence; obtains a depth map sequence by encoding the depth image sequence based on the depth representation information; and encodes the depth map and depth representation information in a single-layer bitstream.

[0089] Embodiments of this disclosure also relate to a method for decoding a depth image sequence from a single-layer bitstream. The method decodes a depth map sequence and depth representation information from the single-layer bitstream, the depth representation information defining information used to obtain a depth image sequence based on the depth map sequence.

[0090] Embodiments of this disclosure also relate to a method for encoding video data having a specific content type in a bitstream, the method comprising encoding single-layer video data information in the bitstream indicating the content type of the video data; and encoding video data in the bitstream based on the content type indicated by the encoded single-layer video data information.

[0091] Embodiments of this disclosure also relate to a method for decoding video data having a specific content type from a bitstream, the method comprising decoding single-layer video data information indicating the content type of the video data from the bitstream; and decoding the video data from the bitstream based on the content type indicated by the decoded single-layer video data information.

[0092] The embodiments of this disclosure are generally described in the context of encoding / decoding systems or devices that acquire or obtain depth information from video images, and / or processes or methods for acquiring, generating, processing, and / or editing such depth information. It should be understood that this disclosure is also applicable to other systems, devices, processes, and / or methods for acquiring, generating, processing, and / or editing depth information. Depth information acquisition systems can be systems / devices designed for film professionals, including full focus control after video capture, and / or systems / devices designed for non-professionals, including, for example, digital SLR cameras or consumer video capture systems aimed at high-end consumers, which perform automatic or semi-automatic focus adjustment control and circuitry during image / video acquisition.

[0093] For example, embodiments of this disclosure may be implemented in conjunction with video capture devices (such as cameras) and / or systems to generate, process, and / or edit depth information.

[0094] Figure 9 A schematic block diagram illustrating the steps of a method 100 for encoding a depth image sequence in a single-layer bitstream according to at least one exemplary embodiment of the present disclosure is shown. The depth image sequence includes at least one depth image, and each depth image value may represent the distance from a sensing device to a point in the scene.

[0095] In step 110, depth representation information is obtained. This depth representation information defines the information used to obtain a depth map sequence based on the depth image sequence.

[0096] In step 120, a depth map sequence is obtained by encoding the depth image sequence based on depth representation information.

[0097] Depth representation information can indicate the inverse of the distance used, causing the encoded / decoded value (depth map value) to decrease and rapidly approach 0 as the target's distance from the camera increases. This maintains higher precision for nearby targets compared to distant ones, consistent with the human visual system's ability to sense depth for near targets better than for distant ones (because the parallax effect, a crucial cue for depth perception, decreases with increasing distance). Furthermore, when using linear quantization of the inverse of the depth image value, targets closer to the camera are described with higher granularity, while targets farther away are described with lower quantization levels. The importance of depth granularity decreases at greater distances because the human eye perceives depth less effectively for the reasons mentioned above.

[0098] In step 130, the depth map sequence and depth representation information are encoded into a single-layer bitstream.

[0099] In some embodiments, method 100 further includes encoding (step 140) single-layer video data information in a single-layer bitstream, the single-layer video data information indicating that the single-layer bitstream is a single-layer bitstream including a depth map sequence.

[0100] Figure 10 A schematic block diagram illustrating the steps of a method 200 for decoding a depth image sequence from a single-layer bitstream according to at least one exemplary embodiment of the present disclosure is shown. The single-layer bitstream can be... Figure 9 Method 100 was obtained.

[0101] The depth image sequence includes at least one depth image, and each depth image value can represent the distance from the sensing device to a point in the scene.

[0102] In step 210, the depth map sequence is decoded from the single-layer bitstream.

[0103] In step 220, depth representation information is decoded from a single-layer bitstream. This depth representation information defines information used to obtain a depth image sequence based on the decoded depth map sequence.

[0104] In some embodiments, method 200 further includes obtaining (step 230) a depth image sequence from a decoded depth map sequence based on decoded depth representation information. In some embodiments, method 200 further includes decoding (step 250) single-layer video data information from a single-layer bitstream, the single-layer video data information indicating that the bitstream is a single-layer bitstream including a depth map sequence.

[0105] Figure 20 A schematic block diagram illustrating the steps of a method 300 for encoding video data associated with video images having a specific content type in a bitstream according to at least one exemplary embodiment of the present disclosure.

[0106] In step 310, single-layer video data information indicating the content type of the video data is encoded into a bitstream.

[0107] In step 320, video data is encoded into a bitstream based on the content type indicated by the encoded single-layer video data information.

[0108] Figure 21 A schematic block diagram illustrating the steps of a method 400 for decoding video data associated with video images having a specific content type from a bitstream according to at least one exemplary embodiment of the present disclosure.

[0109] In step 410, single-layer video data information indicating the content type of the video data is decoded from the bitstream.

[0110] In step 420, video data is decoded from the bitstream based on the content type indicated by the decoded single-layer video data information.

[0111] The embodiments of this disclosure are described in detail below as extensions to the HEVC or VVC standards. However, this disclosure is not limited to these extensions, and can also be implemented as extensions to other standards.

[0112] In some embodiments, single-layer video data information is a depth flag, i.e., a Boolean value.

[0113] This depth flag indicates whether the base layer in an HEVC-compatible bitstream is a depth bitstream.

[0114] In some embodiments, since each layer has at most one Sequence Parameter Set (SPS) active at any given time, the depth flag can be signaled in the sequence-level SPS. However, this approach has the disadvantage of being backward incompatible with existing HEVC-compliant bitstreams, encoders, and decoders because it requires changes to the syntax of the SPS data structure.

[0115] Figure 11 An example of a depth flag notified by a signal in SPS according to an embodiment of the present disclosure is shown.

[0116] A depth flag `sps_depth_sequence_flag` equal to 0 indicates that each CVS referencing an SPS does not represent a depth map sequence associated with a video image. When `sps_depth_sequence_flag` equals 1, it indicates that each CVS referencing an SPS represents a depth map sequence associated with a video image.

[0117] In some embodiments, the depth flag can be signaled as a depth map sequence supplemental enhancement information (SEI) message.

[0118] Compared to previous embodiments, this embodiment has the advantage of backward compatibility with existing HEVC-compliant bitstreams, encoders, and decoders, because traditional decoders can ignore SEI messages, while new decoders can use these messages. SEI messages also mean that new data structures can be registered by defining new payload types.

[0119] The Depth Map Sequence (SEI) message specifies the depth map sequence representing the current layer.

[0120] When a depth map sequence SEI message exists, the message should be associated with the i-th layer, such that DepthLayerFlag[i], DependencyId[i], and AuxId[i] are equal to 0. The following semantics apply to each nuh_layer_id targetLayerId in the nuh_layer_id value applied to the depth map sequence SEI message. The conditions of DepthLayerFlag[i], DependencyId[i], and AuxId[i] mean that the layer is a regular base layer (DependencyId = 0), not a depth layer in the 3D-HEVC sense (DepthLayerFlag[i] = 0), nor an auxiliary layer in the MV-HEVC sense (AuxId[i] = 0).

[0121] When a depth map sequence SEI message is present, the message can be included in any access unit. Preferably, when a depth map sequence SEI message is present, the message is included in the access unit for random access purposes, where the encoded / decoded image with nuh_layer_id equal to targetLayerId is an IRAP image.

[0122] Preferably, the depth map sequence SEI message can be inserted into the bitstream at the point where the decoder begins decoding. This is a common case for several existing SEI messages.

[0123] The depth map sequence SEI message is applied to all video images, which have a nuh_layer_id equal to the targetLayerId, starting from the access unit that includes the depth map sequence SEI message, upwards to, but not including, the next video image associated with the depth map sequence SEI message that can be applied to the targetLayerId in decoding order, or up to the end of the CLVS where nuh_layer_id equals the targetLayerId, whichever is earlier in decoding order.

[0124] The depth map sequence SEI message can then be applied until a new depth map sequence SEI message of the same type appears in the bitstream, until the layer ends. It is highly likely that the properties of the layer (depth or non-depth) will not change over the duration of the layer. However, additional information can be inserted into the depth map sequence SEI message that will cause it to change over the duration of the layer.

[0125] Figure 12 An example of a depth map sequence (SEI) message according to at least one exemplary embodiment of this disclosure is shown. A depth flag dms_depth_flag equal to 0 indicates that each CVS referencing an SPS does not represent a depth map sequence. When sps_depth_sequence_flag equals 1, it indicates that each CVS referencing an SPS represents a depth map sequence.

[0126] In some embodiments, single-layer video data information may be a specific value of the content type parameter.

[0127] In some embodiments, since each layer has at most one Sequence Parameter Set (SPS) active at any given time, the content type parameter can be signaled in the sequence-level SPS, which is a good place to signal the general content type parameter in the SPS. However, a drawback of this embodiment is that it is not backward compatible with existing HEVC-compliant bitstreams, encoders, and decoders.

[0128] Figure 13 An example of a content type parameter notified by signal in SPS according to an embodiment of this disclosure is shown.

[0129] The syntax element `sps_content_type_idc` specifies the content type represented by each CVS referencing this SPS. The value of `sps_content_type_idc` should be in the range of 0 to 15 (inclusive). A value of 0 indicates that the layer is a depth layer, meaning that the single-layer bitstream includes a sequence of depth maps.

[0130] Table 1 provides examples of content types notified by the syntax element sps_content_type_idc signal.

[0131] Table 1

[0132] In some embodiments, the content type parameter is a bit word.

[0133] In the example in Table 1, the content type parameter is a 4-bit word, but the content type parameter can use more bits, such as 8 bits, to allow for more possible content types.

[0134] The list of content types in Table 1 is not exhaustive; other content types can be added.

[0135] The brief information for each content type listed in Table 1 is as follows: -Depth map sequence: Depth maps and depth representation information associated with video images; - Parallax sequence: A sequence of parallax images that describes the relative distance between the same object visible in the left and right views of a stereoscopic view; -α channel sequence: Represents the sequence of α channels used for α mixing, equivalent to the AUX_ALPHA type in MV-HEVC; - Infrared sequence: A sequence of images recorded by an infrared camera sensor. These images are typically single-channel; - Target mask sequence: includes a sequence of masks describing the targets detected in the video image; - Semantic mask sequence: A sequence representing the results of scene semantic analysis, where each image is a semantic mask; - Transparency Mask Sequence: Each image in this sequence is a transparency mask, which is used in conjunction with the associated video to be displayed on a transparent display; - Confidence Image Sequence: A sequence of images representing confidence levels associated with the relevant process.

[0136] In some embodiments, the content type parameter is notified via a signal as a content type information (SEI) message.

[0137] Compared to previous embodiments, this embodiment has the advantage of backward compatibility with existing HEVC-compliant bitstreams, encoders, and decoders, because traditional decoders can ignore SEI messages, while new decoders can use these messages. SEI messages also mean that new data structures can be registered by defining new payload types.

[0138] The Content Type Information (SEI) message specifies the content type represented in the current layer. Specifically, the SEI message specifies that the current layer represents a depth map sequence.

[0139] When a Content Type Information (SEI) message exists, it can be included in any access unit. Preferably, when an SEI message exists, it is included in the access unit for random access purposes, where the encoded / decoded image with nuh_layer_id equal to targetLayerId is an IRAP image.

[0140] Preferably, the Content Type Information (SEI) message can be inserted into the bitstream at the point where the decoder begins decoding. This is a common case for several existing SEI messages.

[0141] The Content Type Information (SEI) message is applied to all video images, which have a nuh_layer_id equal to the targetLayerId, starting from the access unit that includes the SEI message and proceeding upwards until, but not including, the next video image associated with the SEI message that can be applied to the targetLayerId in the decoding order, or until the end of the CLVS where nuh_layer_id equals the targetLayerId, whichever is earlier in the decoding order.

[0142] The Content Type Information (SEI) message can then be applied until a new message of the same type appears in the bitstream, until the layer ends. It's highly likely that the properties of the layer (deep or non-deep) will not change over the layer's duration. However, additional information can be inserted into the SEI message that will cause it to change over the layer's duration.

[0143] Figure 14 An example of a Content Type Information (SEI) message according to at least one exemplary embodiment of this disclosure is shown.

[0144] The syntax element cti_content_type_idc specifies the content type represented by the current layer in the CVS that references this SPS. The value of cti_content_type_idc should be in the range of 0 to 15 (inclusive).

[0145] Table 2 provides examples of content types notified by the syntax element cti_content_type_idc signal.

[0146] Table 2

[0147] In some embodiments, depth representation information can be signaled as a primary depth representation information (SEI) message (output / obtained from / from the depth bitstream (method 100) or bitstream (method 200).

[0148] The Master Depth Representation Information (SEI) message is inspired by the Depth Representation Information (SEI) message, but it is redefined for the background of depth map sequences in a single-layer bitstream. Compared to existing Depth Representation Information (SEI) messages, parallax is no longer reused, as it is only meaningful in a multi-view context of the same scene.

[0149] Figure 15 and Figure 16 An example of a Master Depth Representation Information (SEI) message according to at least one exemplary embodiment of this disclosure is shown.

[0150] The syntax elements in the Master Depth Representation Information (SEI) message specify various parameters representing the master layer of the depth map sequence.

[0151] When the primary depth representation information (SEI) message exists, it should be associated with one or more layers whose AuxId value is equal to 0 and whose sps_content_type_idc value is equal to 0.

[0152] Here, the definition of sps_content_type_idc is the same as... Figure 13 The definition is the same as in [the previous definition]. Alternatively, this condition can also be expressed by changing the depth flag.

[0153] When the Master Depth Representation Information (SEI) message is present, it can be included in any access unit. Preferably, when the SEI message is present, it is included in the access unit for random access purposes, where the encoded / decoded image with nuh_layer_id equal to targetLayerId is an IRAP image.

[0154] The Master Depth Representation Information (SEI) message is applied to all images, which has a nuh_layer_id equal to the targetLayerId, starting from the access unit that includes the SEI message and proceeding upwards until, but not including, the next video image associated with the SEI message that can be applied to the targetLayerId in decoding order, or until the end of the CLVS where nuh_layer_id equals the targetLayerId, whichever is earlier in decoding order.

[0155] For the sake of brevity, the syntax elements used in HEVC's deep representation information messages will not be repeated here.

[0156] Table 3 provides examples of depth representation information notified by the syntax element sps_content_type_idc signal.

[0157] Table 3

[0158] In Table 3, the primary depth representation information message is defined for use with regular images rather than auxiliary images. Additionally, a new depth representation type is defined: the Euclidean distance from the camera center to a distant target. In other cases, the Z-value represents the projected distance on the Z-axis as defined in the introduction of this application.

[0159] When the depth map is encoded and decoded using a multi-channel image format (such as YCbCr 4:2:0), Table 3 should be further adjusted. In this case, the "each decoding depth value" calculated based on the decoded luminance sample value and two chrominance sample values ​​should be used as the decoding depth value, rather than the luminance sample value.

[0160] Since a depth map is by definition a single-channel two-dimensional array, the internal format of the video images being encoded and decoded in the texture bitstream may not be single-channel, i.e., it may only contain luminance samples. When the image format being encoded and decoded is multi-channel, it is necessary to specify how to map between the reconstructed depth map (single-channel) and the decoded video images, or between the original depth map (single-channel) and the video images to be encoded and decoded.

[0161] In some embodiments, such as Figure 16 As shown, Figure 14 The content type information SEI message may also include the syntax element cti_channel_mapping_idc, which specifies the mapping type between the video image format and the depth map image format used in the current layer.

[0162] Table 4 provides examples of values ​​for the syntax element cti_channel_mapping_idc, which should be in the range of 0 to 255 (inclusive).

[0163] Table 4

[0164] Table 4 provides 256 values ​​because many different variants can be defined.

[0165] VVC supports depth map sequences as auxiliary layers and implements signal notification in the form of SEI messages within VSEI. There are two specific methods to extend VSEI to support single-layer CVS in VVC: one is to extend Scalable Dimension Information (SDI) to include depth markers; the other is to extend Scalable Dimension Information (SDI) to include content types.

[0166] In some embodiments, such as Figure 18 As shown, single-layer video data information can be signaled as a Scalable Dimension Information (SDI) SEI message, which, as defined in VVC, can be extended to include a depth flag.

[0167] The syntax element `sdi_depth_sequence_flag[i]` equal to 1 indicates that the i-th layer represents a depth map sequence as the main layer. Therefore, `sdi_depth_sequence_flag[i]` can only exist if `sdi_aux_id[i]` equals 0. When `sdi_depth_sequence_flag[i]` equals 0, it means that the i-th layer does not represent a depth map sequence. When `sdi_depth_sequence_flag[i]` does not exist, its value is inferred to be 0.

[0168] The syntax element sdi_aux_id[i] equal to 0 indicates that the i-th layer in the current CVS does not include auxiliary images. sdi_aux_id[i] greater than 0 indicates the type of auxiliary image in the i-th layer of the current CVS, as specified in Table 15. When sdi_auxiliary_info_flag equals 0, it is inferred that the value of sdi_aux_id[i] is equal to 0.

[0169] Table 5 provides a mapping between sdi_aux_id[i] and auxiliary image types.

[0170] Table 5

[0171] The interpretation of the auxiliary picture associated with sdi_aux_id[i] (in the range of 128 to 159, inclusive) is specified in a way other than the value of sdi_aux_id[i].

[0172] For bitstreams conforming to this version of the specification, sdi_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive). Although the value of sdi_aux_id[i] should be in the range of 0 to 2 (inclusive) or 128 to 159 (inclusive) in this version of the specification, the decoder should also allow other values ​​of sdi_aux_id[i] in the range of 0 to 255 (inclusive).

[0173] If sdi_aux_id[i] equals 0, then the i-th layer is called the main layer. Otherwise, the i-th layer is called the auxiliary layer. When sdi_aux_id[i] equals 1, the i-th layer is also called the α-auxiliary layer. When sdi_aux_id[i] equals 2, the i-th layer is also called the deep auxiliary layer.

[0174] In some embodiments, such as Figure 19 As shown, single-layer video data information can be signaled as a Scalable Dimension Information (SDI) SEI message, which, as defined in the VVC standard, can be expanded to include content type parameters.

[0175] The syntax element sdi_aux_id[i] equal to 0 indicates that the i-th layer in the current CVS does not include auxiliary images. sdi_aux_id[i] greater than 0 indicates the type of auxiliary image in the i-th layer of the current CVS, as specified in Table 15. When sdi_auxiliary_info_flag equals 0, it is inferred that the value of sdi_aux_id[i] is equal to 0.

[0176] Table 5 provides a mapping between sdi_aux_id[i] and auxiliary image types.

[0177] The syntax element sdi_content_type_info_flag equals 1, indicating that one or more layers in the current CVS can be layers carrying specific content, as defined by the semantics of sdi_content_type_idc[i].

[0178] The syntax element `sdi_content_type_idc[i]` specifies the content type of the i-th level in the current CVS. The value of `sdi_content_type_idc[i]` should be in the range of 0 to 15 (inclusive). If `sdi_content_type_idc[i]` does not exist, its value is not specified.

[0179] Table 6 provides examples of values ​​for the syntax element sdi_content_type_idc[i].

[0180] Table 6

[0181] In some embodiments, since SPS also exists in VVC, a similar approach to that in HEVC can be used, where the content type parameter is notified via a signal in the sequence-level SPS. Only the location of the information will differ to suit VVC's SPS syntax.

[0182] In some embodiments, single-layer video data information can be signaled as a content type parameter in the Scalability Dimension Information (SEI) message. VVC uses the same SEI message system as HEVC. Therefore, the content type information (SEI) message for VVC can be received from... Figure 14 The content type information (SEI) used for HEVC is derived from the SEI message shown.

[0183] When the internal format of the encoded video image is not a single-channel image but a multi-channel image, a conversion between the depth map format and the video image format is required. In practical applications, if the receiving device does not support single-channel video decoding, conversion can be performed. At the time of writing, hardware decoders typically do not support single-channel video / images, thus requiring conversion.

[0184] In some embodiments, the conversion can use dummy values ​​of the luminance (Y) and chrominance channels (e.g., Cb and Cr) of a multi-channel video image.

[0185] If m(i,j) is a sample of the depth map, then the samples y(i,j), cb(i,j), and cr(i,j) of the transformed multi-channel image can be defined as follows:

[0186] This conversion is used for the HEVC / H.265 video codec standard (and may also be used for other specifications): When AuxId[lId] equals AUX_ALPHA or AUX_DEPTH, one of the following applies: - In the active SPS of the layer where nuh_layer_id is equal to lId, chroma_format_idc is equal to 0; - In all video images where nuh_layer_id equals lId and the VPS RBSP is an active VPS RBSP, the value of all decoded chroma samples is equal to 1 << ( BitDepthC 1).

[0187] In the 4:4:4 video image format, the three channels (luminance channel and two chroma channels) have the same spatial resolution. Therefore, in some embodiments, these three channels can be used as subranges of the same video image.

[0188] For example, the depth map can have a bit depth of 24 bits, while the encoded video image format is 4:4:4 8 bits per channel. In this case, a mask can be applied, and then shifted to retrieve the first, second, and third parts of the 24-bit value and store them in each channel, as follows:

[0189] For other image formats, such as 4:2:0 and 4:2:2, the spatial resolution of the chroma channel is different from that of the luminance (Y) channel.

[0190] In such cases, different methods can be defined to map a single channel to multiple channels. For example, a single-channel image can be converted to an image in a color space (such as RGB, HSL), and then the converted image can be further converted to an encoded video image format (such as the traditional YCbCr used in video codec standards such as AVC, HEVC, VVC, etc.).

[0191] Other examples of such transformations exist, including in the field of depth image encoding and decoding (https: / / sites.google.com / site / brainrobotdata / home / depth-image-encoding) and depth image compression via coloring.

[0192] Figure 22 A schematic block diagram of an example system 800 is shown, in which various aspects and exemplary embodiments are implemented.

[0193] System 800 can be embedded as one or more devices, including the various components described below. In various exemplary embodiments, system 800 can be configured to implement one or more aspects described in this disclosure.

[0194] Examples of equipment that may constitute all or part of system 800 include personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors, "cave" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or web servers), still or video cameras, encoding or decoding chips, or any other communication devices. The elements of system 800 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of system 800 may be distributed across multiple ICs and / or discrete components. In various exemplary embodiments, system 800 may be communicatively connected to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0195] System 800 may include at least one processor 810 configured to execute instructions loaded therein for implementing various aspects, such as those described in this disclosure. Processor 810 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 800 may include at least one memory 820 (e.g., a volatile memory device and / or a non-volatile memory device). System 800 may include a storage device 840, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 840 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0196] System 800 may include an encoder / decoder module 830 configured to, for example, process data to provide encoded / decoded video image data, and the encoder / decoder module 830 may include its own processor and memory. The encoder / decoder module 830 may represent one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Furthermore, the encoder / decoder module 830 may be implemented as a separate element of system 800, or may be incorporated into processor 810 as a combination of hardware and software known to those skilled in the art.

[0197] Program code to be loaded into processor 810 or encoder / decoder 830 to execute the various aspects described in this disclosure may be stored in storage device 840 and subsequently loaded into memory 820 for execution by processor 810. According to various exemplary embodiments, during the execution of the processes described in this disclosure, one or more of processor 810, memory 820, storage device 840, and encoder / decoder module 830 may store one or more of various items. Such stored items may include, but are not limited to, video image data, information data for encoding / decoding video image data, bitstreams, matrices, variables, and intermediate or final results of equations, formulas, operations, and arithmetic logic processing.

[0198] In several exemplary embodiments, the memory within the processor 810 and / or encoder / decoder module 830 may be used to store instructions and provide working memory for processes that can be performed during encoding or decoding.

[0199] However, in other exemplary embodiments, external memory (e.g., the processing device may be processor 810 or encoder / decoder module 830) is used for one or more of these functions. External memory may be memory 820 and / or storage device 840, such as volatile memory and / or non-volatile flash memory. In several exemplary embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one exemplary embodiment, fast external volatile memory such as RAM can be used as working memory for video encoding / decoding operations, for example, for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 video), AVC, HEVC, EVC, VVC, AV1, etc.

[0200] As indicated in block 890, input to the components of system 800 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section capable of receiving, for example, RF signals transmitted over the air by a broadcasting device, (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, and (v) a bus, such as CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802-3) bus, when this disclosure is implemented in the automotive field.

[0201] In various exemplary embodiments, the input device of block 890 has associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements necessary for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further limiting the band to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some exemplary embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various exemplary embodiments may include one or more elements performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include tuners performing various functions among these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband.

[0202] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and re-filtering to the desired frequency band.

[0203] Various exemplary embodiments may rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0204] Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various exemplary embodiments, the RF portion may include an antenna.

[0205] Furthermore, USB and / or HDMI terminals may include corresponding interface processors for connecting system 800 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 810, as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 810, as needed. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, processor 810 and encoder / decoder 830, which operate in conjunction with memory and storage elements to process the data streams for presentation on an output device as needed.

[0206] Various components of system 800 can be provided within an integrated housing. Within the integrated housing, suitable connection arrangements 890, such as internal buses (including I2C buses), wiring, and printed circuit boards known in the art, can be used to interconnect various components and transfer data between them.

[0207] System 800 may include a communication interface 850 that enables communication with other devices via a communication channel 851. The communication interface 850 may include, but is not limited to, a transceiver configured to send and receive data on the communication channel 851. The communication interface 850 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 851 may be implemented, for example, within a wired and / or wireless medium.

[0208] In various exemplary embodiments, a Wi-Fi network such as IEEE 802.11 can be used to stream data to system 800. The Wi-Fi signals in these exemplary embodiments can be received via a communication channel 851 and a communication interface 850 suitable for Wi-Fi communication. The communication channel 851 in these exemplary embodiments can typically be connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top cloud communications.

[0209] Other exemplary embodiments may use a set-top box to provide streaming data to system 800, the set-top box delivering the data via an HDMI connection in input block 890.

[0210] Other exemplary embodiments may use the RF connection of input block 890 to provide streaming data to system 800.

[0211] Streamed data can be used as signaling information by System 800. Signaling information may include bitstreams and / or information such as video image pixel counts, any encoding / decoding settings parameters, alignment status, alignment reference data, overlap status, resampled data, interpolation data, and / or calibration data.

[0212] It should be recognized that signaling can be implemented in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc., can be used to signal information to the corresponding decoder.

[0213] System 800 can provide output signals to various output devices, including a display 861, a speaker 871, and other peripheral devices 881. In various examples of exemplary embodiments, other peripheral devices 881 may include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 800.

[0214] In various exemplary embodiments, control signals may be communicated between system 800 and display 861, speaker 871 or other peripheral device 881 using signaling such as AV.Link (audio / video link), CEC (consumer electronics control), or other communication protocols that enable device-to-device control with or without user intervention.

[0215] Output devices can be connected to system 800 via dedicated connections through the corresponding interfaces 860, 870 and 880.

[0216] Alternatively, the output device can be connected to the system 800 via communication interface 850 using communication channel 851. The display 861 and speaker 871 can be integrated with other components of the system 800 into a single unit in an electronic device, such as a television set.

[0217] In various exemplary embodiments, the display interface 860 may include a display driver, such as, for example, a timing controller (T Con) chip.

[0218] For example, if the RF portion of input 890 is part of a separate set-top box, then display 861 and speaker 871 may optionally be separate from one or more other components. In various exemplary embodiments where display 861 and speaker 871 can be external components, output signals may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0219] exist Figure 1-21This document describes various methods, each comprising one or more steps or actions to implement the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined.

[0220] Examples of block diagrams and / or operation flowcharts are described. Each block represents a portion of circuitry, a module, or code, which includes one or more executable instructions for implementing one or more specified logical functions. It should also be noted that in other implementations, the functions (one or more) marked in the blocks may occur out of order. For example, depending on the functions involved, two blocks shown sequentially may actually execute substantially concurrently, or sometimes these blocks may be executed in reverse order.

[0221] The embodiments and aspects described herein may be implemented in, for example, methods or processes, apparatus, computer programs, data streams, bit streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), implementations of the discussed features may be implemented in other forms (e.g., apparatus or computer programs).

[0222] The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices.

[0223] Furthermore, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values ​​generated by the implementation) can be stored on a computer-readable storage medium, such as storage device 840. Figure 9 Computer-readable storage media may take the form of a computer-readable program product implemented in one or more computer-readable media and having computer-readable program code executed thereon. Given the inherent ability to store information therein and the inherent ability to retrieve information provided therefrom, computer-readable storage media as used herein can be considered non-transitory storage media. Computer-readable storage media may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. It should be understood that while more specific examples of computer-readable storage media to which this exemplary embodiment may be applied are provided below, they are merely illustrative and not exhaustive, as will be readily apparent to those skilled in the art: portable computer floppy disks; hard disks; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable optical disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination of the foregoing.

[0224] Instructions can form applications that are tangibly implemented on processor-readable media.

[0225] For example, instructions can be found in hardware, firmware, software, or a combination thereof. Instructions can be found, for example, in an operating system, a standalone application, or a combination of both. Therefore, a processor can be characterized as, for example, a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Additionally, in addition to or instead of instructions, the processor-readable medium can store data values ​​generated by the implementation.

[0226] The device can be implemented, for example, in appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, head-mounted display devices (HMDs, see-through glasses), projectors, "caves" (systems comprising multiple displays), servers, video encoders, video decoders, post-processors that process the output from the video decoder, pre-processors that provide input to the video encoder, web servers, set-top boxes, and any other devices for processing video images, or other communication devices. It should be clear that the equipment can be mobile and even mounted in mobile vehicles.

[0227] The computer software may be implemented by the processor 810 or by hardware, or by a combination of hardware and software. As a non-limiting exemplary embodiment, the embodiment may also be implemented by one or more integrated circuits. The memory 820 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 810 may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures, as non-limiting examples.

[0228] As will be apparent to those skilled in the art based on this disclosure, implementations can generate various signals formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described exemplary embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

[0229] The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. As used herein, the singular forms “an,” “a,” and “the” may also be intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms “include / comprise” and / or “including / comprising” may specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is referred to as “in response to,” “connected to,” or “associated with,” another element, it may be directly responsive to, connected to, or associated with another element, or there may be intermediate elements. In contrast, when an element is referred to as “directly responsive to,” “directly connected to,” or “directly associated with,” another element, there are no intermediate elements.

[0230] It should be recognized that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the symbols / terms “ / ,” “and / or,” and “at least one of” can be intended to cover the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to cover the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or the selection of all three options (A, B, and C). As will be clear to those skilled in the art and related fields, this can be extended to as many items as are listed.

[0231] Various numerical values ​​may be used in this disclosure. Specific values ​​may be used for illustrative purposes and the aspects described are not limited to these specific values.

[0232] It will be understood that while the terms first, second, etc., may be used herein to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the teachings of this disclosure, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. There is no implied order between the first element and the second element.

[0233] References to “an exemplary embodiment” or “an exemplary embodiment” or “an implementation” or “implementation” and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Therefore, the phrases “in an exemplary embodiment” or “in an exemplary embodiment” or “in one implementation” or “in one implementation” appearing throughout this disclosure, as well as any other variations, do not necessarily refer to the same exemplary embodiment.

[0234] Similarly, the references to "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" and their variations are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an exemplary embodiment / example / implementation) may be included in at least one exemplary embodiment / example / implementation. Therefore, the expressions "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" appearing throughout this disclosure do not necessarily refer to the same exemplary embodiment / example / implementation, nor are individual or alternative exemplary embodiments / examples / implementations necessarily mutually exclusive with other exemplary embodiments / examples / implementations.

[0235] The reference numerals appearing in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, these exemplary embodiments / examples and variations may be employed in any combination or subcombination.

[0236] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0237] While some diagrams include arrows along the communication path to indicate the main direction of communication, it should be understood that communication can occur in the opposite direction to the arrows depicted.

[0238] Various implementations involve decoding. As used in this disclosure, "decoding" can encompass all or part of a process performed on a received video sequence (which may include a received bitstream encoded with one or more video sequences) to produce a final output suitable for display or further processing in a reconstructed video domain. In various exemplary embodiments, such processes include one or more processes typically performed by a decoder. In various exemplary embodiments, such processes also include, or optionally include, processes performed by a decoder of the various embodiments described in this disclosure.

[0239] As a further example, in one exemplary embodiment, "decoding" may refer only to dequantization; in another exemplary embodiment, "decoding" may refer to entropy decoding; in yet another exemplary embodiment, "decoding" may refer only to differential decoding; and in yet another exemplary embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. It will be clear, and believed to be well understood by those skilled in the art, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, depending on the context of the specific description.

[0240] Various implementations involve encoding. In a manner similar to the above discussion of “decoding,” the term “encoding” as used in this disclosure can encompass all or part of a process performed on an input video sequence to produce an output bitstream. In various exemplary embodiments, such a process includes one or more processes typically performed by an encoder. In various exemplary embodiments, such a process also includes, or optionally includes, a process performed by an encoder of the various embodiments described in this disclosure.

[0241] As a further example, in one exemplary embodiment, "encoding" may refer only to quantization; in another exemplary embodiment, "encoding" may refer only to entropy encoding; in yet another exemplary embodiment, "encoding" may refer only to differential encoding; and in still another exemplary embodiment, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. It will be clear, and believed to be well understood, by those skilled in the art, whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, depending on the context of the particular description.

[0242] Furthermore, this disclosure may refer to "obtaining" various types of information. Obtaining information may include one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0243] Furthermore, this disclosure may refer to "receiving" various messages. Receiving messages may include one or more of the following, such as access information or receiving information from a communication network.

[0244] Moreover, as used herein, the term "notify by signal" specifically refers to instructing the corresponding decoder to do something. For example, in some exemplary embodiments, the encoder notifies specific information, such as encoding / decoding parameters or encoded video image data, by signal. In this way, in exemplary embodiments, the same parameter can be used on both the encoder and decoder sides. Thus, for example, the encoder can transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, then signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various exemplary embodiments by avoiding the transmission of any actual functionality. It should be recognized that signal notification can be accomplished in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc., are used to notify the corresponding decoder of information by signal. Although the verb form of the term "notify by signal" is mentioned above, the term "notify by signal" can also be used as a noun herein.

[0245] Several implementations have been described. However, it should be understood that various modifications can be made. For example, elements of different implementations can be combined, supplemented, modified, or removed to produce other implementations. Furthermore, those skilled in the art will understand that other structures and processes can replace the disclosed structures and processes, and the resulting implementations will perform at least substantially the same functions in at least substantially the same manner (one or more) to achieve at least substantially the same results (one or more) as the disclosed implementations. Thus, these and other implementations are contemplated in this disclosure.

Claims

1. A method (100) for encoding a depth image sequence in a single-layer bitstream, wherein the method comprises: - Obtain (110) depth representation information, the depth representation information being defined for obtaining a depth map sequence based on a depth image sequence; - The depth image sequence is obtained by encoding the depth image sequence based on the depth representation information; - Encode (130) the depth map sequence and the depth representation information in a single-layer bitstream.

2. The method according to claim 1, wherein the method further comprises encoding (140) single-layer video data information in the single-layer bitstream, the single-layer video data information indicating that the single-layer bitstream is a single-layer bitstream including a depth map sequence.

3. A method for decoding a depth image sequence from a single-layer bitstream, wherein the method comprises: -Decode the (210) depth map sequence from the single-layer bitstream; and - Decode (220) depth representation information from the single-layer bitstream, the depth representation information defining information for obtaining a depth image sequence based on the decoded depth map sequence.

4. The method according to claim 3, wherein the method further comprises obtaining (230) the depth image sequence from the decoded depth map sequence based on the decoded depth representation information.

5. The method according to claim 3 or 4, wherein the method further comprises decoding (step 230) single-layer video data information from the single-layer bitstream, the single-layer video data information indicating that the bitstream is a single-layer bitstream including a depth map sequence.

6. The method according to any one of claims 1 to 5, wherein the depth representation information includes the minimum and maximum values ​​of the depth image sequence.

7. The method according to any one of claims 2 or 5, wherein the single-layer video data information is a depth marker.

8. The method of claim 7, wherein the depth flag is signaled in a sequence-level parameter set.

9. The method of claim 7, wherein the depth marker is notified via a signal as a depth map sequence supplementary enhancement information message.

10. The method according to claim 2 or 4, wherein the single-layer video data information is a specific value of the content type parameter.

11. The method of claim 10, wherein the content type parameter is signaled in a sequence-level parameter set.

12. The method of claim 10, wherein the content type parameter is notified via a signal as a supplementary enhancement message to the content type information.

13. The method of claim 12, wherein the content type information supplementary enhancement message further includes a syntax element specifying a mapping type between the depth map format and the video image format.

14. The method according to any one of claims 1 to 6, wherein the depth representation information is notified via a signal as a supplementary enhancement information message to the main depth representation information.

15. The method according to claim 2 or 4, wherein the single-layer video data information is signaled in a scalability dimension information supplementation and enhancement information message.

16. The method according to claim 2 or 4, wherein the content type parameter in the single-layer video data information as a supplementary enhancement information message for scalability dimension information is notified via a signal.

17. A bitstream formatted according to any one of claims 1-2 or 6-16.

18. A system comprising means for performing the method according to any one of claims 1 to 16.

19. A computer program product comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Improvement in step-spindles

    US101109A