Scalability Dimension Information in Video Coding
The proposed technique addresses issues in video coding by using a modified syntax element for SDI view identifiers, ensuring the absence of certain SEI messages without SDI, and restricting scalable nesting, thereby enhancing the reliability and accuracy of video coding processes.
Patent Information
- Application Number
- JP2023559809
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-02
- Filing Date
- 2022-04-02
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-04-02
AI Technical Summary
Existing video coding standards face challenges in preventing the length of the syntax element for scalability dimension information (SDI) view identifier from becoming zero, ensuring the absence of multi-view acquisition information SEI messages without SDI messages, and preventing scalable nesting of multi-view acquisition information SEI messages.
The technique involves using a syntax element obtained by subtracting L from the length of the SDI view identifier to prevent it from becoming zero, ensuring that multi-view acquisition information SEI messages are not present without SDI messages, and restricting the scalable nesting of these messages.
This approach effectively prevents errors in video coding, ensures accurate layer identification, and maintains the integrity of supplementary enhancement information messages within the bitstream.
Smart Images

Figure 0007693826000346 
Figure 0007693826000347 
Figure 0007693826000348
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This This application claims the priority and benefits of an application filed on April 2, 2021, and is based on International Patent Application No. PCT / CN2022 / 084992, filed on April 2, 2022. All of the aforementioned patent applications are, International Patent Application No. PCT / CN2021 / 085292 in their entirety, is hereby incorporated by reference incorporated herein by reference. into this specification.
[0002] The present disclosure generally relates to video coding, and more particularly to supplementary enhancement information (SEI) messages used in image / video coding.
Background Art
[0003] In the Internet and other digital communication networks, digital video occupies the maximum usage of bandwidth. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is expected to increase in the future.
Summary of the Invention
[0004] The disclosed aspects / embodiments provide a technique to prevent the length of the syntax element obtained by subtracting L from the length of the scalability dimension information (SDI) view identifier from becoming zero for specifying the view identifier of the i - th layer in the bitstream. The disclosed aspects / embodiments further provide a technique to ensure that when there is no SDI message in the bitstream, there is no multi - view acquisition information supplementary enhancement information (SEI) message or auxiliary information SEI message in the bitstream. Also, the disclosed aspects / embodiments provide a technique to prevent the multi - view acquisition information SEI message from being scalable - nested.
[0005] The first aspect relates to a method for processing video data. This method uses a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate a syntax element obtained by subtracting L from the length of the SDI view identifier, and based on the SDI SEI message, performs a conversion between a video media file and a bitstream.
[0006] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that the syntax element obtained by subtracting L from the length of the SDI view identifier is configured such that the length of the SDI view identifier value syntax element that specifies the view identifier of the i-th layer in the bitstream does not become zero.
[0007] Optionally, in any of the foregoing aspects, other embodiments of this aspect specify that L is equal to 1.
[0008] Optionally, in any of the foregoing aspects, other embodiments of this aspect specify that the syntax element obtained by subtracting L from the length of the SDI view identifier is designated as sdi_view_id_len_minus1.
[0009] Optionally, in any of the foregoing aspects, other embodiments of this aspect specify the SDI view identifier value syntax element as sdi_view_id_val[i].
[0010] Optionally, in any of the foregoing aspects, other embodiments of this aspect specify that the value obtained by adding 1 to the syntax element obtained by subtracting L from the length of the SDI view identifier specifies the length of the SDI view identifier value syntax element.
[0011] Optionally, in any of the foregoing aspects, other embodiments of this aspect specify that the syntax element obtained by subtracting L from the SDI view identifier length is coded as an unsigned integer using N bits.
[0012] Optionally, in any of the foregoing aspects, other embodiments of this aspect define that N is equal to 4.
[0013] Optionally, in any of the foregoing aspects, other embodiments of this aspect define that the syntax element obtained by subtracting L from the SDI view identifier length is coded as a fixed pattern bit string using N bits, a signed integer using N bits, a truncated binary, a K-th exponential Golomb coding syntax element of a signed integer with K equal to 0, or an M-th exponential Golomb coding syntax element of an unsigned integer with M equal to 0.
[0014] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that the bitstream is a bitstream within the scope.
[0015] Optionally, in any of the foregoing aspects, other embodiments of this aspect define that the multi-view information SEI message and the auxiliary information SEI message do not exist in a coded video sequence (CVS) unless the SDI SEI message exists in the CVS.
[0016] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that the multi-view information SEI message includes a multi-view acquisition information SEI message.
[0017] Optionally, in any of the preceding aspects, other embodiments of this aspect provide that the auxiliary information SEI message includes a depth representation information SEI message.
[0018] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that the auxiliary information SEI message includes an alpha channel information SEI message.
[0019] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that when a multi-view information SEI message or an auxiliary information SEI message is present in the bitstream, one or more of the SDI multi-view information flag and the SDI auxiliary information flag are equal to 1.
[0020] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that the multi-view information SEI message includes a multi-view acquisition information SEI message, and the multi-view acquisition information SEI message is not scalable nested.
[0021] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that an SEI message in the bitstream having a payload type equal to 179 is restricted from being included in the scalable nested SEI message.
[0022] Optionally, in any of the foregoing aspects, other embodiments of this aspect provide that an SEI message in the bitstream having a payload type equal to 3, 133, 179, 180, or 205 is restricted from being included in the scalable nesting SEI message.
[0023] A second aspect relates to an apparatus for processing video data including a processor and a non-transitory memory having instructions recorded thereon, which when executed by the processor cause the processor to use a scalability dimension information (SDI) auxiliary extension information (SEI) message indicating a syntax element obtained by subtracting L from the SDI view identifier length, and based on the SDI SEI message, to perform conversion between a video media file and a bitstream.
[0024] A third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by one or more processors, cause the coding device to use a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate an SDI view identifier length minus L syntax element and to convert a video media file and a bitstream based on the SDI SEI message.
[0025] A fourth aspect relates to a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to use a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate a syntax element obtained by subtracting L from the length of an SDI view identifier and to perform conversion between a video media file and a bitstream based on the SDI SEI message.
[0026] A fifth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method executed by a video processing device, the method including using a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate a syntax element obtained by subtracting L from the length of an SDI view identifier and converting between a video media file and a bitstream based on the SDI SEI message.
[0027] A sixth aspect relates to a method of storing a bitstream of video, the method including using a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate a syntax element obtained by subtracting L from the length of an SDI view identifier, generating a bitstream based on the SDI SEI message, and storing the bitstream on a non-transitory computer-readable recording medium.
[0028] For the purpose of clarification, any one of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to constitute a new embodiment within the scope of the present disclosure.
[0029] These features and other features can be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.
[0030] Next, for a more complete understanding of the present disclosure, reference may be made to the following brief description in connection with the accompanying drawings and detailed description.
Brief Description of the Drawings
[0031]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0032] At the outset, examples of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether presently known or existing. The present disclosure should in no way be limited to the embodiments, drawings, and techniques illustrated below, including the designs and embodiments exemplified and described herein, but may be varied within the scope of the appended claims and the full scope of their equivalents.
[0033] Video coding standards have mainly evolved through the development of well-known International Telecommunication Union (ITU-T) and International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) standards. The ITU-T has developed H.261 and H.263, the ISO / IEC has developed Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and both organizations have jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC) standards. See ITU-T and ISO / IEC, "High efficiency video coding", Rec. ITU-T H.265 | ISO / IEC 23008-2 (published version). Since H.262, video coding standards have been based on a hybrid video coding structure of temporal prediction + transform coding. To explore future video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new techniques have been adopted by the JVET and incorporated into a reference software named Joint Exploration Model (JEM). See J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, J. Boyce, "Algorithm description of Joint Exploration Test Model 7 (JEM7)", JVET-G1001, Aug. 2017. The JVET was later renamed the Joint Video Experts Team (JVET) when the Versatile Video Coding (VVC) project was officially launched. VVC is a new coding standard targeting a 50% bitrate reduction compared to HEVC, which was finalized at the 19th meeting of the JVET that ended on July 1, 2020. Rec. ITU-T H.266 | ISO / IEC 23090-3, "Versatile Video Coding", 2020.
[0034] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed to be used in a maximum wide range of applications, including not only conventional applications such as television broadcasting, video conferencing, and playback from storage media, but also new applications. This includes not only conventional applications such as television broadcasting, video conferencing, and playback from storage media, but also more recent and advanced use cases such as adaptive bitrate streaming, video region extraction, content synthesis and merging from multiple coded video bitstreams, multi-view video, scalable layer coding, 360° immersive media adapted to viewports, etc., and are designed to be used in a maximum wide range of applications. B. Bross, J. Chen, S. Liu, Y. Wang (editors), “Versatile Video Coding (Draft10)”, JVET-S2001, Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams”, 2020, and J. Boyce, V. Drugeon, G. Sullivan, Y.-K. Wang (editors), “Versatile Video Coding (Draft10)”, JVET-S2001, Rec. Wang (editors), “Versatile supplemental enhancement information messages for coded video bitstreams (Draft5)”, JVET-S2007.
[0035] The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard recently developed by MPEG.
[0036] FIG. 1 is a schematic diagram showing an example of layer - based prediction 100. The layer - based prediction 100 corresponds to one - direction inter - prediction and / or bi - direction inter - prediction and is also executed between pictures in different layers.
[0037] The layer - based prediction 100 is applied between pictures 111, 112, 113, and 114 in different layers and pictures 115, 116, 117, and 118. In the illustrated example, pictures 111, 112, 113, and 114 are part of layer N + 1 132, and pictures 115, 116, 117, and 118 are part of layer N 131. Layers such as layer N 131 and / or layer N + 1 132 are groups of pictures all associated with similar characteristic values such as similar size, quality, resolution, signal - to - noise ratio, capabilities, etc. In the illustrated example, layer N + 1 132 is associated with a larger image size than layer N 131. Thus, pictures 111, 112, 113, and 114 of layer N + 1 132 have, in this example, a larger picture size (e.g., greater height and width, so more samples) than pictures 115, 116, 117, and 118 of layer N 131. However, such pictures can be separated between layer N + 1 132 and layer N 131 by other characteristics. Only two layers, layer N + 1 132 and layer N 131, are shown, but the set of pictures can be separated into any number of layers based on the relevant characteristics. Layer N + 1 132 and layer N 131 may also be indicated by a layer ID. The layer ID is an item of data associated with a picture and indicating that the picture is part of the layer in which it is shown. Thus, each of pictures 111 - 118 may be associated with a corresponding layer ID to indicate that the corresponding layer N + 1 132 or layer N 131 contains the corresponding picture.
[0038] Pictures 111 to 118 in different layers 131 to 132 are configured to be alternatively displayed. Thus, pictures 111 to 118 in different layers 131 to 132 can share the same time identifier (ID) and can be included in the same access unit (AU) 106. As used herein, an AU is a set of one or more coded pictures associated with the same display time for output from a decoded picture buffer (DPB). For example, if a smaller picture is desired, the decoder can decode and display picture 115 at the current display time, or if a larger picture is desired, the decoder can decode and display picture 111 at the current display time. Thus, pictures 111 to 114 in the upper layer N+1 132 include substantially the same image data as the corresponding pictures 115 to 118 in the lower layer N 131 (regardless of the difference in picture size). Specifically, picture 111 includes substantially the same image data as picture 115, picture 112 includes substantially the same image data as picture 116, and so on.
[0039] Pictures 111 to 118 can be coded by referring to other pictures 111 to 118 within the same layer N 131 or N+1 132. Pictures coded by referring to another picture within the same layer are brought about by inter prediction 123 and conform to one-way inter prediction and / or bidirectional inter prediction. Inter prediction 123 is depicted by solid arrows. For example, picture 113 may be coded by adopting inter prediction 123 that uses one or two of pictures 111, 112, and / or 114 within layer N+1 132 as references, where one picture is referred to for one-way inter prediction and / or two pictures are referred to for bidirectional inter prediction. Further, picture 117 may be coded by adopting inter prediction 123 that uses one or two of pictures 115, 116, and / or 118 within layer N 131 as references, where one picture is referred to for one-way inter prediction and / or two pictures are referred to for bidirectional inter prediction. When a picture is used as a reference to another picture within the same layer when performing inter prediction 123, that picture can be referred to as a reference picture. For example, picture 112 may be a reference picture used to code picture 113 according to inter prediction 123. Inter prediction 123 may also be referred to as in-layer prediction in a multi-layer context. Thus, inter prediction 123 is a mechanism for coding samples of the current picture by referring to samples shown in a reference picture different from the current picture when the reference picture and the current picture are in the same layer.
[0040] Pictures 111 to 118 may be coded with reference to other pictures 111 to 118 of different layers. This process is well-known as inter-layer prediction 121 and is depicted by dashed arrows. Inter-layer prediction 121 is a mechanism in which samples of the current picture are coded by referring to samples represented by a reference picture in a layer different from the current picture, that is, a reference picture having a different layer ID. For example, a picture of the lower layer N 131 can be used as a reference picture for coding the corresponding picture of the upper layer N+1 132. As a specific example, picture 111 may be coded with reference to picture 115 by inter-layer prediction 121. In such a case, picture 115 is used as an inter-layer reference picture. An inter-layer reference picture is a reference picture used for inter-layer prediction 121. In most cases, inter-layer prediction 121 is restricted such that the current picture, such as picture 111, can only use inter-layer reference picture(s) in a lower layer, such as picture 115, that are included in the same AU106. When multiple layers (for example, two or more) are available, inter-layer prediction 121 can encode / decrypt the current picture based on multiple inter-layer reference picture(s) at a level lower than the current picture.
[0041] The video encoder can encode pictures 111-118 through many different combinations and / or permutations of inter prediction 123 and inter-layer prediction 121 by adopting layer-based prediction 100. For example, picture 115 can be coded according to intra prediction. Thereafter, pictures 116-118 can be coded according to inter prediction 123 by using picture 115 as a reference picture. Further, picture 111 can be coded according to inter-layer prediction 121 by using picture 115 as an inter-layer reference picture. Pictures 112-114 can be coded according to inter-layer prediction 123 by using picture 111 as a reference picture. In this way, the reference picture can serve both as a single-layer reference picture and an inter-layer reference picture for different coding mechanisms. By coding the upper-layer N+1 132 picture based on the lower-layer N 131 picture, the upper-layer N+1 132 can avoid adopting intra prediction with much lower coding efficiency than inter-layer prediction 123 and inter-layer prediction 121. In this way, the intra prediction with low coding efficiency is limited to the picture with the minimum / lowest quality, and thus can be limited to the coding of the minimum amount of video data. The reference picture and / or the picture used as an inter-layer reference picture can be indicated by the entry of the reference picture list included in the reference picture list structure.
[0042] Each access unit 106 in FIG. 1 may include a plurality of pictures. For example, one AU 106 may include pictures 111 and 115. Another AU 106 may include pictures 112 and 116. In fact, each AU 106 is a set of one or more coded pictures associated with the same presentation time (e.g., the same time ID) for output from the decoded picture buffer (DPB) (e.g., for display to the user). Each access unit delimiter (AUD) 108 is an indicator or data structure used to indicate the start of an AU (e.g., AU 108) or the boundary between AUs.
[0043] Previous H.26x video coding families have provided support for scalability in a profile separate from the profile for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 and provides support for spatial, temporal, and quality scalability. In SVC, for each macroblock (MB) of an Enhancement Layer (EL) picture, a flag indicating whether the ELMB is predicted using blocks arranged side by side from the lower layer is signaled. Prediction from the blocks arranged side by side can include texture, motion vectors, and / or coding modes. The implementation of SVC cannot directly reuse the execution of unmodified H.264 / AVC in its design. The syntax and decoding process of the EL macroblocks of SVC are different from those of H.264 / AVC.
[0044] Scalable High Efficiency Video Coding (SHVC) is an extension of the HEVC / H.265 standard that supports spatial and quality scalability, Multiview HEVC (MV-HEVC) is an extension of HEVC / H.265 that supports multiview scalability, and 3DHEVC (3D-HEVC) is an extension of HEVC / H.264 that supports more advanced and efficient three-dimensional (3D) video coding than MV-HEVC. Note that temporal scalability is included as an integral part of the single-layer HEVC codec. The design of the multi-layer extension of HEVC adopts the idea that the decoded pictures used for inter-layer prediction appear only from the same Access Unit (AU), are treated as Long-Term Reference Pictures (LTRP), and reference index values in the reference picture list are assigned for other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to the inter-layer reference pictures (if any) within the reference picture list (if any).
[0045] In particular, both resampling of reference pictures and the spatial scalability function require resampling of a reference picture or a part thereof. Reference picture resampling (RPR) can be implemented either at the picture level or at the coding block level. However, when RPR is referenced as a coding function, RPR is a function for single-layer coding. Nevertheless, it is possible or preferable from the codec design perspective to use the same resampling filter for both the RPR function for single-layer coding and the spatial scalability function for multi-layer coding.
[0046] Figure 2 shows an example of layer-based prediction 200 using an output layer set (OLS). Layer-based prediction 100 is compatible with unidirectional inter prediction and / or bidirectional inter prediction, and is also performed between pictures in different layers. The layer-based prediction in Figure 2 is the same as that in Figure 1. Therefore, for the sake of brevity, a full description of layer-based prediction will not be repeated.
[0047] A part of the layers of the coded video sequence (CVS) 290 in FIG. 2 is included in the OLS. The OLS is a set of layers in which one or more layers are designated as output layers. The output layer is the layer output by the OLS. FIG. 2 shows three different OLSs, namely OLS1, OLS2, and OLS3. As shown, OLS1 includes layer N 231 and layer N+1 232. Layer N 231 includes pictures 215, 216, 217, and 218, and layer N+1 232 includes pictures 211, 212, 213, and 214. OLS2 includes layer N 231, layer N+1 232, layer N+2 233, and layer N+3 234. Layer N+2 233 includes pictures 241, 242, 243, and 244, and layer N+3 234 includes pictures 251, 252, 253, and 254. OLS3 includes layer N 231, layer N+1 232, and layer N+2 233. Although three OLSs are shown, different numbers of OLSs may be used in actual applications. In the illustrated embodiment, none of the OLSs includes layer N+4 235 that includes pictures 261, 262, 263, and 264.
[0048] Different OLSs can each include any number of layers. Different OLSs are generated to accommodate the different coding capabilities of different devices with various coding capabilities. For example, OLS1 that includes only two layers is generated to accommodate a mobile phone with relatively limited coding capabilities. On the other hand, OLS2 that includes four layers may be generated to accommodate a large-screen TV that can decode higher layers than a mobile phone. OLS3 that includes three layers may be generated to accommodate a personal computer, a laptop computer, or a tablet computer, which may be able to decode higher layers than a mobile phone but may not be able to decode the topmost layers like a large-screen TV.
[0049] The layers shown in Figure 2 can all be independent of each other. That is, each layer can be coded without using inter-layer prediction (ILP). In this case, the layer is referred to as a simulcast layer. One or more of the layers shown in Figure 2 may be coded using ILP. Whether a layer is a simulcast layer or some of the layers are coded using ILP may be signaled by a flag in the video parameter set (VPS). When some of the layers use ILP, the layer dependencies between the layers are also signaled in the VPS.
[0050] In an embodiment, when multiple layers are simulcast layers, only one layer is selected for decoding and output. In an embodiment, when some of the layers use ILP, all the layers (e.g., the entire bitstream) are specified to be decoded, and a specific one of the layers is specified to be the output layer. The output layer or multiple layers can be, for example, any of 1) only the top layer, 2) all layers, or 3) a set of the top layer + the specified lower layers. For example, when the top layer + the specified lower layers are output by a flag in the VPS, layer N+3 234 (top layer) and layers N 231 and N+1 232 (lower layers) are output from OLS2.
[0051] Some of the layers shown in Figure 2 may be referred to as primary layers, and other layers may be referred to as auxiliary layers. For example, layer N 231 and layer N+1 232 may be referred to as primary layers, and layer N+2 233 and layer N+3 234 may be referred to as auxiliary layers. The auxiliary layer may be referred to as an alpha auxiliary layer or a depth auxiliary layer. When auxiliary information is present in the bitstream, the primary layer may be associated with the auxiliary layer.
[0052] Unfortunately, the existing standard has drawbacks. 1. Currently, the syntax element sdi_view_id_len is coded as u(4), and its value is required to be in the range from 0 to 15. This value specifies the bit length of the sdi_view_id_val[i] syntax element that designates the view ID of the i-th layer in the bitstream. However, the length of sdi_view_id_val[i] must not be 0.
[0053] 2. For example, when there is some auxiliary information in the bitstream, as indicated by the SDI SEI message (scalability dimension SEI message), depth representation information SEI message, or alpha channel information SEI message, it is unclear to which non-auxiliary layer or primary layer the auxiliary information applies.
[0054] 3. It makes no sense for the multi-view acquisition information SEI message, depth representation information SEI message, or alpha channel information SEI message to be present in the bitstream, but the scalability dimension information SEI message is not present in the bitstream.
[0055] 4. The multi-view acquisition information SEI message contains information about all views present in the bitstream. Therefore, it is meaningless to make it scalable and nested as currently permitted.
[0056] Disclosed herein is a technology that solves one or more of the aforementioned problems. For example, the present disclosure provides a technology that prevents the length of an SDI view ID value syntax element, which specifies a view identifier of the i-th layer in a bitstream, from becoming zero by using a syntax element obtained by subtracting L from the length of a scalability dimension information (SDI) view identifier. The disclosed aspect / embodiment further provides a technology that, when there is no SDI message in the bitstream, ensures that there is no multi-view acquisition information supplementary enhancement information (SEI) message or supplementary information SEI message in the bitstream. Also, the disclosed aspect / embodiment provides a technology that prevents a multi-view acquisition information SEI message from being scalably nested.
[0057] FIG. 3 is a diagram showing an embodiment of a video bitstream 300. As used herein, the video bitstream 300 may also be referred to as a coded video bitstream, a bitstream, or a variation thereof. As shown in FIG. 3, the bitstream 300 is composed of one or more of decoding capability information (DCI) 302, video parameter set (VPS) 304, sequence parameter set (SPS) 306, picture parameter set (PPS) 308, picture header (PH) 312, picture 314, and SEI message 322. Each of DCI 302, VPS 304, SPS 306, and PPS 308 may generally be referred to as a parameter set. In an embodiment, other parameter sets not shown in FIG. 3 may also be included in the bitstream 300. For example, an adaptive parameter set (APS) is a syntax structure that includes syntax elements applicable to zero or more slices determined by zero or more syntax elements found in a slice header.
[0058] DCI302 may also be referred to as a decoded parameter set (DPS) or a decoding parameter set, and is a syntax structure that contains syntax elements applied to the entire bitstream. DCI302 contains certain parameters during the validity period of a video bitstream (e.g., bitstream 300), which can be converted to the validity period of a session. DCI302 contains profile, level, and sub-profile information, and can determine the maximum complexity interlace point that is guaranteed never to be exceeded even if splicing of the video sequence occurs within the session. Additionally, if necessary, it contains constraint flags, which indicate that the video bitstream is restricted in the use of certain functions as indicated by the values of these flags. This can indicate that certain tools are not used in the bitstream, enabling resource allocation in the decoder implementation, etc. Similar to all parameter sets, DCI302 is present when first referenced, is referenced by the very first picture in the video sequence, and implies that it must be transmitted between the first network abstraction layer (NAL) units in the bitstream. It is possible for multiple DCI302s to exist in the bitstream, but the values of the syntax elements within them must not conflict when referenced.
[0059] VPS304 contains decoding dependencies or decoding information for the construction of reference picture sets in the auxiliary layer. VPS304 provides an overall perspective or view of the scalable sequence, including what types of operation points are provided, the profile, layer, and level of the operation points, and other high-level properties of the bitstream that can be used as a basis for session negotiation and content selection.
[0060] In an embodiment, when it is shown that some layers use ILP, VPS304 indicates that the total number of OLSs specified by the VPS is equal to the number of layers, that the i-th OLS includes the layers including layer indices from 0 to i, and that for each OLS, only the top layer within the OLS is output.
[0061] SPS306 includes data common to all pictures within a sequence of pictures (SOP). SPS306 is a syntax structure including syntax elements applied to zero or more entire CLVs as determined by the content of the syntax elements found in the PPS referred to by the syntax elements found in each picture header. In contrast, PPS308 includes data common to an entire picture. PPS308 is a syntax structure including syntax elements applied to zero or more entire coded pictures as determined by the syntax elements found in each picture header (e.g., PH312).
[0062] DCI302, VPS304, SPS306, and PPS308 are included in different types of network abstraction layer (NAL) units. An NAL unit is a syntax structure including an indication of the type of subsequent data (e.g., coded video data). NAL units are classified into video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include data representing the values of samples within a video picture, and non-VCL NAL units include related additional information such as parameter sets (important data applicable to multiple VCL NAL units) and auxiliary enhancement information (timing information and other auxiliary enhancement information that may improve the usability of the decoded video signal but is not necessary for decoding the values of samples within the video picture).
[0063] In an embodiment, DCI302 is included in a non-VCLNAL unit designated as a DCINAL unit or a DPSNAL unit. That is, a DCINAL unit has a DCINAL unit type (NUT), and a DPSNAL unit has a DPSNUT. In an embodiment, VPS304 is included in a non-VCLNAL unit designated as a VPSNAL unit. Thus, a VPSNAL unit has a VPSNUT. In an embodiment, SPS306 is a non-VCLNAL unit designated as an SPSNAL unit. Therefore, an SPSNAL unit has an SPSNUT. In an embodiment, PPS308 is included in a non-VCLNAL unit designated as a PPSNAL unit. Thus, a PPSNAL unit has a PPSNUT.
[0064] PH312 is a syntax structure that includes syntax elements applied to all slices (e.g., slice 318) of a coded picture (e.g., picture 314). In an embodiment, PH312 is a type of non-VCLNAL unit designated as a PHNAL unit. Thus, a PHNAL unit has a PHNUT (e.g., PH_NUT).
[0065] In an embodiment, the PHNAL unit related to PH312 has a temporal ID and a layer ID. The temporal ID identifier indicates the position of the PHNAL unit relative to other PHNAL units in the bitstream (e.g., bitstream 300). The layer ID indicates the layer (e.g., layer 131 or layer 132) that contains the PHNAL unit. In an embodiment, the temporal ID is similar to but different from the picture order count (POC). The POC uniquely identifies each picture in order. In a single-layer bitstream, the temporal ID and the POC are the same. In a multi-layer bitstream (e.g., see FIG. 1), pictures within the same AU have different POCs but the same temporal ID.
[0066] In an embodiment, the PHNAL unit precedes the VCLNAL unit that includes the first slice 318 of the associated picture 314. As a result, the picture header ID is signaled within PH312, and an association between PH312 and the slice 318 of the picture 314 associated with PH312 is established without the need to be referenced from the slice header 320. As a result, it can be inferred that all VCLNAL units between two PH312s belong to the same picture 314, and the picture 314 is associated with the first PH312 between the two PH312s. In an embodiment, the first VCLNAL unit following PH312 includes the first slice 318 of the picture 314 associated with PH312.
[0067] In an embodiment, the PHNAL unit follows a parameter set at the picture level (e.g., PPS), or a higher-level parameter set such as DCI (also known as DPS), VPS, SPS, PPS, etc., that has both a temporal ID and a layer ID smaller than the temporal ID and layer ID of the PHNAL unit, respectively. As a result, these parameter sets are not repeated within a picture or access unit. Due to this order, PH312 can be resolved immediately. That is, a parameter set that includes parameters related to the entire picture is placed before the PHNAL unit in the bitstream. Those that include parameters for a part of the picture are placed after the PHNAL unit.
[0068] In an alternative example, the PHNAL unit follows a picture-level parameter set and a prefix auxiliary enhancement information (SEI) message, or a higher-level parameter set such as DCI (also known as DPS), VPS, SPS, PPS, APS, SEI message, etc.
[0069] The picture 314 is an array of monochrome format luminance samples, or an array of luminance samples and two corresponding arrays of chrominance samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0070] Picture 314 may be either a frame or a field. However, in one CVS 316, all pictures 314 are either all frames or all fields. CVS 316 is a coded video sequence for all coded layer video sequences (CLVS) within video bitstream 300. In particular, when video bitstream 300 includes one layer, CVS 316 and CLVS are the same. CVS 316 and CLVS are different only when video bitstream 300 includes multiple layers (for example, as shown in FIGS. 1 and 2).
[0071] Each picture 314 includes one or more slices 318. A slice 318 is an integer number of complete tiles or an integer number of consecutive complete coding tree unit (CTU) rows within a tile of a picture (for example, picture 314). Each slice 318 is exclusively included in a single NAL unit (for example, a VCL NAL unit). A tile (not shown) is a rectangular region of CTUs within a specific tile column and a specific tile row in a picture (for example, picture 314). A CTU (not shown) is a coding tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of a picture having three sample arrays, or a CTB of samples of a monochrome picture, or a picture coded using three separate color planes and syntax structures used to code samples. A CTB (not shown) is an N×N block of samples for some value of N such that partitioning is the splitting of a component into CTBs. A block (not shown) is an M×N (M columns × N rows) array of samples (for example, pixels) or an M×N array of transform coefficients.
[0072] In an embodiment, each slice 318 includes a slice header 320. The slice header 320 is a portion that includes data elements related to all tiles or CTU rows within the tile represented by the slice 318 among the coded slices 318. That is, the slice header 320 includes information regarding the slice 318, such as, for example, the slice type and which reference picture is used.
[0073] Pictures 314 and their slices 318 constitute data related to an image or video to be encoded or decoded. Therefore, pictures 314 and their slices 318 may simply be referred to as the payload or data carried in the bitstream 300.
[0074] The bitstream 300 also includes one or more SEI messages such as SEI message 322, SEI message 326, and SEI message 328. The SEI message includes supplementary auxiliary information. The SEI message can include various types of data that indicate the timing of video pictures, or describe various characteristics of the coded video, or describe how the coded video can be used or extended. The SEI message can also include any user-defined data. The SEI message does not affect the core decoding process, but can indicate how it is recommended to post-process or display the video. Other high-level characteristics of the video content, such as an indication of the color space for the interpretation of the video content, are conveyed by video user usability information (VUI). In addition, VUI identifiers for indicating new color spaces, such as high dynamic range and wide color gamut video, have been added due to the development of new color spaces.
[0075] In an embodiment, the SEI message 322 is an SDI SEI message. The SDI SEI message is used to indicate which primary layer is associated with the auxiliary layer when auxiliary information is present in the bitstream. For example, the SDI SEI message includes one or more syntax elements 324 for indicating which primary layer is associated with the auxiliary layer when auxiliary information is present in the bitstream. Below, various SEI messages and the syntax elements included in those SEI messages will be described.
[0076] In an embodiment, the SEI message 326 is a multi-view information SEI message such as a multi-view acquisition information SEI message. When present in the bitstream 300, the multi-view information SEI message includes one or more syntax elements 324 that specify various parameters of the acquisition environment, such as intrinsic and extrinsic camera parameters. These parameters are useful for view warping and interpolation.
[0077] In an embodiment, the SEI message 328 may be an auxiliary information SEI message such as a depth representation information SEI message or an alpha channel information SEI message. If present in the bitstream 300, the depth representation information SEI message includes one or more syntax elements 324 that specify various depth representations for a depth view for the purpose of processing the decoded texture and depth view components before rendering on a three-dimensional (3D) display such as view synthesis. The SEI message may be associated with an instantaneous decoder refresh (IDR) access unit for the purpose of random access. If present in the bitstream 300, the alpha channel information SEI message includes alpha channel sample values and post-processing applied to the decoded alpha plane auxiliary picture, and one or more syntax elements 324 that provide information regarding one or more related primary pictures. Blending is a process of synthesizing two images into one image. One of the images to be blended is associated with an auxiliary image identified as an alpha plane. The alpha channel information SEI message may be used to specify how the pixel values of the blended image are to be converted into another image including interpretation values.
[0078] Those skilled in the art will appreciate that the bitstream 300 may contain other parameters and information.
[0079] To solve the above problems, methods as summarized below are disclosed. These techniques should be considered as an example for explaining general concepts and should not be construed narrowly. Further, these techniques may be applied individually or combined in any manner.
[0080] Example 1
[0081] 1) To solve Problem 1, as an example, instead of signaling the length of the view ID syntax element via, for example, the syntax element sdi_view_id_len, a value obtained by subtracting L from the length (e.g., L = 1) is signaled via, for example, the syntax element sdi_view_id_len_minusL.
[0082] a. As an example, further, the syntax element may be coded as an unsigned integer using N bits.
[0083] i. As an example, N is equal to 4.
[0084] ii. Alternatively, the syntax may be coded as a fixed pattern bit string using N bits, or a signed integer using N bits, or truncated binary, or a K (e.g., K = 0) - th exponential Golomb coding syntax element of a signed integer, or an M (e.g., M = 0) - th exponential Golomb coding syntax element of an unsigned integer.
[0085] b. As an example, alternatively, for example, the length is still signaled via the syntax element sdi_view_id_len, but there is a constraint that the value of the syntax element must not be equal to 0.
[0086] Example 2
[0087] 2) To solve Problem 2, it is proposed that an auxiliary layer (i.e., the layer corresponding to which sdi_aux_id[i] is 1 or 2) can be applied to one or more related layers.
[0088] a. As an example, one or more syntax elements indicating the related layers of each auxiliary layer may be signaled in the scalability dimension information SEI message.
[0089] i. As an example, the related layer is specified by a layer ID.
[0090] ii. In other examples, the relevant layer is specified by a layer index.
[0091] iii. In other examples, an indication of whether an auxiliary layer is applied to one or more relevant layers may be specified by one or more syntax elements for the relevant layer.
[0092] 1. As an example, syntax elements may be used to indicate whether to apply an auxiliary layer to all relevant layers.
[0093] 2. As an example, the syntax element may indicate whether the auxiliary layer is applied to a particular relevant layer.
[0094] a. As an example, one or more primary layers are indicated by a syntax element.
[0095] i. As an example, all primary layers may be indicated by a syntax element.
[0096] ii. As an example, only primary layers with a layer index smaller than the layer index of the auxiliary layer may be indicated by a syntax element.
[0097] iii. As an example, only primary layers with a layer index larger than the layer index of the auxiliary layer may be indicated by a syntax element.
[0098] b. As an example, the syntax element is coded as a flag.
[0099] b. Alternatively, it is proposed that one or more layers associated with each auxiliary layer may be derived without being explicitly signaled.
[0100] i. As an example, the related layer of each auxiliary layer may be a layer having a nuh_layer_id obtained by adding N1, N2, …, Nk to the nuh_layer_id of the auxiliary layer, where k is an integer, and when i, j (i!= j) are within the range of 1 to k, Ni!= Nj.
[0101] 1. As an example, k is equal to 1, and N1 is equal to 1, or 2, or -1, or -2.
[0102] 2. As an example, k is greater than 1.
[0103] a. As an example, k is equal to 2, N1 = 1, and N2 = 2.
[0104] ii. As an example, the related layer of each auxiliary layer may be a layer having a layer index obtained by adding N1, N2, …, Nk to the layer index of the auxiliary layer, where k is an integer, and when i, j (i!= j) are within the range of 1 to k, Ni!= Nj.
[0105] 1. As an example, k is equal to 1, and N1 is equal to 1, or 2, or -1, or -2.
[0106] 2. As an example, k is greater than 1.
[0107] a. As an example, k is equal to 2, N1 = 1, and N2 = 2.
[0108] c. Alternatively, the indication of the related layer of each auxiliary layer may be explicitly signaled as one of the scalability dimension information SEI messages or a group of syntax elements.
[0109] d. Alternatively, the display of the related layer (e.g., depth representation information or alpha channel information) of the auxiliary information SEI message may be explicitly signaled by one or more syntax elements within the auxiliary information SEI message.
[0110] i. As an example, the auxiliary information SEI message may refer to the depth representation information SEI message or the alpha channel information SEI message.
[0111] ii. As an example, one or more syntax elements may indicate the layer ID value of the associated layer.
[0112] 1. As an example, the layer ID indicated by the syntax element may be required to be less than or equal to the maximum value of the layer ID, i.e., vps_layer_id[vps_max_layers_minus1] or vps_layer_id[sdi_max_layers_minus1].
[0113] iii. As an example, one or more syntax elements may indicate the layer index value of the associated layer.
[0114] 1. As an example, the layer index indicated by the syntax element may be required to be less than the maximum number of layers in the bitstream (e.g., sdi_max_layers_minus1 + 1 or vps_max_layers_minus1 + 1).
[0115] iv. As an example, a signal may be signaled indicating whether one or more layers are associated with an auxiliary layer.
[0116] 1. As an example, one syntax element may be used to specify whether the auxiliary information SEI message applies to all layers.
[0117] a. As an example, when auxiliary_all_layer_flag is equal to X (X is 1 or 0), it can be specified that the auxiliary information SEI message applies to all associated primary layers.
[0118] 2. As an example, one or more syntax elements may be used to specify whether the supplementary information SEI message is applied to one or more layers.
[0119] a. As an example, N syntax elements may be used to specify whether the supplementary information SEI message is applied to N layers.
[0120] i. As an example, the syntax element may be coded as a flag using 1 bit.
[0121] b. As an example, one syntax element may be used to specify whether the supplementary information SEI message is applied to one or more layers.
[0122] i. As an example, the syntax element may be an exponential Golomb coding of the K-th (e.g., K = 0).
[0123] ii. As an example, a syntax element equal to 5 specifies that the supplementary information SEI message is applied to the 0-th and 2-nd layers but not to the 1-st layer.
[0124] 1. When the syntax element is equal to 5, the supplementary information SEI message is applied to the (N - 1)-th and (N - 3)-th layers and not to the (N - 2)-th layer.
[0125] c. The above syntax element may be signaled conditionally, for example, only when the supplementary information SEI message is not applied to all layers.
[0126] e. As an example, the indication of the number of related layers of the supplementary picture for one layer may be signaled in the bitstream.
[0127] f. As an example, the above syntax element may be signaled using an unsigned integer using N bits, or a fixed pattern bit sequence using N bits, or a signed integer using N bits, or a truncated binary, or an exponential Golomb coding syntax element with a signed integer K (e.g., K = 0) next, or an exponential Golomb coding syntax element with an unsigned integer M (e.g., M = 0) next.
[0128] g. As an example, the indication of the number of related layers of the auxiliary picture and / or the number of related layers of the auxiliary picture can be conditionally signaled only when, for example, the auxiliary picture is included in the i-th layer of bitstreamInScope (e.g., sdi_aux_id[i] > 0). bitstreamInScope (also known as the scope of the bitstream) is defined as a sequence of AUs consisting of the first AU containing the SDI SEI message and zero or more subsequent AUs in decoding order (however, subsequent AUs containing other SDI SEI messages are not included).
[0129] Example 3
[0130] 3) To solve Problem 3, a bitstream compliance requirement is added that a CVS without a scalability dimension information SEI message shall not have a multi-view or auxiliary information SEI message.
[0131] a. Further, the multi-view information SEI message may refer to the multi-view acquisition information SEI message.
[0132] b. Further, the auxiliary information SEI message may refer to the depth representation information SEI message or the alpha channel information SEI message.
[0133] c. Alternatively, when a multi-view or auxiliary information SEI message exists in the bitstream, it is added as a requirement for bitstream compliance that at least one of the sdi_multiview_info_flag and sdi_auxiliary_info_flag of the scalability dimension information SEI message is equal to 1.
[0134] Example 4
[0135] 4) To solve Problem 4, as an example, a requirement for bitstream compliance is added that the multi-view acquisition information SEI message must not be scalable nested.
[0136] a. Alternatively, it is specified that an SEI message with a payloadType equal to 179 (multi-view acquisition) must not be included in a scalable nesting SEI message.
[0137] Examples of several embodiments summarized above are shown below. Each embodiment is applicable to VVC. Most of the relevant parts that are added or changed are drawn in bold italic, and some of the deleted parts are drawn in italic. There may also be those not highlighted due to editorial changes.
[0138] Each scalability dimension SEI message syntax described below includes one or more syntax elements. A syntax element is, for example, one or more values, flags, variables, phrases, instructions, indexes, mappings, data elements, or combinations thereof included in the scalability dimension SEI message syntax disclosed herein. In an embodiment, the syntax elements may be organized into groups of values, flags, variables, phrases, instructions, indexes, mappings, and / or data elements.
[0139] Embodiment 1
[0140]
Table 1
[0141] Scalability Dimension SEI Message Semantics
[0142] The scalability dimension SEI message provides, as the scalability dimension information of each layer within bitstreamInScope (defined below), 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0143] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0144]
Chemical formula
[0145]
Chemical formula
[0146]
Chemical formula
[0147]
Chemical formula
[0148]
Chemical formula
[0149]
Chemical formula
[0150]
Table 2
[0151] The interpretation of the auxiliary picture related to sdi_aux_id in the range of 1 - 128 to 159 is specified by means other than the sdi_aux_id value.
[0152] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range from 0 to 2 or from 128 to 159. The value of sdi_aux_id[i] shall be in the range from 0 to 2 or from 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range from 0 to 255.
[0153] Embodiment 2
[0154]
Table 3
Table 4
[0155] Scalability Dimension SEI Message Semantics
[0156] The scalability dimension SEI message provides, as the scalability dimension information of each layer within bitstreamInScope (defined below), 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth and alpha values) is transmitted in one or more layers within bitstreamInScope.
[0157] The bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0158]
Chem.
[0159]
Chem.
[0160]
Chem.
[0161]
Chem.
[0162]
Chem.
[0163]
Chem.
[0164]
Table 5
[0165] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id included in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0166] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0167] Embodiment 3 Scalability Dimension SEI Message Syntax
[0168] [Table 6]
[0169] Scalability Dimension SEI Message Semantics
[0170] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, for example, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0171] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message) in decoding order.
[0172] [Figure]
[0173] [Figure]
[0174] [Chemistry]
[0175] [Chemistry]
[0176] [Chemistry]
[0177] [Chemistry]
[0178] [Table 7]
[0179] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0180] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range of 0 to 255.
[0181] Embodiment 4
[0182] [Table 8]
[0183] Scalability Dimension SEI Message Semantics
[0184] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0185] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0186]
Chem.
[0187]
Chem.
[0188]
Chem.
[0189]
Chem.
[0190]
Chem.
[0191]
Chem.
[0192]
Chem.
[0193]
Chem.
[0194]
Table 9
[0195] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0196] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range of 0 to 255.
[0197]
Chem.
[0198] Embodiment 5
[0199]
Table 10
[0200] Scalability Dimension SEI Message Semantics
[0201] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0202] bitstreamInScope is a sequence of AUs that, in decoding order, consists of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0203]
Chemical formula
[0204]
Chemical formula
[0205]
Chemical formula
[0206]
Chemical formula
[0207]
Chemical formula
[0208]
Chemical formula
[0209]
Table 11
[0210] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0211] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0212]
Chem.
[0213]
Chem.
[0214] Embodiment 6
[0215]
Table 12
Table 13
[0216] Scalability Dimension SEI Message Semantics
[0217] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, for example, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0218] The bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0219]
Chem.
[0220]
Chem.
[0221]
Chem.
[0222]
Chem.
[0223]
Chem.
[0224]
Chem.
[0225]
Table 14
[0226] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0227] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0228]
Chemical formula
[0229]
Chemical formula
[0230]
Chemical formula
[0231] Embodiment 7
[0232]
Table 15
Table 16
[0233] Scalability Dimension SEI Message Semantics
[0234] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, for example, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0235] The bitstreamInScope is a sequence of AUs that, in decoding order, consists of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding the subsequent AUs that contain the scalability dimension SEI message).
[0236]
Chem.
[0237]
Chem.
[0238]
Chem.
[0239]
Chem.
[0240]
Chem.
[0241]
Chem.
[0242]
Table 17
[0243] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0244] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. However, for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0245]
Chemical formula
[0246]
Chemical formula
[0247]
Chemical formula
[0248]
Chemical formula
[0249] Embodiment 8
[0250]
Table 18
Table 19
[0251] Scalability Dimension SEI Message Semantics
[0252] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer if bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0253] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0254]
Chemical formula
[0255]
Chemical formula
[0256]
Chemical formula
[0257]
Chemical formula
[0258]
Chemical formula
[0259]
Chemical formula
[0260]
Table 20
[0261] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0262] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range of 0 to 255.
[0263]
Chem.
[0264]
Chem.
[0265]
Chem.
[0266] Embodiment 9
[0267]
Table 21
Table 22
[0268] Scalability Dimension SEI Message Semantics
[0269] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0270] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0271]
[0272]
[0273]
[0274]
[0275]
[0276]
[0277]
[0278] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0279] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range of 0 to 255.
[0280]
Chem.
[0281]
Chem.
[0282]
Chem.
[0283] Embodiment 10
[0284]
Table 24
Table 25
[0285] Scalability Dimension SEI Message Semantics
[0286] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer if bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0287] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0288] [Chem.]
[0289] [Chem.]
[0290] [Chem.]
[0291] [Chem.]
[0292] [Chem.]
[0293] [Chem.]
[0294] [Table 26]
[0295] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0296] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0297]
Chem.
[0298]
Chem.
[0299]
Chem.
[0300]
Chem.
[0301] Embodiment 11
[0302]
Table 27
[0303] Scalability Dimension SEI Message Semantics
[0304] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, 1) the view ID of each layer if bitstreamInScope is a multi-view bitstream, 2) the auxiliary ID of each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope, etc.
[0305] bitstreamInScope is a sequence of AUs consisting of, in decoding order, the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing the scalability dimension SEI message).
[0306]
Chemical formula
[0307]
Chemical formula
[0308]
Chemical formula
[0309]
Chemical formula
[0310]
Chemical formula
[0311]
Chemical formula
[0312]
Table 28
[0313] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0314] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0315] Embodiment 12
[0316] [Table 29]
[0317] Scalability Dimension SEI Message Semantics
[0318] The scalability dimension SEI message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides, for example, 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0319] bitstreamInScope is a sequence of AUs that, in decoding order, consists of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs that contain the scalability dimension SEI message).
[0320] [Diagram]
[0321]
Chem.
[0322]
Chem.
[0323]
Chem.
[0324]
Chem.
[0325]
Chem.
[0326]
Chem.
[0327]
Table 30
[0328] Note 1 - The interpretation of the auxiliary picture related to sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0329] For the bit stream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but the value of sdi_aux_id[i] in this version of the decoder shall be in the range of 0 to 255.
[0330] Embodiment 13
[0331] Depth Representation Information SEI Message
[0332]
Table 31
Table 32
Table 33
[0333] Semantics of Depth Representation Information SEI Message
[0334] The syntax elements of the depth representation information SEI message specify various parameters of the auxiliary picture of type AUX_DEPTH for the purpose of processing the decoded first picture and auxiliary pictures before rendering on a 3D display such as view synthesis. Specifically, the range of the depth or disparity of the depth picture is specified.
[0335] If present, the depth representation information SEI message shall be associated with one or more layers for which the sdi_aux_id value is equal to AUX_DEPTH. The following semantics apply individually to each nuh_layer_id targetLayerId among the nuh_layer_id values to which the depth representation information SEI message applies.
[0336] If present, the depth representation information SEI message may be included in any access unit. If present, it is recommended to include the SEI message for the purpose of random access in an access unit where the nuh_layer_id is equal to targetLayerId and the coded picture is an Intra Random Access Picture (IRAP) picture.
[0337] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is a related first picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j in the range including 0 to 2 and 4 to 15.
[0338] The information indicated by the SEI message is applied, in decoding order, to all pictures having a nuh_layer_id equal to targetLayerId, excluding from the access unit containing the SEI message up to the next picture associated with the applicable depth representation information SEI message, and is applied to the earlier of either targetLayerId or the end of the CLVS of the picture having a nuh_layer_id equal to targetLayerId.
[0339]
Chem.
[0340]
Chem.
[0341]
Chem.
[0342]
Chem.
[0343]
Chem.
[0344]
Chem.
[0345] The variable maxVal is set equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is a value included in, or inferred for, the active SPS of the layer for which nuh_layer_id is equal to targetLayerId.
[0346]
Table 34
Table 35
[0347]
Chem.
[0348] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is useful when the value of depth_representation_type is 1 or 3.
[0349] The variables in the x column of Table Y2 are derived from the respective variables in the s, e, n, and v columns of Table Y2 as follows: - Except when the value of e is in the range 0 to 127, x is set equal to (-1) s * 2 e-31 * (1 + n ÷ 2 v ) - Otherwise (e = 0), x is set equal to (-1) s * 2 -(30+v) * n
[0350] Note 1 - The above specification is the same as that described in IEC60559:1989.
[0351]
Table 36
[0352] The values of DMin and DMax, if present, have the same ViewId as the ViewId of the auxiliary picture and are specified in units of the luminance sample width of the coded picture.
[0353] The units of the values of ZNear and ZFar are the same if present, but are not specified.
[0354]
Chemical formula
[0355]
Chemical formula
[0356] Note 2 - When depth_representation_type = 3, the auxiliary picture contains non-linearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample values from non-linear representation to linear representation, i.e., to uniformly quantized disparity values. The shape of this conversion is defined by a line segment approximation in the non-linear disparity space from 2D linear disparity. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of the additional nodes are transmitted in the form of deviations (depth_nonlinear_representation_model[i]) from the straight line curve. These deviations are uniformly distributed along the entire range of 0 to maxVal at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0357] The variable DepthLUT[i] in the range of 0 to maxVal is specified as follows: for(k = 0; k <= depth_nonlinear_representation_num_minus1 + 1; k++){ pos1 = (maxVal * k) / (depth_nonlinear_representation_num_minus1 + 2) dev1 = depth_nonlinear_representation_model[k] pos2 = (maxVal * (k + 1)) / (depth_nonlinear_representation_num_minus1 + 2) dev2 = depth_nonlinear_representation_model[k + 1](X) x1 = pos1 - dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x = Max(x1, 0); x <= Min(x2, maxVal); x++) DepthLUT[x] = Clip3(0, maxVal, Round(((x - x1) * (y2 - y1)) ÷ (x2 - x1) + y1)) }
[0358] When depth_representation_type = 3, for all decoded luminance sample values dS of the auxiliary picture in the range of 0 to maxVal, DepthLUT[dS] represents the disparity uniformly quantized in the range of 0 to maxVal.
[0359] The syntax structure specifies the value of the element of the depth representation information SEI message.
[0360] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen that represent floating-point values. When the syntax structure is included in another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced by the variable names used when the syntax structure is included.
[0361]
Chem.
[0362]
Chem.
[0363]
Chem.
[0364]
Chem.
[0365] Embodiment 14 Depth representation information SEI message
[0366]
Table 37
Table 38
[0367]
Table 39
[0368] Depth representation information SEI message semantics
[0369] The syntax elements of the depth representation information SEI message specify various parameters of the AUX_DEPTH type auxiliary picture for the purpose of decoding the first picture and the auxiliary picture before rendering on a 3D display such as view synthesis. Specifically, the range of the depth or parallax of the depth picture is specified.
[0370] If present, the depth representation information SEI message shall be associated with one or more layers for which the sdi_aux_id value is equal to AUX_DEPTH. The following semantics apply individually to each nuh_layer_id targetLayerId among the nuh_layer_id values to which the depth representation information SEI message applies.
[0371] If present, the depth representation information SEI message may be included in any access unit. If present, in an access unit where the coded picture with nuh_layer_id equal to targetLayerId is an IRAP picture, it is recommended to include the SEI message for random access purposes.
[0372] In an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is a related first picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j in the range including 0 to 2 and including 4 to 15.
[0373] The information indicated by the SEI message applies to all pictures with nuh_layer_id equal to targetLayerId, excluding from the access unit containing the SEI message up to the next picture associated with the applicable depth representation information SEI message in decoding order, to either targetLayerId or the end of the CLVS of nuh_layer_id equal to targetLayerId, whichever is earlier in decoding order.
[0374]
Chemical formula
[0375] [Chemistry]
[0376] [Chemistry]
[0377] [Chemistry]
[0378] [Chemistry]
[0379] [Chemistry]
[0380] [Chemistry]
[0381] The variable maxVal is set equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is the value included in, or inferred for, the active SPS of the layer for which nuh_layer_id is equal to targetLayerId.
[0382] [Table 40] [Table 41]
[0383] [Chemistry]
[0384] Note 1 - The disparity_ref_view_id exists only when d_min_flag = 1 or d_max_flag = 1, and is useful when the value of depth_representation_type is 1 or 3.
[0385] The variables in the x column of Table Y2 are derived from the variables in the s, e, n, and v columns of Table Y2 as follows: - Except when the value of e is in the range of 0 to 127, x is (-1) s * 2 e-31 *(1 + n ÷ 2 v ) and is set equal to it. - Otherwise (when e = 0), x is (-1) s * 2 -(30+v) * n and is set equal to it.
[0386] Note 1 - The above specification is the same as that described in IEC60559:1989.
[0387]
Table 42
[0388] If the values of DMin and DMax exist, they have the same ViewId as the ViewId of the auxiliary picture and are specified in units of the luminance sample width of the coded picture.
[0389] The units of the values of ZNear and ZFar are the same if they exist, but are not specified.
[0390]
Diagram
[0391]
Diagram
[0392] When depth_representation_type = 3, the auxiliary picture includes non-linearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample value from a non-linear representation to a linear representation, that is, to a uniformly quantized disparity value. The shape of this conversion is defined by a line segment approximation in the two-dimensional linear disparity - non-linear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of the additional nodes are transmitted in the form of deviations (depth_nonlinear_representation_model[i]) from the straight line curve. These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals corresponding to the value of nonlinear_depth_representation_num_minus1.
[0393] The variable DepthLUT[i] is specified as follows when i is in the range from 0 to maxVal: for(k = 0; k <= depth_nonlinear_representation_num_minus1 + 1; k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1 + 2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k + 1)) / (depth_nonlinear_representation_num_minus1 + 2) dev2=depth_nonlinear_representation_model[k + 1](X) x1=pos1 - dev1 y1=pos1 + dev1 x2=pos2 - dev2 y2=pos2 + dev2 for(x = Max(x1,0); x <= Min(x2,maxVal); x++) DepthLUT[x] = Clip3(0, maxVal, Round(((x - x1) * (y2 - y1)) ÷ (x2 - x1) + y1)) }
[0394] When depth_representation_type = 3, for all decoded luminance sample values dS of the auxiliary picture in the range of 0 to maxVal, DepthLUT[dS] represents the parallax uniformly quantized in the range of 0 to maxVal.
[0395] The syntax structure specifies the values of the elements of the depth representation information SEI message.
[0396] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen that represent floating-point values. When the syntax structure is included in another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced by the variable names used when the syntax structure is included.
[0397]
Chemical
[0398]
Chemical
[0399]
Chemical
[0400]
Chemical
[0401] Embodiment 15 Alpha channel information SEI message
[0402]
Table 43
[0403] Alpha Channel Information SEI Message Semantics
[0404] The Alpha Channel Information SEI message provides alpha channel sample values coded in an AUX_ALPHA type auxiliary picture and one or more associated primary pictures, and information regarding post-processing applied to the decoded alpha plane.
[0405] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, if there is an associated primary picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j from 0 to 2 inclusive and from 4 to 15 inclusive.
[0406] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA continue in output order until one or more of the following conditions are true: - The next picture where nuh_layer_id is equal to nuhLayerIdA is output in output order. - The CLVS containing the auxiliary picture picA ends. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer where nuh_layer_id is equal to nuhLayerIdA ends.
[0407] The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_ids to which the alpha channel information SEI message is applied.
[0408]
Chem.
[0409]
Chem.
[0410] Let currPic be the picture associated with the alpha channel information SEI message. The semantics of the alpha channel information SEI message persist for the current layer in the output order until one or more of the following conditions become true: - A new CLVS of the current layer starts. - The bitstream ends. - In an access unit containing an alpha channel information SEI message where nuh_layer_id is equal to targetLayerId, a picture picB where nuh_layer_id is equal to targetLayerId is output with a PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic, respectively, immediately after the call to decode the picture order count of picB.
[0411]
Chem.
[0412]
Chem.
[0413] [Chemistry]
[0414] [Chemistry]
[0415] [Chemistry]
[0416] [Chemistry]
[0417] [Chemistry]
[0418] Note When both alpha_channel_incr_flag and alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by alpha_channel_clip_type_flag should be applied first, and then the change specified by alpha_channel_incr_flag should be applied to obtain the interpreted sample value of the luminance samples of the auxiliary picture.
[0419] Embodiment 16 Alpha Channel Information SEI Message
[0420] [Table 44]
[0421] Semantics of Alpha Channel Information SEI Message
[0422] The alpha channel information SEI message provides alpha channel sample values and information regarding post-processing applied to decoded alpha channel samples coded in an AUX_ALPHA type auxiliary picture and one or more associated primary pictures.
[0423] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the associated primary picture (if any) is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0, and scalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to scalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j including 0 to 2 and 4 to 15.
[0424] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA persist in output order until one or more of the following conditions become true: - The next picture where nuh_layer_id is equal to nuhLayerIdA is output in output order. - The CLVS containing the auxiliary picture picA ends. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer where nuh_layer_id is equal to nuhLayerIdA ends.
[0425] The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_ids to which the alpha channel information SEI message is applied.
[0426]
Chemical
[0427] Set currPic to the picture associated with the alpha channel information SEI message. The semantics of the alpha channel information SEI message persist for the current layer in output order until one or more of the following conditions are true: - A new CLVS for the current layer begins. - The bitstream ends. - A picture picB with nuh_layer_id equal to targetLayerId that contains an alpha channel information SEI message with nuh_layer_id equal to targetLayerId is output with PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic, respectively, immediately after calling the decoding process for the picture order count of picB and currPic.
[0428]
Chemical formula
[0429]
Chemical formula
[0430]
Chemical formula
[0431]
Chemical formula
[0432]
Chemical formula
[0433]
Chem.
[0434]
Chem.
[0435]
Chem.
[0436]
Chem.
[0437] Note: When both the alpha_channel_incr_flag and the alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by the alpha_channel_clip_type_flag should be applied first, and then the change specified by the alpha_channel_incr_flag should be applied to obtain the interpreted sample value of the luminance samples of the auxiliary picture.
[0438] Embodiment 17 Multi-view acquisition information SEI message
[0439]
Table 45
Table 46
Table 47
[0440] Semantics of the multi-view acquisition information SEI message
[0441]
Chem.
[0442] The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id values to which the multi-view acquisition information SEI message is applied.
[0443] If it exists, the multi-view acquisition information SEI message applied to the current layer shall be included in the access unit containing the IRAP picture which is the first picture of the CLVS of the current layer. The information signaled in the SEI message is applied to the CLVS.
[0444]
Chem.
[0445]
Chem.
[0446]
Chem.
[0447]
Chem.
[0448] If the multi-view acquisition information SEI message is included in the scalable nesting SEI message, the syntax elements sn_ols_flag and sn_all_layers_flag of the scalable nesting SEI message shall be equal to 0.
[0449] The variable numViewsMinus1 is derived as follows: - If the multi-view acquisition information SEI message is not included in the scalable nesting SEI message, numViewsMinus1 is set to 0. - Otherwise (if the multi-view acquisition information SEI message is included in the scalable nesting SEI message), numViewsMinus1 is set equal to sn_num_layers_minus1.
[0450] Among the views in which the multi-view acquisition information SEI message contains multi-view acquisition information, there are also those that do not exist.
[0451] In the following semantics, the index i indicates the syntax elements and variables applied to the layer where nuh_layer_id is equal to NestingLayerId[i].
[0452] The external camera parameters are specified according to a right-handed coordinate system with the upper left corner of the image as the origin, that is, the (0, 0) coordinates, and the other corners of the image having non-negative coordinates. With these specifications, the 3D world point wP = [x y z] is mapped to the 2D camera point cP[i] = [u v 1] for the i-th camera as follows:
[0453]
Number
[0454] Here, A[i] represents the internal camera parameter matrix, R -1 [i] represents the inverse matrix of the rotation matrix R[i], T[i] represents the translation vector, and s (scalar value) is an arbitrary scale factor selected so that the third coordinate of cP[i] becomes 1. The elements of A[i], R[i], and T[i] are determined as specified below according to the syntax elements signaled in this SEI message.
[0455]
Chemical
[0456]
Chem.
[0457]
Chem.
[0458]
Chem.
[0459]
Chem.
[0460]
Chem.
[0461]
Chem.
[0462]
Chem.
[0463]
Chem.
[0464]
Chem.
[0465]
Chem.
[0466] The length of the mantissa_focal_length_y[i] syntactic element is variable and is determined as follows: - If exponent_focal_length_y[i] is 0, the length is Max(0, prec_focal_length - 30). - Otherwise (when exponent_focal_length_y[i] is in the range from 0 to 63, exclusive), the length is Max(0, exponent_focal_length_y[i] + prec_focal_length - 31).
[0467]
Chemical formula
[0468]
Chemical formula
[0469]
Chemical formula
[0470]
Chemical formula
[0471]
Chemical formula
[0472]
Chemical formula
[0473]
Chemical formula
[0474]
Chemical formula
[0475]
Chem.
[0476]
Chem.
[0477] The intrinsic matrix A[i] of the i-th camera is represented by the following equation.
[0478]
Math.
[0479] prec_rotation_param specifies the exponent of the maximum allowable truncation error of r[i][j][k] given by 2 -prec_rotation_param . The value of prec_rotation_param shall be in the range from 0 to 31.
[0480]
Chem.
[0481]
Chem.
[0482]
Chem.
[0483]
Chem.
[0484] The rotation matrix R[i] of the i-th camera is represented as follows:
[0485]
Math.
[0486]
Chem.
[0487]
Chem.
[0488]
Chem.
[0489] The translation vector T[i] of the i-th camera is expressed as follows:
[0490]
Math.
[0491] The relationship between the camera parameter variables and the corresponding syntax elements is specified in Table ZZ. Each component of the intrinsic matrix, rotation matrix, and translation vector is obtained as a variable x calculated as follows from the variables specified in Table ZZ: - When e is in the range from 0 to 63, x is set equal to (-1) s * 2 e-31 * (1 + n ÷ 2 v ). - Otherwise (e = 0), x is set equal to (-1) s * 2 -(30+v) * n.
[0492] Note - The above specification is similar to that described in IEC60559:1989.
[0493]
Table 48
[0494] Embodiment 18 Depth Representation Information SEI Message
[0495] [Table 49] [Table 50]
[0496] [Table 51]
[0497] Semantics of Depth Representation Information SEI Message
[0498] The syntax elements of the depth representation information SEI message specify various parameters of the auxiliary picture of the AUX_DEPTH type for the purpose of decoding the first picture and the auxiliary picture before rendering on a 3D display such as view synthesis. Specifically, the range of the depth or parallax of the depth picture is specified.
[0499] When the depth representation information SEI message exists, the depth representation information SEI message must be associated with one or more layers whose sdi_aux_id value is equal to AUX_DEPTH. The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id values to which the depth representation information SEI message is applied.
[0500] In this case, the depth representation information SEI message can be included in any access unit. When the SEI message exists, it is recommended that the SEI message be included for the purpose of random access in an access unit in which the coded picture having an nuh_layer_id equal to the targetLayerId is an IRAP picture.
[0501] [Chemistry]
[0502] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is a related first picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j from 0 to 2 inclusive and from 4 to 15 inclusive.
[0503] The information indicated in the SEI message is applied to all pictures having a nuh_layer_id equal to targetLayerId that are earlier in the decoding order, either until the next picture in the decoding order that is excluded from the depth representation information SEI message associated with the targetLayerId and included in the access unit containing the SEI message, or until the end of the CLVS for a nuh_layer_id equal to targetLayerId.
[0504] [Chemistry]
[0505] [Chemistry]
[0506] [Chemistry]
[0507] [Chemistry]
[0508]
Chem.
[0509] The variable maxVal is set equal to (1<<(8 + sps_bitdepth_minus8)) - 1. Here, sps_bitdepth_minus8 is a value included in or inferred from the active SPS of the layer where nuh_layer_id is equal to targetLayerId.
[0510]
Table 52
Table 53
[0511]
Chem.
[0512] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is valid when the value of depth_representation_type is equal to 1 and 3.
[0513] The variables in the x column of Table Y2 are derived from the variables in the s, e, n, and v columns of Table Y2 as follows: - If the value of e is in the range (exclusive) from 0 to 127, x is (-1) s * 2 e-31 * (1 + n ÷ 2 v ) and is set equal to it. - Otherwise (when e is equal to 0), x is (-1) s * 2 -(30+v) * n and is set equal to it.
[0514] Note 1 - The above specification is the same as that described in IEC60559:1989.
[0515]
Table 54
[0516] The DMin and DMax values, if present, are specified in units of the luminance sample width of the coded picture having a ViewId equal to the ViewId of the auxiliary picture.
[0517] The units of the ZNear and ZFar values are the same if present, but are not specified.
[0518]
Fig.
[0519]
Fig.
[0520] Note 2 - When depth_representation_type is equal to 3, the auxiliary picture contains non-linearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample values from a non-linear representation to a linear representation, i.e., a uniformly quantized disparity value. The shape of this conversion is defined by a line segment approximation in the two-dimensional linear disparity space - non-linear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of the additional nodes are communicated in the form of the deviation from the straight line curve (depth_nonlinear_representation_model[i]). These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0521] The variable DepthLUT[i] for i from 0 to maxVal is specified as follows: for(k = 0; k <= depth_nonlinear_representation_num_minus1 + 1; k++){ pos1 = (maxVal * k) / (depth_nonlinear_representation_num_minus1 + 2) dev1 = depth_nonlinear_representation_model[k] pos2 = (maxVal * (k + 1)) / (depth_nonlinear_representation_num_minus1 + 2) dev2 = depth_nonlinear_representation_model[k + 1](X) x1 = pos1 - dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x = Max(x1, 0); x <= Min(x2, maxVal); x++) DepthLUT[x] = Clip3(0, maxVal, Round(((x - x1) * (y2 - y1)) / (x2 - x1)) + y1)) }
[0522] When depth_representation_type is equal to 3, for all decoded luminance sample values dS of the auxiliary picture in the range of 0 to maxVal, DepthLUT[dS] represents the uniformly quantized disparity in the range of 0 to maxVal.
[0523] The syntax structure specifies the values of the elements of the depth representation information SEI message.
[0524] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen that represent floating-point values. When the syntax structure is included in another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced by the variable names used when the syntax structure is included.
[0525]
Chem.
[0526]
Chem.
[0527]
Chem.
[0528]
Chem.
[0529] Embodiment 19 Depth Representation Information SEI Message
[0530]
Table 55
Table 56
[0531]
Table 57
[0532] Depth Representation Information SEI Message Semantics
[0533]
Chem.
[0534]
Chem.
[0535]
Chem.
[0536]
Chem.
[0537] If present, the depth representation information SEI message can be included in any access unit. In an access unit where the coded picture with nuh_layer_id equal to targetLayerId is an IRAP picture, if the SEI message is present, it is recommended to include the SEI message for the purpose of random access.
[0538] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is a related first picture, that picture is in the same access unit and is such that sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j from 0 to 2 inclusive and from 4 to 15 inclusive.
[0539] The information indicated by the SEI message applies to the pictures in the access unit containing the SEI message, up to but excluding the next picture related to the depth representation information SEI message applicable to targetLayerId among all pictures having a nuh_layer_id equal to targetLayerId, or up to the end of the CLVS of nuh_layer_id equal to targetLayerId, whichever is earlier in the decoding order.
[0540]
Chem.
[0541] [Chemistry]
[0542] [Chemistry]
[0543] [Chemistry]
[0544] [Chemistry]
[0545] The variable maxVal is set equal to (1<<(8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is the value included in, or inferred for, the active SPS of the layer for which nuh_layer_id is equal to targetLayerId.
[0546] [Table 58] [Table 59]
[0547] [Chemistry]
[0548] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is valid when the value of depth_representation_type is equal to 1 and 3.
[0549] The variable in the x column of Table Y2 is derived as follows from the variables in the s, e, n, and v columns of Table Y2: - When the value of e is in the range (exclusive) from 0 to 127, x is (-1) s *2 e-31 *(1 + n÷2 v ) and is set equal to it. - Otherwise (e = 0), x is (-1) s *2 -(30+v) *n and is set equal to it.
[0550] Note 1 - The above specification is similar to that described in IEC60559:1989.
[0551]
Table 60
[0552] The DMin and DMax values, if present, are specified in units of the luminance sample width of the coded picture having a ViewId equal to the ViewId of the auxiliary picture.
[0553] The units of the ZNear and ZFar values are the same if present, but are not specified.
[0554]
Chem.
[0555]
Chem.
[0556] Note 2—When depth_representation_type is equal to 3, the auxiliary picture includes non-linearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample value from a non-linear representation to a linear representation, i.e., a uniformly quantized disparity value. The shape of this conversion is defined by a line segment approximation in a two-dimensional linear disparity space - non-linear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of the additional nodes are transmitted in the form of deviations (depth_nonlinear_representation_model[i]) from a straight line curve. These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0557] The variable DepthLUT[i] for i from 0 to maxVal is specified as follows: for(k=0;k<=depth_nonlinear_representation_num_minus1+1;k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1+2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k+1)) / (depth_nonlinear_representation_num_minus1+2) dev2=depth_nonlinear_representation_model[k+1](X) x1=pos1-dev1 y1=pos1+dev1 x2=pos2-dev2 y2=pos2+dev2 for(x=Max(x1,0);x<=Min(x2,maxVal);x++) DepthLUT[x] = Clip3(0, maxVal, Round(((x - x1) * (y2 - y1)) ÷ (x2 - x1) + y1)) }
[0558] When depth_representation_type is equal to 3, for all decoded luminance sample values dS of the auxiliary picture in the range of 0 to maxVal, DepthLUT[dS] represents the parallax uniformly quantized in the range of 0 to maxVal.
[0559] The syntax structure specifies the values of the elements of the depth representation information SEI message.
[0560] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen that represent floating-point values. When the syntax structure is included in another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced by the variable names used when the syntax structure is included.
[0561]
Chemical
[0562]
Chemical
[0563]
Chemical
[0564]
Chemical
[0565] Embodiment 20 Alpha channel information SEI message
[0566]
Table 61
[0567] Alpha Channel Information SEI Message Semantics
[0568] The Alpha Channel Information SEI message provides alpha channel sample values coded with an AUX_ALPHA type auxiliary picture and one or more associated primary pictures, and information regarding post-processing applied to the decoded alpha plane.
[0569] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, if there is an associated primary picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j from 0 to 2 inclusive and from 4 to 15 inclusive.
[0570] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA continue in output order until one or more of the following conditions are true: - The next picture where nuh_layer_id is equal to nuhLayerIdA is output in output order. - The CLVS containing the auxiliary picture picA ends. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer where nuh_layer_id is equal to nuhLayerIdA ends.
[0571] [Chemistry]
[0572] The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_ids to which the alpha channel information SEI message is applied.
[0573] [Chemistry]
[0574] Let currPic be the picture associated with the alpha channel information SEI message. The semantics of the alpha channel information SEI message persist for the current layer in output order until one or more of the following conditions are true: - A new CLVS of the current layer starts. - The bitstream ends. - A picture picB whose nuh_layer_id is equal to targetLayerId and which contains an alpha channel information SEI message with nuh_layer_id equal to targetLayerId is output with PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic, respectively, immediately after the call to decode the picture order count of picB.
[0575] [Chemistry]
[0576] [Chemistry]
[0577] [Chemistry]
[0578]
Chem.
[0579]
Chem.
[0580]
Chem.
[0581]
Chem.
[0582] Note: When both the alpha_channel_incr_flag and the alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by the alpha_channel_clip_type_flag should be applied first, and then the change specified by the alpha_channel_incr_flag should be applied to obtain the interpreted sample value of the luminance samples of the auxiliary picture.
[0583] Embodiment 21 Alpha Channel Information SEI Message
[0584]
Table 62
[0585] Alpha Channel Information SEI Message Semantics
[0586]
Chem.
[0587]
Chem.
[0588]
Chem.
[0589]
Chem.
[0590]
Chem.
[0591] The picture where nuh_layer_id is equal to nuhLayerIdA is output in the output order. - The CLVS including the auxiliary picture picA ends. - The bitstream ends. - The CLVS of the related primary layer of the auxiliary picture layer where nuh_layer_id is equal to nuhLayerIdA ends.
[0592] The following semantics are individually applied to the targetLayerId of each nuh_layer_id among the nuh_layer_ids to which the alpha channel information SEI message is applied.
[0593]
Chem.
[0594]
Chem.
[0595]
Chem.
[0596]
Chem.
[0597]
Chem.
[0598]
Chem.
[0599]
Chem.
[0600]
Chem.
[0601]
Chem.
[0602] Note: When both the alpha_channel_incr_flag and the alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by the alpha_channel_clip_type_flag should be applied first, and then the change specified by the alpha_channel_incr_flag should be applied to obtain the interpreted sample value of the luminance samples of the auxiliary picture.
[0603] Embodiment 22 Scalability Dimension Information (SDI) SEI Message
[0604]
Table 63
[0605] Scalability Dimension SEI Message Semantics
[0606] The scalability dimension SEI message provides, as the scalability dimension information of each layer within bitstreamInScope (defined below), 1) the view ID of each layer if bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer, etc., if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0607] bitstreamInScope is a sequence of zero or more AUs that, in decoding order, includes up to the AU in which the current scalability dimension SEI message is contained and does not include subsequent AUs in which the scalability dimension SEI message is contained.
[0608]
Chemical formula
[0609]
Chemical formula
[0610]
Chemical formula
[0611]
Chemical formula
[0612]
Chemical formula
[0613]
Chemical formula
[0614] [Chemical formula]
[0615] [Table 64]
[0616] The interpretation of the auxiliary picture related to sdi_aux_id in the range of 3 - 128 to 159 is specified by means other than the sdi_aux_id value.
[0617] For a bitstream compliant with this specification, sdi_aux_id[i] shall be in the range from 0 to 2 or from 128 to 159. The value of sdi_aux_id[i] shall be in the range from 0 to 2 or from 128 to 159. However, in this version of this specification, the decoder shall accept values of sdi_aux_id[[i] in the range including 0 to 255.
[0618] Multi-view acquisition information SEI message
[0619] [Table 65] [Table 66] [Table 67]
[0620] Semantics of multi-view acquisition information SEI message
[0621] [Chemical formula]
[0622] [Chemical formula]
[0623]
Chem.
[0624]
Chem.
[0625]
Chem.
[0626] The following semantics are applied individually to the targetLayerId of each nuh_layer_id among the nuh_layer_ids to which the multi-view acquisition information SEI message is applied.
[0627] If it exists, the multi-view acquisition information SEI message applied to the current layer must be included in the access unit containing the IRAP picture which is the first picture of the CLVS of the current layer. The information signaled in the SEI message is applied to the CLVS.
[0628]
Chem.
[0629]
Chem.
[0630] Among the views in which the multi-view acquisition information SEI message contains multi-view acquisition information, there are also views that do not exist.
[0631] In the following semantics, the index i refers to the syntax elements and variables applied to the layer where nuh_layer_id is equal to NestingLayerId[i].
[0632] The external parameters of the camera are specified according to a right-handed coordinate system with the upper left corner of the image as the origin, i.e., the (0, 0) coordinates, and the other corners of the image having non-negative coordinates. With these specifications, a 3D world point wP = [x y z] is mapped to a 2D camera point cP[i] = [u v 1] for the i-th camera as follows:
[0633]
Number
[0634] Here, A[i] represents the matrix of the internal parameters of the camera, R -1 [i] represents the inverse matrix of the rotation matrix R[i], T[i] represents the translation vector, and s (a scalar value) is an arbitrary magnification factor selected such that the third coordinate of cP[i] becomes 1. The elements of A[i], R[i], and T[i] are determined as specified below according to the syntax elements signaled in this SEI message.
[0635]
Chemical
[0636]
Chemical
[0637]
Chemical
[0638]
Chemical
[0639]
Chemical
[0640] [Chemistry]
[0641] [Chemistry]
[0642] [Chemistry]
[0643] [Chemistry]
[0644] [Chemistry]
[0645] [Chemistry]
[0646] [Chemistry]
[0647] The length of the mantissa_focal_length_y[i] syntactic element is variable and is determined as follows: If exponent_focal_length_y[i] is 0, the length is Max(0, prec_focal_length - 30). - Otherwise (when exponent_focal_length_y[i] is in the range from 0 to 63, exclusive), the length is Max(0, exponent_focal_length_y[i] + prec_focal_length - 31).
[0648] [Chemistry]
[0649]
Chem.
[0650]
Chem.
[0651]
Chem.
[0652]
Chem.
[0653]
Chem.
[0654]
Chem.
[0655]
Chem.
[0656]
Chem.
[0657]
Chem.
[0658] The intrinsic matrix A[i] of the i-th camera is represented by the following equation:
[0659]
Math.
[0660] The prec_rotation_param specifies the exponent of the maximum allowable truncation error of r[i][j][k] given by 2. The value of prec_rotation_param shall be in the range from 0 to 31. -prec_rotation_param
[0661]
Chem.
[0662]
Chem.
[0663]
Chem.
[0664]
Chem.
[0665] The rotation matrix R[i] of the i-th camera is expressed as follows:
[0666]
Math.
[0667]
Chem.
[0668]
Chem.
[0669]
Chem.
[0670] The translation vector T[i] of the i-th camera is expressed as follows:
[0671]
Number
[0672] The relationship between the parameter variables of the camera and the corresponding syntax elements is specified in Table ZZ. Each component of the intrinsic matrix, rotation matrix, and translation vector is obtained as a variable x calculated as follows from the variables specified in Table ZZ: - When e is in the range from 0 to 63, x is (-1) s * 2 e-31 * (1 + n ÷ 2 v ) and is set equal to it. - Otherwise (when e is equal to 0), x is (-1) s * 2 -(30+v) * n and is set equal to it.
[0673] Note - The above specification is similar to that described in IEC60559:1989.
[0674]
Table 68
[0675] FIG. 4 is a block diagram showing an exemplary video processing system 400 in which various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of the video processing system 400. The video processing system 400 may include an input 402 for receiving video content. The video content may be received in a raw format or an uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or may be in a compressed format or an encoded format. The input 402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi and cellular interfaces.
[0676] The imaging processing system 400 may include a coding component 404 that can implement various coding or encoding methods described herein. The coding component 404 may reduce the average bit rate of the video from the input 402 to the output of the coding component 404 to generate a coded representation of the video. Thus, coding techniques may also be referred to as video compression techniques or video transcoding techniques. The output of the encoding component 404 is either stored or transmitted via a communication connection as represented by component 406. The stored or communicated bitstream (or coded) representation of the video received at input 402 may be used by component 408 to generate pixel values or a displayable video that is transmitted to the display interface 410. The process of generating a video viewable by a user from the bitstream representation may be referred to as video restoration. Further, certain video processing operations are referred to as "coding" operations or tools, but it will be understood that coding tools or operations are used in an encoder and corresponding decoding tools or operations that reverse the results of coding are executed in a decoder.
[0677] Examples of a peripheral bus interface and a display interface include USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface) (registered trademark), Displayport, and the like. Examples of a storage interface include SATA (serial advanced technology attachment), PCI (Peripheral Component Interconnect), IDE (Integrated Drive Electronics) interface, and the like. The techniques described herein can be embodied in various electronic devices capable of performing digital data processing and / or video display, such as mobile phones, laptops, smartphones, and the like.
[0678] FIG. 5 is a block diagram of a video processing apparatus 500. The apparatus 500 can be used to implement one or more methods described herein. The apparatus 500 can be embodied in a smartphone, a tablet, a computer, an IoT (Internet of Things) receiver, etc. The apparatus 500 may include one or more processors 502, one or more memories 504, and video processing hardware 506 (also known as a video processing circuit). The processor(s) 502 may be configured to implement one or more methods described herein. The memory (storage device) 504 may be used to store data and coding used to implement the methods and techniques described herein. The video processing hardware 506 may be used to implement some of the techniques described herein in a hardware circuit. In some embodiments, the hardware 506 may be partially or fully disposed in the processor 502, such as a graphics processor, etc.
[0679] FIG. 6 is a block diagram showing an exemplary video coding system 600 that can utilize the technology of the present disclosure. As shown in FIG. 6, the video coding system 600 may include a source device 610 and a destination device 620. The source device 610 generates encoded video data, which may be referred to as a video encoding device. The destination device 620 may decode the encoded video data generated by the source device 610, which may be referred to as a video decoding device.
[0680] The source device 610 may include a video source 612, a video encoder 614, and an input / output (I / O) interface 616.
[0681] The video source 612 may include a video capture device, an interface for receiving video data from a video content provider, and / or a source such as a computer graphics system for generating video data, or a combination of such sources. The video data may be composed of one or more pictures. The video encoder 614 encodes the video data from the video source 612 to generate a bitstream. The bitstream may include a bit sequence forming a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded picture is a coded representation of the picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 616 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 620 via the I / O interface 616 over the network 630. The encoded video data may also be stored in the storage medium / server 640 for access by the destination device 620.
[0682] The destination device 620 may include an I / O interface 626, a video decoder 624, and a display device 622.
[0683] The I / O interface 626 may include a receiver and / or a modem. The I / O interface 626 may obtain the encoded video data from the source device 610 or the storage medium / server 640. The video decoder 624 may decode the encoded video data. The display device 622 may display the decoded video data to the user. The display device 622 may be integrated with the destination device 620, or may be external to the destination device 620 and configured to connect to an external display device.
[0684] The video encoder 614 and the video decoder 624 can operate according to video compression standards such as the HEVC (High Efficiency Video Coding) standard, the VVC (Versatile Video Coding) standard, and other current and / or future standards.
[0685] FIG. 7 is a block diagram showing an example of a video encoder 700, which may be the video encoder 614 in the video coding system 600 shown in FIG. 6.
[0686] The video encoder 700 can be configured to execute any or all of the techniques of the present disclosure. In the example of FIG. 7, the video encoder 700 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 700. In some examples, the processor may be configured to execute any or all of the techniques described in the present disclosure.
[0687] The functional components of the video encoder 700 may include a prediction unit 702 that may include a splitting unit 701, a mode selection unit 703, a motion estimation unit 704, a motion compensation unit 705, and an intra prediction unit 706, a residual generation unit 707, a conversion unit 708, a quantization unit 709, an inverse quantization unit 710, an inverse conversion unit 711, a reconstruction unit 712, a buffer 713, and an entropy encoding unit 714.
[0688] In other examples, the video encoder 700 can include more, fewer, or different functional components. In one example, the prediction unit 702 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in the IBC mode where at least one reference picture is the picture in which the current video block is located.
[0689] Furthermore, some components such as the motion estimation unit 704 and the motion compensation unit 705 may be highly integrated, but are shown separately for illustrative purposes in the example of FIG. 7.
[0690] The splitting unit 701 can split a picture into one or more video blocks. The video encoder 614 and the video decoder 624 in FIG. 6 can support various video block sizes.
[0691] The mode selection unit 703 selects either an intra or an inter coding mode, for example, based on an error result, and provides the resulting intra or inter coded block to a residual generation unit 707 that generates residual block data and a reconstruction unit 712 that reconstructs the coded block for use as a reference picture. In some examples, the mode selection unit 703 may select a combination of intra prediction and inter prediction (CIIP) modes based on a prediction being based on an inter prediction signal and an intra prediction signal. Also, the mode selection unit 703 may select the resolution of the motion vector for a block (e.g., sub-pixel or integer pixel accuracy) in the case of inter prediction.
[0692] To perform inter prediction on the current video block, the motion compensation unit 704 may generate motion information for the current video block by comparing one or more reference frames from the buffer 713 with the current video block. The motion compensation unit 705 may determine a predicted video block of the current video block based on the motion information and the decoded samples of the picture from the buffer 713 other than the picture associated with the current video block.
[0693] The motion estimation unit 704 and the motion compensation unit 705 can execute different operations on the current video block according to, for example, whether the current video block is an I slice, a P slice, or a B slice. An I slice (or an I frame) has the lowest compressibility but does not require other video frames for decoding. An S slice (or a P frame) can use the data of the previous frame for decompression and has higher compressibility than an I frame. A B slice (or a B frame) can use both the previous frame and the next frame for data reference to obtain the highest data compression ratio.
[0694] In some examples, the motion estimation unit 704 may perform unidirectional prediction on the current video block, and the motion estimation unit 704 may search for a reference picture in list 0 or list 1 for the reference video block for the current video block. Next, the motion estimation unit 704 may generate a reference index indicating the reference picture in list 0 or list 1 including the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 704 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 705 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.
[0695] In another example, the motion estimation unit 704 may perform bidirectional prediction on the current video block. The motion estimation unit 704 may search for the reference picture in list 0 for the reference video block for the current video block, and may also search for the reference picture in list 1 for another reference video block for the current video block. The motion estimation unit 704 may generate a reference index indicating the reference pictures in list 0 and list 1 including the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 704 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 705 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0696] In some examples, the motion estimation unit 704 can output a full set of motion information for decoder decoding processing.
[0697] In some examples, the motion estimation unit 704 may not need to output a full set of motion information of the current video. Rather, the motion estimation unit 704 may refer to the motion information of another video block and signal the motion information of the current video block. For example, the motion estimation unit 704 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0698] As an example, the motion estimation unit 704 can indicate a value that indicates to the video decoder 624 that the current video block has the same motion information as another video block in the syntax structure related to the current video block.
[0699] In another example, the motion estimation unit 704 may identify another video block and a motion vector difference (MVD) in a syntax structure related to the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 624 can use the difference between the motion vector of the indicated video block and the motion vector to determine the motion vector of the current video block.
[0700] As described above, the video encoder 614 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 614 are advanced motion vector prediction (AMVP) and merge mode signaling.
[0701] The intra prediction unit 706 may perform intra prediction on the current video block. When the intra prediction unit 706 performs intra prediction on the current video block, the intra prediction unit 706 may generate prediction data for the current video block based on the decoded samples of other video blocks within the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0702] The residual generation unit 707 may generate residual data for the current video block by subtracting the predicted video block(s) of the current video block from the current video block (e.g., indicated by a minus sign). The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples of the current video block.
[0703] In another example, there may be no residual data for the current video block. For example, in skip mode, the residual generation unit 707 may not perform the subtraction operation.
[0704] The conversion unit 708 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0705] After the conversion unit 708 generates the transform coefficient video block associated with the current video block, the quantization unit 709 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0706] The inverse quantization unit 710 and the inverse conversion unit 711 may respectively apply inverse quantization and inverse conversion to the transform coefficient video block to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 712 can generate a reconstructed video block associated with the current block for storage in the buffer 713 by adding the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 702.
[0707] After the reconstruction unit 712 reconstructs the video block, loop filtering processing may be performed to reduce video block artifacts within the video block.
[0708] The entropy encoding unit 714 may receive data from other functional components of the video encoder 700. When the entropy encoding unit 714 receives the data, the entropy encoding unit 714 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0709] FIG. 8 is a block diagram showing an example of a video decoder 800, which may be the video decoder 624 in the video coding system 600 shown in FIG. 6.
[0710] The video decoder 800 may be configured to execute some or all of the techniques of the present disclosure. In the example of FIG. 8, the video decoder 800 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the components of the video decoder 800. In some examples, a processor may be configured to execute some or all of the techniques described in the present disclosure.
[0711] In the example of FIG. 8, the video decoder 800 includes an entropy decoder 801, a motion compensation unit 802, an intra prediction unit 803, an inverse quantization unit 804, an inverse transform unit 805, a reconstruction unit 806, and a buffer 807. In some examples, the video decoder 800 can execute a decoding path generally inverse to the encoding path described with respect to the video encoder 614 (FIG. 6).
[0712] The entropy decoder 801 may obtain an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., blocks of encoded video data). The entropy decoder 801 may decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 802 may determine motion information including motion vectors, motion vector precision, indices of reference picture lists, and other motion information. The motion compensation unit 802 may determine such information, for example, by performing AMVP and merge mode signaling.
[0713] The motion compensation unit 802 may generate motion-compensated blocks and perform interpolation based on an interpolation filter when possible. An identifier of the interpolation filter used with sub-pixel precision may be included in a syntax element.
[0714] The motion compensation unit 802 may calculate the interpolation values of the sub-integer pixels of the reference block using the interpolation filter used by the video encoder 614 during the encoding of the video block. The motion compensation unit 802 may determine the interpolation filter used by the video encoder 614 according to the received syntax information, and generate a prediction block using the interpolation filter.
[0715] The motion compensation unit 802 can use a part of the syntax information to determine the size of the block used to encode the frame(s) and / or slice(s) of the encoded video sequence, the partition information explaining how each macroblock of the picture of the encoded video sequence is divided, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) of each inter-encoded block, and other information for decoding the encoded video sequence.
[0716] The intra prediction unit 803 can form a prediction block from spatially adjacent blocks using, for example, the intra prediction mode received in the bitstream. The inverse quantization unit 804 inverse quantizes, i.e., de-quantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 801. The inverse transform unit 805 applies an inverse transform.
[0717] The reconstruction unit 806 can sum the residual block with the corresponding prediction block generated by the motion compensation unit 802 or the intra prediction unit 803 to form a decoded block. Optionally, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is stored in the buffer 807, provides a reference block for subsequent motion compensation / intra prediction, and generates a decoded video for display on a display device.
[0718] FIG. 9 shows a method 900 for coding video data according to an embodiment of the present disclosure. The method 900 may be executed by a coding device (e.g., an encoder) having a processor and a memory. The method 900 may be executed when using an SEI message to transmit information within a bitstream.
[0719] In block 902, the coding device uses a scalability dimension information (SDI) supplementary enhancement information (SEI) message to indicate a syntax element obtained by subtracting L from the length of the SDI view identifier. The SDI SEI message is of a type such as, for example, the SEI message of the bitstream 300 in FIG. 3. The syntax element obtained by subtracting L from the length of the SDI view identifier is a kind of syntax element such as, for example, the syntax element 324 of the bitstream 300 in FIG. 3. The SEI message including the SDI SEI message can transmit any of the syntax elements disclosed herein.
[0720] In block 904, the coding device performs a conversion between a video media file and a bitstream based on the SDI SEI message.
[0721] When implemented in an encoder, the conversion includes receiving a media file (e.g., a video unit) and encoding the SEI message into the bitstream. When implemented in a decoder, the conversion includes receiving a bitstream including the SEI message and decoding the SEI message in the bitstream to generate a video media file.
[0722] In an embodiment, the syntax element obtained by subtracting L from the length of the SDI view identifier is configured such that the length of the SDI view identifier value syntax element that specifies the view identifier of the i-th layer in the bitstream does not become zero. In an embodiment, L is equal to 1. In an embodiment, the syntax element obtained by subtracting L from the length of the SDI view identifier is designated as sdi_view_id_len_minus1. In an embodiment, the SDI view identifier value syntax element is designated as sdi_view_id_val[i]. In an embodiment, adding 1 to the syntax element obtained by subtracting L from the length of the SDI view identifier specifies the length of the SDI view identifier value syntax element.
[0723] In an embodiment, the syntax element obtained by subtracting L from the length of the SDI view identifier is coded as an unsigned integer using N bits. As an example, an unsigned integer is an integer (e.g., a natural number) that does not have a sign (e.g., positive or negative). In an embodiment, N is equal to 4.
[0724] In an embodiment, the syntax element obtained by subtracting L from the length of the SDI view identifier is coded as a fixed pattern bit string using N bits, a signed integer using N bits, a truncated binary, a K-th exponential Golomb coding syntax element of a signed integer with K equal to 0, or an M-th exponential Golomb coding syntax element of an unsigned integer with M equal to 0. A bit string is an array data structure that stores bits compactly. A fixed pattern bit string is an array data structure that has a fixed pattern. An unsigned integer is an integer (e.g., a natural number) that does not have a sign (e.g., positive or negative). Truncated binary, or truncated binary encoding, is an entropy encoding generally used for a uniform probability distribution with a finite alphabet. The exponential Golomb code is a type of universal code.
[0725] In an embodiment, the bitstream is a bitstream within a scope. In an embodiment, the bitstream within the scope is a sequence of access units (AUs) composed of the first AU including an SDI SEI message and zero or more subsequent AUs following it (subsequent AUs including another SDI SEI message are not included) in the decoding order.
[0726] In an embodiment, the multi-view information SEI message and the auxiliary information SEI message do not exist in the coded video sequence (CVS) unless the SDI SEI message exists in the CVS.
[0727] In an embodiment, the multi-view information SEI message includes a multi-view acquisition information SEI message. In an embodiment, the auxiliary information SEI message includes a depth representation information SEI message. In an embodiment, the auxiliary information SEI message includes an alpha channel information SEI message.
[0728] In an embodiment, when a multi-view information SEI message or an auxiliary information SEI message exists in the bitstream, one or more of the SDI multi-view information flag (e.g., sdi_multiview_info flag) and the SDI auxiliary information flag (e.g., sdi_auxiliary_info_flag) are equal to 1. The flag is a variable or a single-bit syntax element that takes one of two possible values, 0 and 1.
[0729] In an embodiment, the multi-view information SEI message includes a multi-view acquisition information SEI message, and the multi-view acquisition information SEI message is not scalable nested. A scalable nested SEI message is an SEI message within a scalable nested SEI message. A scalable nested SEI message is a message that includes one or more output layer sets or a plurality of scalable nested SEI messages corresponding to one or more layers within a multi-layer bitstream.
[0730] In an embodiment, an SEI message in a bitstream with a payload type equal to 179 is restricted from being included in a scalable nesting SEI message. In an embodiment, an SEI message in a bitstream with a payload type equal to 3, 133, 179, 180, or 205 is restricted from being included in a scalable nesting SEI message.
[0731] In an embodiment, method 900 can utilize or incorporate one or more features or steps of other methods disclosed herein.
[0732] Next, a list of preferred solutions in some embodiments is shown.
[0733] The following solutions show examples of embodiments of the technology discussed in this disclosure (e.g., Example 1).
[0734] 1. A video processing method, comprising: performing a conversion between a video and a bitstream of the video; the bitstream conforming to format rules; the format rules specifying that a syntax element indicates a length obtained by subtracting L from the length of a syntax element of a view identifier, where L is an integer.
[0735] 2. The method according to claim 1, wherein the syntax element is coded as an unsigned integer using N bits.
[0736] 3. The method according to any one of claims 1 to 2, wherein L is a positive integer.
[0737] 4. The method according to claim 1, wherein L = 0 and it is prohibited for the syntax element to have a zero value.
[0738] 5. A video processing method, comprising the step of performing conversion between a video including a plurality of layers and a bitstream of the video, the bitstream conforming to a format rule, the format rule specifying that the bitstream includes auxiliary layers associated with one or more associated layers of the video.
[0739] 6. The method according to claim 5, wherein the format rule further specifies whether or how the bitstream includes one or more syntax elements indicating a relationship between the auxiliary layer and the one or more associated layers, and the one or more syntax elements are included in a scalability dimension auxiliary extension information syntax structure.
[0740] 7. The method according to claim 6, wherein the format rule specifies that the one or more associated layers are indicated by corresponding layer identifiers (IDs).
[0741] 8. The method according to claim 6, wherein the format rule specifies that the one or more associated layers are indicated by corresponding layer indices.
[0742] 9. The method according to any one of claims 5 to 8, wherein the format rule specifies that the bitstream includes one or more syntax elements indicating whether the auxiliary layer is applicable to the one or more associated layers.
[0743] 10. The method according to claim 9, wherein the one or more syntax elements include a syntax element indicating that the auxiliary layer is applicable to all of the one or more associated layers.
[0744] 11. The method according to claim 9, wherein the format rule specifies that each associated layer includes a syntax element indicating whether the auxiliary layer is applicable to the corresponding associated layer.
[0745] 12. The method according to claim 11, wherein the syntactic element indicates all primary layers related to the auxiliary layer.
[0746] 13. The method according to claim 11, wherein the syntactic element indicates all primary layers related to the auxiliary layer and having a layer index smaller than the layer index of the auxiliary layer.
[0747] 14. The method according to claim 11, wherein the syntactic element indicates all primary layers related to the auxiliary layer and having a layer index larger than the layer index of the auxiliary layer.
[0748] 15. The method according to any one of claims 11 to 14, wherein the syntactic element is a flag.
[0749] 16. The method according to claim 6, wherein the formatting rule specifies that the bitstream does not include an explicit syntactic element indicating the applicability of the auxiliary layer to the one or more related layers, and the applicability is derived during conversion.
[0750] 17. The method according to claim 16, wherein the formatting rule defines that the related layer of the auxiliary layer has a layer ID obtained by adding N1, N2... Nk to the layer ID of the auxiliary layer, where k is an integer and for i = 1,... k, two Ni are not equal to each other.
[0751] 18. The method according to claim 17, wherein k = 1 and N1 is any one of 1, -1, 2 or -2.
[0752] 19. The method according to claim 17, wherein k is greater than 1.
[0753] 20. The method according to claim 19, wherein k is equal to 2, N1 = 1, and N2 = 2.
[0754] 21. The method according to claim 5, wherein the formatting rule further specifies that the bitstream omits one or more syntax elements indicating a relationship between the auxiliary layer and the one or more associated layers, and the relationship is derived based on a predetermined rule.
[0755] 22. The method according to claim 5, wherein the formatting rule further defines that the bitstream includes one or more syntax elements indicating a relationship between the auxiliary layer and the one or more associated layers, and the one or more syntax elements are included in an auxiliary information auxiliary extension information syntax structure.
[0756] 23. The method according to any one of claims 5 to 22, wherein the formatting rule specifies that the bitstream includes a syntax element indicating the number of associated layers of the auxiliary picture of the layer.
[0757] 24. The method according to any one of claims 5 to 22, wherein the formatting rule specifies that when a condition is satisfied, the bitstream includes a syntax element indicating the number of associated layers of the auxiliary picture of the associated layer of the layer or the auxiliary picture.
[0758] 25. The method according to claim 24, wherein the condition includes that the i-th layer in the bitstream InScope includes an auxiliary picture.
[0759] 26. A video processing method, which performs conversion between a video including a plurality of video layers and a bitstream of the video, the bitstream conforms to a formatting rule, and the formatting rule specifies that the bitstream includes a multi-view auxiliary extension information (SEI) message or an auxiliary information SEI message according to whether a scalable dimension information SEI message is included in the encoded video sequence of the bitstream.
[0760] 27. The method according to claim 26, wherein the format rule specifies that the multi-view information SEI message refers to the multi-view acquisition information SEI message.
[0761] 28. The method according to any one of claims 26 to 27, wherein the format rule specifies that the auxiliary information SEI message refers to the depth representation information SEI message or the alpha channel information SEI message.
[0762] 29. A video processing method, comprising: executing conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream conforms to a format rule, and the format rule stipulates that at least one of a first flag indicating the presence of multi-view information in the scalability dimension information SEI message or a second flag indicating the presence of auxiliary information is equal to 1 in response to the presence of a multi-view or auxiliary information auxiliary extension information (SEI) message in the bitstream.
[0763] 30. A video processing method, comprising: executing conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream conforms to a format rule, and the format rule specifies that the multi-view acquisition information auxiliary extension information message included in the bitstream is not scalable nested or is included in the scalable nesting auxiliary extension information message.
[0764] 31. The method according to any one of claims 1 to 30, wherein the conversion includes generating the video from the bitstream or generating the bitstream from the video.
[0765] 32. A method for storing a bitstream in a computer-readable medium, the method comprising generating a bitstream according to any one or more of claims 1 to 31, and storing the bitstream in the computer-readable medium.
[0766] 33. A computer-readable medium having a bitstream of video, wherein when the bitstream is processed by a processor of a video decoder, the video decoder is caused to generate the video, and the bitstream is generated according to any one or more of claims 1 to 31.
[0767] 34. A video decoding apparatus comprising a processor configured to execute the method according to any one or more of claims 1 to 31.
[0768] 35. A video encoding apparatus comprising a processor configured to execute the method according to any one or more of claims 1 to 31.
[0769] 36. A computer program product having stored computer code, which when executed by a processor causes the processor to perform the method according to any one of claims 1 to 31.
[0770] 37. A computer-readable medium on which a bitstream conforming to a bitstream format generated according to any one of claims 1 to 31 is recorded.
[0771] 38. A bitstream generated according to the method, apparatus, disclosed method or system described herein.
[0772] The following documents may contain additional details related to the technology disclosed herein:
[0773] [1] ITU-T and ISO / IEC, "High efficiency video coding", Rec. ITU-T H.265 | ISO / IEC 23008-2 (current version).
[0774] [2] J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, J. Boyce, "Algorithm description of Joint Exploration Test Model 7 (JEM7)", JVET-G1001, Aug. 2017.
[0775] [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, "Versatile Video Coding", 2020.
[0776] [4] B. Bross, J. Chen, S. Liu, Y.-K. Wang (editors), "Versatile Video Coding (Draft 10)", JVET-S2001.
[0777] [5] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, "Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams", 2020.
[0778] [6] J. Boyce, V. Drugeon, G. Sullivan, Y.-K. Wang (editors), "Versatile supplemental enhancement information messages for coded video bitstreams (Draft 5)", JVET-S2007.
[0779] Other solutions, examples, embodiments, modules, and functional operations disclosed herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that constructs an execution environment for the computer program, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to an appropriate receiver device.
[0780] A computer program (also known as a program, software, software application, script, or coding) can be written in any form of programming language, including compiled languages or interpreted languages, and can be deployed in any form as a stand-alone program, module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), a single file dedicated to the program, or multiple coordinated files (such as files that store one or more modules, subprograms, or parts of the code). A computer program can be deployed to be executed on one computer, or deployed to be executed on multiple computers located at one site, or deployed to be executed on multiple computers distributed across multiple sites and interconnected by a communication network.
[0781] The processes and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs that operate on input data and generate output to perform functions. The processes and logical flows can also be executed by, and a device can also be implemented with, special-purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0782] Processors suitable for the execution of a computer program include, by way of example, any one or more processors of general and special purpose microprocessors, and any of various kinds of digital computers. In general, a processor receives instructions and data from a read only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. In general, a computer is operatively coupled to receive data from, transfer data to, or both transfer and receive data from one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, and optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing the instructions and data of a computer program include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices (such as, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and compact disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0783] Although this patent document describes many specific details, these should not be construed as limiting the scope of any subject matter or the claims that may be made, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Specific features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable partial combination in multiple embodiments. Further, even if a feature is described above as acting in a particular combination and is initially claimed as such, one or more features may be deleted from the claimed combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[0784] Similarly, although the operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the operations shown be performed, in order to achieve a desirable result. Further, the separation of various system components in the embodiments described in this patent document should not be understood as necessary in all embodiments.
[0785] Only a few implementations and examples are described, and other implementations, extensions, and variations are possible based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, performing conversion between a video and a bitstream of the video according to rules, the rules specifying that an auxiliary enhancement information (SEI) message having a payload type equal to 179 is not included in a scalable nested SEI message, the rules further specifying that when the coded video sequence (CVS) of the video does not include a scalability dimension information (SDI) SEI message, the CVS does not include a multi-view acquisition information (MAI) SEI message, method.
2. The SEI message having the payload type equal to 179 is the multi-view acquisition information (MAI) SEI message, The method according to claim 1.
3. The MAI SEI message is not scalable nested, The method according to claim 2.
4. the rules specifying that the length of a first syntax element indicating a view identifier of a current layer within the current coded video sequence (CVS) of the video is based on a value of a second syntax element within the scalability dimension information (SDI) SEI message, adding 1 to the second syntax element specifies the length of the first syntax element in bits, The method according to claim 1.
5. the second syntax element is coded as an unsigned integer using N bits, where N is an integer, The method according to claim 4.
6. N is equal to 4, The method according to claim 5.
7. The first syntax element is specified as sdi_view_id_val[], The second syntax element is specified as sdi_view_id_len_minus1, The method according to claim 4.
8. The first syntax element and the second syntax element are conditionally included in the bitstream based on the value of a third syntax element indicating whether the current CVS has multiple views, The method according to claim 4.
9. When the value of the third syntax element is equal to 1, the first syntax element and the second syntax element are included in the SDI SEI message of the bitstream, When the value of the third syntax element is equal to 0, the first syntax element and the second syntax element are not included in the SDI SEI message of the bitstream, The method according to claim 8.
10. The bitstream is a bitstream within the scope, The method according to claim 1.
11. The rule further specifies that an SEI message having a payload type equal to 3, 133, 180, or 205 is not included in the scalable nested SEI message, The method according to claim 1.
12. The conversion includes encoding the video into the bitstream, The method according to any one of claims 1 to 11.
13. The conversion includes decoding the video from the bitstream, The method according to any one of claims 1 to 11.
14. An apparatus for processing video data, comprising a processor and a non-transitory memory storing instructions that cause the processor to perform the following when executed by the processor: performing a conversion between a video and a bitstream of the video according to rules; the rules specify that an auxiliary enhancement information (SEI) message having a payload type equal to 179 is not included in a scalable nested SEI message; the rules further specify that when the coded video sequence (CVS) of the video does not include a scalability dimension information (SDI) SEI message, the CVS does not include a multi-view acquisition information (MAI) SEI message; apparatus.
15. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the following: performing a conversion between a video and a bitstream of the video according to rules; the rules specify that an auxiliary enhancement information (SEI) message having a payload type equal to 179 is not included in a scalable nested SEI message; the rules further specify that when the coded video sequence (CVS) of the video does not include a scalability dimension information (SDI) SEI message, the CVS does not include a multi-view acquisition information (MAI) SEI message; non-transitory computer-readable storage medium.
16. A method for storing a bitstream of a video, comprising: generating the bitstream of the video according to rules; storing the bitstream of the video in a non-transitory computer-readable recording medium; the rules specify that an auxiliary enhancement information (SEI) message having a payload type equal to 179 is not included in a scalable nested SEI message; The rule further specifies that when the coded video sequence (CVS) of the video does not include a scalability dimension information (SDI) SEI message, the CVS does not include a multi-view acquisition information (MAI) SEI message. Method.