Utilization of scalability dimension information
By using SDI SEI messages to specify primary and auxiliary layer associations, the ambiguity in existing video coding standards is resolved, enhancing clarity and efficiency in video encoding and decoding processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-02
- Publication Date
- 2026-04-01
AI Technical Summary
Existing video coding standards lack clarity on which primary layer is associated with auxiliary layers when auxiliary information is present in the bitstream, leading to inefficiencies and ambiguity in scalability dimension information (SDI) SEI messages.
Utilize Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages to identify which primary layer is associated with an auxiliary layer, by including layer identifiers and syntactic elements that specify the association, enabling clear conversion between video media files and bitstreams.
Enhances clarity and efficiency in video coding by clearly defining layer associations, facilitating effective conversion processes and improving scalability in video encoding and decoding.
Smart Images

Figure 0007839188000346 
Figure 0007839188000347 
Figure 0007839188000348
Abstract
Description
[Technical Field]
[0001] (Cross-reference of related applications) Book The application was filed on April 2, 2021. International patent application PCT / CN2021 / 085292 Based on International Patent Application PCT / CN2022 / 085030, filed on 2 April 2022, claiming priority and interest. All of the aforementioned patent applications are, By reference The whole This specification is incorporated herein.
[0002] This disclosure generally relates to video coding, and more particularly to the use of Supplemental Enhancement Information (SEI) messages for transmitting scalability dimension information in image / video coding. [Background technology]
[0003] On the internet and other digital communication networks, digital video accounts for the largest share of bandwidth usage. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to rise. [Overview of the project]
[0004] The disclosed embodiments provide a technique for identifying which primary layer (or non-auxiliary layer) is associated with an auxiliary layer by utilizing Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages when auxiliary information is present in the bitstream.
[0005] The first aspect relates to a method for processing video data. This method includes using Scalability Dimension Information (SDI) Supplementary Extension Information (SEI) messages to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream, and performing a conversion between a video media file and a bitstream based on the SDI SEI messages.
[0006] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that, if auxiliary information is present in the bitstream, one or more syntactic elements in the SDI SEI message indicate which primary layer is associated with the auxiliary layer.
[0007] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that an auxiliary layer has a layer identifier (ID) specifying sdi_aux_id[i], where an auxiliary layer identifier equal to zero indicates that the i-th layer in the bitstream does not contain an auxiliary picture, and an auxiliary layer identifier greater than zero indicates that the i-th layer in the bitstream contains an auxiliary picture of type.
[0008] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that a layer index is included in the SDI SEI message to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream.
[0009] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that one or more syntactic elements of the SDI SEI message indicate whether or not an auxiliary layer is applied to one or more primary layers.
[0010] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that the syntactic elements of the SDI SEI message indicate whether or not an auxiliary layer is applied from a primary layer to a particular primary layer.
[0011] If necessary, in any of the embodiments described above, other embodiments of the embodiment specify that the syntactic elements of the SDI SEI message indicate whether or not an auxiliary layer is applied to one or more primary layers.
[0012] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that when the auxiliary layer is one of a plurality of auxiliary layers in the bitstream and the auxiliary information exists in the bitstream, one or a group of syntax elements for indicating which primary layer is associated with each auxiliary layer within the plurality of auxiliary layers are included in the SDI SEI message.
[0013] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that the indication of the number of primary layers associated with the auxiliary picture of the auxiliary layer is signaled in the bitstream.
[0014] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that the indication of the number of primary layers is specified as sdi_num_associated_primary_layers_minus1.
[0015] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that sdi_num_associated_primary_layers_minus1 is signaled as a 6-bit unsigned integer.
[0016] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that the indication of the number of primary layers associated with the auxiliary layer or associated with the auxiliary picture of the auxiliary layer is signaled conditionally in the bitstream.
[0017] Optionally, in any of the foregoing aspects, other embodiments of the aspect provide that the bitstream constitutes a bitstream within a scope, and the conditional signal signals an indication of the number of primary layers only when the i-th layer within the bitstream within the scope includes an auxiliary picture.
[0018] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that the i-th layer in the bitstream in scope contains an auxiliary picture when the layer identifier (ID) specifying sdi_aux_id[i] is greater than 0.
[0019] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that the bitstream consists of a bitstream in scope, and the bitstream in scope is a sequence of access units (AUs) in decoding order, which does not include subsequent AUs containing another SDI SEI message, but which include an initial AU containing an SDI SEI message followed by zero or more subsequent AUs.
[0020] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that if auxiliary information is present in the bitstream, or if the bitstream constitutes a scoped bitstream and the scoped bitstream is a multiview bitstream, the SDI SEI message includes an auxiliary identifier (ID) for each layer.
[0021] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that if the layer identifier (ID) specifying sdi_aux_id[i] is equal to zero, the i-th layer is referred to as the primary layer, and otherwise, the i-th layer is referred to as the auxiliary layer.
[0022] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that when the layer identifier (ID) specifying sdi_aux_id[i] is equal to 1, the i-th layer is referred to as an alpha auxiliary layer, and when the layer ID specifying sdi_aux_id[i] is equal to 2, the i-th layer is referred to as a depth auxiliary layer.
[0023] If necessary, in any of the aforementioned embodiments, other embodiments of the embodiment provide that which primary layer is associated with an auxiliary layer is derived instead of being shown in the bitstream.
[0024] If necessary, in any of the embodiments described above, other embodiments of the embodiment provide that, if auxiliary information is present in the bitstream, auxiliary auxiliary extended information messages indicate which primary layer is associated with the auxiliary layer.
[0025] If necessary, in any of the aforementioned embodiments, other embodiments of the embodiment provide that the conversion encodes a video media file into a bitstream.
[0026] If necessary, in any of the aforementioned embodiments, other embodiments of the embodiment provide that the conversion decodes the bitstream to obtain a video media file.
[0027] The second aspect is a device for coding video data, comprising a processor and a non-transient memory on which instructions are recorded, wherein, during execution by the processor, the instructions cause the processor to indicate which primary layer is associated with an auxiliary layer using Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages, if auxiliary information is present in the bitstream, and to perform conversion between a video media file and a bitstream based on the SDI SEI messages.
[0028] A third aspect is a non-transient computer-readable medium containing a computer program product for use in a coding device, the computer program product containing computer-executable instructions stored in the non-transient computer-readable medium, which, when executed by one or more processors, causes the coding device to use Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream, and to perform conversions between a video media file and a bitstream based on the SDI SEI messages.
[0029] A fourth aspect is a non-transient, computer-readable storage medium storing instructions to be executed by a processor, the instructions causing the processor to use Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream, and to perform conversions between a video media file and a bitstream based on the SDI SEI messages.
[0030] A fifth aspect is a non-transient, computer-readable recording medium for storing a bitstream of video generated by a method performed by a video processing device, the method comprising the steps of indicating which primary layer is associated with an auxiliary layer using Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages, if auxiliary information is present in the bitstream, and converting between a video media file and a bitstream based on the SDI SEI messages.
[0031] A sixth aspect is a method for storing a video bitstream, which includes the steps of: indicating which primary layer is associated with an auxiliary layer using Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages, if auxiliary information is present in the bitstream; generating a bitstream based on the SDI SEI messages; and storing the bitstream on a non-transient, computer-readable recording medium.
[0032] For the purpose of clarification, any one of the embodiments described above can be combined with any one or more of the other embodiments described above to constitute a new embodiment within the scope of this disclosure.
[0033] These and other features can be better understood from the following detailed description, which is taken in conjunction with the attached drawings and claims.
[0034] Next, in order to fully understand this disclosure, please refer to the following brief description in conjunction with the attached drawings and detailed description. [Brief explanation of the drawing]
[0035] [Figure 1] Figure 1 shows an example of multi-layer coding for spatial scalability. [Figure 2] Figure 2 shows an example of multilayer coding using an Output Layer Set (OLS). [Figure 3] Figure 3 shows an embodiment of a video bitstream. [Figure 4] Figure 4 is a block diagram showing an example of an image processing system. [Figure 5] Figure 5 is a block diagram of the video processing device. [Figure 6] Figure 6 is a block diagram showing an example of a video coding system. [Figure 7] Figure 7 is a block diagram showing an example of a video encoder. [Figure 8] Figure 8 is a block diagram showing an example of a video decoder. [Figure 9] Figure 9 shows a method for coding video data according to an embodiment of the present disclosure. [Modes for carrying out the invention]
[0036] While examples of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. This disclosure is not to be limited in any way to the embodiments, drawings, and techniques illustrated below, including the designs and embodiments illustrated and described herein, and may be modified within the scope of the appended claims to the full extent of their equivalents.
[0037] Video coding standards have primarily developed through the development of well-known International Telecommunication Union (ITU-T) and International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) standards. The ITU-T developed H.261 and H.263, while ISO / IEC developed Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC) standards. See ITU-T and ISO / IEC, "High efficiency video coding," Rec. ITU-T H.265 | ISO / IEC 23008-2 (published version). Since H.262, video coding standards have been based on a hybrid video coding structure of time prediction + transformation coding. In 2015, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET) to explore future video coding technologies that surpass HEVC. Since then, many new techniques have been adopted by JVET and incorporated into reference software named Joint Exploration Model (JEM). See J. Chen, E. Alshina, G. Sullivan, J.-R. Ohm, J. Boyce, "Algorithm description of Joint Exploration Test Model 7 (JEM7)", JVET-G1001, Aug. 2017. JVET was later renamed the Joint Video Experts Team (JVET) when the Versatile Video Coding (VVC) project was officially launched. VVC is a new coding standard that aims for a 50% bitrate reduction compared to HEVC, and was finalized at the 19th JVET meeting, which concluded on July 1, 2020. Rec.ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Coding”, 2020.
[0038] The VVC standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for the widest possible range of applications, including not only traditional uses such as television broadcasting, video conferencing, and playback from storage media, but also newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, content synthesis and merging of multiple coded video bitstreams, multiview video, enhanced layer coding, and viewport-adapted 360° immersive media. B.Bross, J.Chen, S.Liu, Y.Wang(editors), "Versatile Video Coding (Draft10)", JVET-S2001, Rec.ITU-TRec.H.274|ISO / IEC23002-7, "Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams'',2020,andJ.Boyce,V.Drugeon,G.Sullivan,Y.-K.Wang(editors),``Versatile Video Coding(Draft10)'',JVET-S2001,Rec.Wang(editors),``Versatile supplemental enhancement information messages for coded video bitstreams(Draft5),''JVET-S2007.
[0039] The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard recently developed by MPEG.
[0040] Figure 1 is a schematic diagram showing an example of layer-based prediction 100. Layer-based prediction 100 supports unidirectional interpretation and / or bidirectional interpretation, but can also be performed between pictures in different layers.
[0041] Layer-based prediction 100 is applied between pictures 111, 112, 113, and 114 in different layers, and between pictures 115, 116, 117, and 118. In the illustrated example, pictures 111, 112, 113, and 114 are part of layer N+1 132, and pictures 115, 116, 117, and 118 are part of layer N 131. Layers such as layer N 131 and / or layer N+1 132 are groups of pictures all associated with similar characteristics such as similar size, quality, resolution, signal-to-noise ratio, and capability. In the illustrated example, layer N+1 132 is associated with larger image sizes than layer N 131. Therefore, pictures 111, 112, 113, and 114 in layer N+1 132 have larger picture sizes (e.g., larger height and width, resulting in more samples) than pictures 115, 116, 117, and 118 in layer N 131 in this example. However, such pictures may be separated between layer N+1 132 and layer N 131 by other characteristics. Although only two layers, layer N+1 132 and layer N 131, are shown, a set of pictures can be separated into any number of layers based on relevant characteristics. Layers N+1 132 and N 131 may also be indicated by layer IDs. A layer ID is a data item associated with a picture that indicates that the picture is part of the layer to which it is shown. Thus, each picture 111-118 may be associated with a corresponding layer ID to indicate that the corresponding layer N+1 132 or layer N 131 contains the corresponding picture.
[0042] Pictures 111-118 in different layers 131-132 are configured to be displayed alternately. Thus, pictures 111-118 in different layers 131-132 can share the same time identifier (ID) and may be contained in the same access unit (AU) 106. As used herein, an AU is a set of one or more coded pictures associated with the same display time for output from a decoded picture buffer (DPB). For example, the decoder can decode and display picture 115 at the current display time if a smaller picture is desired, or the decoder can decode and display picture 111 at the current display time if a larger picture is desired. Thus, pictures 111-114 in the upper layer N+1 132 contain substantially the same image data as the corresponding pictures 115-118 in the lower layer N 131 (regardless of differences in picture size). Specifically, picture 111 contains substantially the same image data as picture 115, picture 112 contains substantially the same image data as picture 116, and so on.
[0043] Pictures 111-118 may be coded by referencing other pictures 111-118 in the same layer N 131 or N+1 132. Pictures coded by referencing other pictures in the same layer are brought about by interpretation 123, which fits unidirectional interpretation and / or bidirectional interpretation. Interpretation 123 is drawn with solid arrows. For example, picture 113 may be coded by employing interpretation 123 using one or two of pictures 111, 112, and / or 114 in layer N+1 132 as references, where one picture is referenced for unidirectional interpretation and / or two pictures are referenced for bidirectional interpretation. Furthermore, picture 117 may also be coded by employing interpretation 123, which uses one or two of pictures 115, 116, and / or 118 in layer N 131 as references, where one picture is referenced for one-way interpretation and / or two pictures are referenced for two-way interpretation. A picture may be referenced as a reference picture if it is used as a reference to another picture in the same layer when performing interpretation 123. For example, picture 112 may be a reference picture used to code picture 113 according to interpretation 123. Interpretation 123 is sometimes referred to as intra-layer prediction in a multi-layer context. Thus, interpretation 123 is a mechanism for coding a sample in the current picture by referencing a sample indicated in a reference picture that is different from the current picture, when the reference picture and the current picture are in the same layer.
[0044] Pictures 111-118 may be coded by referencing other pictures 111-118 on different layers. This process is known as inter-layer prediction 121 and is depicted by a dashed arrow. Inter-layer prediction 121 is a mechanism by which a sample of the current picture is coded by referencing a sample indicated by a reference picture where the current picture and the reference picture are on different layers, i.e., have different layer IDs. For example, a picture in a lower layer N 131 can be used as a reference picture to code the corresponding picture in a higher layer N+1 132. Specifically, picture 111 may be coded by inter-layer prediction 121 by referencing picture 115. In such a case, picture 115 is used as the inter-layer reference picture. The inter-layer reference picture is the reference picture used in inter-layer prediction 121. In most cases, the inter-layer prediction 121 is constrained to use only inter-layer reference pictures (or more) that are in the same AU 106 and are located in lower layers, such as picture 115, for the current picture, such as picture 111. If multiple layers (e.g., two or more) are available, the inter-layer prediction 121 can encode / decode the current picture based on multiple inter-layer reference pictures (or more) that are at a lower level than the current picture.
[0045] The video encoder employs layer-based prediction 100 to encode pictures 111-118 through many different combinations and / or permutations of inter-prediction 123 and inter-layer prediction 121. For example, picture 115 may be coded according to intra-prediction. Then, pictures 116-118 may be coded according to inter-prediction 123 by using picture 115 as a reference picture. Furthermore, picture 111 may be coded according to inter-layer prediction 121 by using picture 115 as an inter-layer reference picture. Pictures 112-114 may be coded according to inter-layer prediction 123 by using picture 111 as a reference picture. In this way, the reference picture can play the role of both a single-layer reference picture and an inter-layer reference picture for different coding mechanisms. By coding the upper layer N+1 132 picture based on the lower layer N 131 picture, the upper layer N+1 132 can avoid employing intra prediction, which is far less coding efficient than inter-layer prediction 123 and inter-layer prediction 121. Thus, the less coding efficient intra prediction may be limited to the smallest / lowest quality picture and, therefore, limited to coding the smallest amount of video data. Pictures used as reference pictures, and / or inter-layer reference pictures, may be indicated by entries in the reference picture list contained in the reference picture list structure.
[0046] Each AU106 in Figure 1 may contain multiple pictures. For example, one AU106 may contain pictures 111 and 115. Another AU106 may contain pictures 112 and 116. In fact, each AU106 is a set of one or more coded pictures associated with the same display time (e.g., the same time ID) for output from the decoded picture buffer (DPB) (e.g., for display to the user). Each access unit delimiter (AUD)108 is an indicator or data structure used to indicate the start of an AU (e.g., AU108) or a boundary between AUs.
[0047] The H.26x video coding family has provided support for scalability through a separate profile from the one used for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and qualitative scalability. In SVC, each macroblock (MB) of an auxiliary layer (EL) picture is signaled with a flag indicating whether the ELMB is predicted using blocks arranged in a row from the lower layers. Predictions from the arranged blocks can include textures, motion vectors, and / or coding modes. An implementation of SVC cannot directly reuse an unmodified H.264 / AVC execution in its design. The syntax and decoding process for EL macroblocks in SVC differs from that of H.264 / AVC.
[0048] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that supports spatial and qualitative scalability, Multiview HEVC (MV-HEVC) is an extension of HEVC / H.265 that supports multiview scalability, and 3DHEVC (3D-HEVC) is an extension of HEVC / H.264 that supports more advanced and efficient three-dimensional (3D) video coding than MV-HEVC. Note that temporal scalability is included as an integral part of the single-layer HEVC codec. The design of the multi-layer extension of HEVC employs the idea that the decoded picture used for inter-layer prediction only appears from the same AU and is treated as a Long-Term Reference Picture (LTRP), and is assigned a reference index in the reference picture list along with other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the value of the reference index to refer to the inter-layer reference picture(s) in the reference picture list(s).
[0049] In particular, both reference picture resampling and spatial scalability features require resampling of the reference picture or a portion thereof. Reference picture resampling (RPR) can be implemented at either the picture level or the coding block level. However, when RPR is referred to as a coding feature, RPR is a feature for single-layer coding. Nevertheless, from a codec design perspective, it is possible, or even preferable, to use the same resampling filter for both the RPR feature for single-layer coding and the spatial scalability feature for multi-layer coding.
[0050] Figure 2 shows an example of layer-based prediction 200 using an output layer set (OLS). Layer-based prediction 100 is compatible with unidirectional interpretation and / or bidirectional interpretation, but can also be performed between pictures in different layers. The layer-based prediction in Figure 2 is similar to that in Figure 1. Therefore, for brevity, a full explanation of layer-based prediction will not be repeated.
[0051] Some of the layers of the coded video sequence (CVS) 290 in Figure 2 are included in the OLS. An OLS is a set of layers in which one or more layers are designated as output layers. Output layers are the layers that are output by the OLS. Figure 2 shows three different OLSs, namely OLS1, OLS2, and OLS3. As illustrated, OLS1 includes layers N 231 and N+1 232. Layer N 231 includes pictures 215, 216, 217, and 218, and layer N+1 232 includes pictures 211, 212, 213, and 214. OLS2 includes layers N 231, N+1 232, N+2 233, and N+3 234. Layer N+2 233 includes pictures 241, 242, 243, and 244, and Layer N+3 234 includes pictures 251, 252, 253, and 254. OLS3 includes layer N 231, layer N+1 232, and layer N+2 233. Although three OLSs are shown, different numbers of OLSs may be used in actual applications. In the illustrated embodiments, none of the OLSs include layer N+4 235, which includes pictures 261, 262, 263, and 264.
[0052] Different OLSs can each contain any number of layers. Different OLSs are generated to accommodate the varying coding capabilities of different devices with different coding capabilities. For example, OLS1, which contains only two layers, is generated to accommodate a mobile phone with relatively limited coding capabilities. On the other hand, OLS2, which contains four layers, may be generated to accommodate a large-screen television capable of decoding higher layers than a mobile phone. OLS3, which contains three layers, may be generated to accommodate a personal computer, laptop computer, or tablet computer, which may be able to decode higher layers than a mobile phone but not the highest layers, such as a large-screen television.
[0053] The layers shown in Figure 2 can all be independent of each other; that is, each layer can be coded without using inter-layer prediction (ILP). In this case, the layer is referred to as a simulcast layer. One or more layers shown in Figure 2 may be coded using ILP. Whether a layer is a simulcast layer or whether some layers are coded using ILP may be signaled by a flag in the video parameter set (VPS). If some layers use ILP, the layer dependencies between layers are also signaled in the VPS.
[0054] In an embodiment, if multiple layers are simulcast layers, only one layer is selected for decoding and output. In an embodiment, if several layers use ILP, all layers (e.g., the entire bitstream) are specified to be decoded, and a specific layer among the layers is specified to be the output layer. The output layer or multiple layers can be, for example, 1) only the top layer, 2) all layers, or 3) the top layer plus a set of specified lower layers. For example, if a flag in the VPS outputs the top layer plus specified lower layers, then layers N+3 234 (top layer) and layers N 231 and N+1 232 (lower layers) will be output from OLS2.
[0055] Some of the layers shown in Figure 2 may be referred to as primary layers, while others may be referred to as auxiliary layers. For example, layers N 231 and N+1 232 may be referred to as primary layers, and layers N+2 233 and N+3 234 may be referred to as auxiliary layers. Auxiliary layers may be referred to as alpha auxiliary layers or depth auxiliary layers. If auxiliary information is present in the bitstream, primary layers may be associated with auxiliary layers.
[0056] Unfortunately, the existing standard has shortcomings. 1. Currently, the syntax element sdi_view_id_len is coded as u(4), and its value is required to be in the range of 0 to 15. This value specifies the bit length of the sdi_view_id_val[i] syntax element, which specifies the view ID of the i-th layer in the bitstream. However, the length of sdi_view_id_val[i] must not be 0.
[0057] 2. For example, if there is some auxiliary information in the bitstream, such as when an SDI SEI message (Scalability Dimension SEI message) is indicated by a Depth Representation Information SEI message or an Alpha Channel Information SEI message, it is unclear which non-auxiliary layer or primary layer that the auxiliary information applies to.
[0058] 3. The presence of multiview acquisition information SEI messages, depth representation information SEI messages, or alpha channel information SEI messages in the bitstream is meaningless, but the scalability dimension information SEI message is not present in the bitstream.
[0059] 4. The multi-view acquisition information SEI message contains information about all views present in the bitstream. Therefore, making it scalable nested, which is currently permitted, is pointless.
[0060] Disclosed herein are techniques for solving one or more of the aforementioned problems. For example, the disclosure provides a technique for using Scalability Dimension Information (SDI) Auxiliary Extension Information (SEI) messages to identify which first (or non-auxiliary) layer is associated with an auxiliary layer, when auxiliary information is present in the bitstream.
[0061] Figure 3 shows an embodiment of the video bitstream 300. As used herein, the video bitstream 300 may also be referred to as a coded video bitstream, a bitstream, or a variation thereof. As shown in Figure 3, the bitstream 300 consists of one or more of the following: Decoding Capability Information (DCI) 302, Video Parameter Set (VPS) 304, Sequence Parameter Set (SPS) 306, Picture Parameter Set (PPS) 308, Picture Header (PH) 312, Picture 314, and SEI Message 322. Each of the DCI 302, VPS 304, SPS 306, and PPS 308 may be generally referred to as a parameter set. In embodiments, other parameter sets not shown in Figure 3 may also be included in the bitstream 300, for example, an Adaptive Parameter Set (APS) is a syntactic structure containing syntactic elements that apply to zero or more slices determined by zero or more syntactic elements found in the slice header.
[0062] DCI302, sometimes referred to as the Decoded Parameter Set (DPS) or Decoded Parameter Set, is a syntactic structure containing syntactic elements that apply to the entire bitstream. DCI302 contains constant parameters for the lifetime of the video bitstream (e.g., bitstream 300), which can be translated into a session lifetime. DCI302 includes profile, level, and subprofile information, allowing for the determination of a maximum complexity interlop point that is guaranteed never to be exceeded, even if splicing of the video sequence occurs within a session. It also includes constraint flags, as needed, indicating that the video bitstream is constrained from using certain features, as indicated by the values of these flags. This allows for indicating that certain tools should not be used on the bitstream, enabling resource allocation in decoder implementations, etc. Like all parameter sets, DCI302 implies that it must exist when first referenced, be referenced by the very first picture in the video sequence, and be transmitted between the first Network Abstraction Layer (NAL) units in the bitstream. Multiple DCI302s can exist in a bitstream, but the values of their syntactic elements must not contradict each other when referenced.
[0063] VPS304 includes decoding dependencies or decoding information for constructing the auxiliary layer's reference picture set. VPS304 provides an overall perspective or view of the scalable sequence, including the types of operation points provided, the profile, hierarchy, and levels of the operation points, as well as other high-level properties of the bitstream that can be used as the basis for session negotiation and content selection.
[0064] In an embodiment where it is indicated that some layers use ILP, VPS304 indicates that the total number of OLS specified by VPS is equal to the number of layers, that the i-th OLS contains layers with layer indices from 0 to i, and that for each OLS, only the top layer within the OLS is output.
[0065] SPS306 contains data common to all pictures within a set of pictures (SOP). SPS306 is a syntactic structure containing syntactic elements that apply to zero or more CLVSs, as determined by the content of syntactic elements found in the PPS referenced by syntactic elements found in each picture header. In contrast, PPS308 contains data common to all pictures. PPS308 is a syntactic structure containing syntactic elements that apply to zero or more coded pictures, as determined by the syntactic elements found in each picture header (e.g., PH312).
[0066] DCI302, VPS304, SPS306, and PPS308 belong to different types of Network Abstraction Layer (NAL) units. A NAL unit is a syntactic structure that contains an indication of the type of data that follows (e.g., coded video data). NAL units are classified into Video Coding Layer (VCL) and non-VCLNAL units. VCLNAL units contain data representing the values of samples in a video picture, while non-VCLNAL units contain related additional information such as parameter sets (important data applicable to multiple VCLNAL units) and auxiliary extensions (timing information and other auxiliary extensions that may improve the usability of the decoded video signal but are not necessary for decoding the values of samples in the video picture).
[0067] In the embodiment, DCI302 is included in a non-VCLNAL unit designated as a DCINAL unit or a DPSNAL unit. That is, a DCINAL unit has a DCINAL unit type (NUT), and a DPSNAL unit has a DPSNUT. In the embodiment, VPS304 is included in a non-VCLNAL unit designated as a VPSNAL unit. Therefore, a VPSNAL unit has a VPSNUT. In the embodiment, SPS306 is a non-VCLNAL unit designated as an SPSNAL unit. Therefore, an SPSNAL unit has an SPSNUT. In the embodiment, PPS308 is included in a non-VCLNAL unit designated as a PPSNAL unit. Therefore, a PPSNAL unit has a PPSNUT.
[0068] PH312 is a syntactic structure containing syntactic elements that apply to all slices (e.g., slice 318) of a coded picture (e.g., picture 314). In embodiments, PH312 is a type of non-VCLNAL unit designated as a PHNAL unit. Thus, a PHNAL unit has a PHNUT (e.g., PH_NUT).
[0069] In the embodiment, the PHNAL unit associated with PH312 has a time ID and a layer ID. The time ID identifier indicates the position of the PHNAL unit relative to other PHNAL units in the bitstream (e.g., bitstream 300). The layer ID indicates the layer containing the PHNAL unit (e.g., layer 131 or layer 132). In the embodiment, the time ID is similar to, but different from, the picture order count (POC). The POC uniquely identifies each picture in sequence. In a single-layer bitstream, the time ID and POC are the same. In a multi-layer bitstream (see, for example, Figure 1), pictures within the same AU have different POCs but the time ID is the same.
[0070] In this embodiment, the PHNAL unit precedes the VCLNAL unit which contains the first slice 318 of the associated picture 314. This establishes an association between the PH312 and the slice 318 of the picture 314 associated with the PH312, without the need for the picture header ID to be signaled within the PH312 and referenced from the slice header 320. As a result, it can be inferred that all VCLNAL units between the two PH312s belong to the same picture 314, and that the picture 314 is associated with the first PH312 between the two PH312s. In this embodiment, the first VCLNAL unit following the PH312 contains the first slice 318 of the picture 314 associated with the PH312.
[0071] In this embodiment, the PHNAL unit follows a picture-level parameter set (e.g., PPS), or a higher-level parameter set such as DCI (also known as DPS), VPS, SPS, or PPS, each having both a time ID and a layer ID smaller than the time ID and layer ID of the PHNAL unit. As a result, these parameter sets are not repeated within the picture or access unit. Due to this order, PH312 can be resolved immediately. That is, parameter sets containing parameters related to the entire picture are placed before the PHNAL unit in the bitstream. Those containing parameters for only a portion of the picture are placed after the PHNAL unit.
[0072] In one alternative scenario, the PHNAL unit follows a picture-level parameter set and prefixed auxiliary extended information (SEI) messages, or a higher-level parameter set such as DCI (also known as DPS), VPS, SPS, PPS, APS, or SEI messages.
[0073] Picture 314 is an array of luminance samples in monochrome format, or two corresponding arrays of luminance samples and saturation samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0074] A picture 314 can be either a frame or a field. However, in a single CVS 316, either all picture 314s are frames, or all picture 314s are fields. A CVS 316 is a coded video sequence for all coded layer video sequences (CLVSs) within the video bitstream 300. In particular, if the video bitstream 300 contains one layer, the CVS 316 and CLVS are the same. The CVS 316 and CLVS are different only when the video bitstream 300 contains multiple layers (for example, as shown in Figures 1 and 2).
[0075] Each picture 314 contains one or more slices 318. A slice 318 is an integer number of complete tiles or an integer number of consecutive complete coding tree unit (CTU) rows within a tile of a picture (e.g., picture 314). Each slice 318 is exclusively contained within a single NAL unit (e.g., VCLNAL unit). A tile (not shown) is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture (e.g., picture 314). A CTU (not shown) is a coding tree block (CTB) of luminance samples, two corresponding CTBs of saturation samples in a picture with three sample arrays, or a CTB of samples in a monochrome picture, or a picture coded using three separate color planes and syntactic structures used to code the samples. A CTB (not shown) is an N×N block of samples for some value N such that dividing a component into CTBs is partitioning. A block (not shown) is an M×N (M columns × N rows) array of samples (e.g., pixels), or an M×N array of transformation coefficients.
[0076] In this embodiment, each slice 318 includes a slice header 320. The slice header 320 is the portion of the coded slice 318 that contains data elements relating to all tiles or CTU rows within the tile represented by slice 318. That is, the slice header 320 includes information about slice 318, such as the slice type and whether a reference picture is used.
[0077] The pictures 314 and their slices 318 constitute the data related to the image or video being encoded or decoded. Therefore, the pictures 314 and their slices 318 may simply be referred to as the payload or data carried in the bitstream 300.
[0078] Bitstream 300 also includes one or more SEI messages, such as SEI message 322, which contains supplementary information. SEI messages can contain various types of data that indicate the timing of video pictures, describe various characteristics of the coded video, or describe how the coded video can be used or extended. SEI messages can also contain arbitrary user-defined data. SEI messages do not affect the core decoding process but can indicate how the video is recommended to be post-processed or displayed. Other high-level characteristics of the video content, such as color space instructions for interpreting the video content, are communicated in Video Usability Information (VUI). With the development of new color spaces, such as high dynamic range and wide color gamut video, VUI identifiers indicating them have also been added.
[0079] In one embodiment, SEI message 322 is an SDI SEI message. SDI SEI messages are used to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream. For example, an SDI SEI message includes one or more syntactic elements 324 to indicate which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream. The following describes various SEI messages and the syntactic elements they contain.
[0080] Those skilled in the art will understand that bitstream 300 may contain other parameters and information.
[0081] To address the above challenges, the following methods are disclosed. These techniques should be considered examples to illustrate general concepts and should not be interpreted narrowly. Furthermore, these techniques may be applied individually or combined in any way.
[0082] Example 1
[0083] 1) To solve Problem 1, as an example, instead of signaling the length of the view ID syntax element via, for example, the syntax element sdi_view_id_len, we can signal the value obtained by subtracting L from the length (for example, L=1) via, for example, the syntax element sdi_view_id_len_minusL.
[0084] a. As an example, syntactic elements may also be coded as unsigned integers using N bits.
[0085] i. For example, N is equal to 4.
[0086] ii. Alternatively, the syntax may be coded as a fixed pattern bit sequence using N bits, or as a signed integer using N bits, or as a truncated binary, or as a K (e.g., K=0)-th exponential Golomb coding syntax element of a signed integer, or as an M (e.g., M=0)-th exponential Golomb coding syntax element of an unsigned integer.
[0087] b. As an alternative, for example, the length can still be signaled via the syntax element sdi_view_id_len, but with the constraint that the value of the syntax element cannot be equal to 0.
[0088] Example 2
[0089] 2) To solve problem 2, it is proposed that an auxiliary layer (i.e., a layer whose corresponding sdi_aux_id[i] is 1 or 2) can be applied to one or more related layers.
[0090] a. As an example, one or more syntactic elements indicating the associated layers of each auxiliary layer may be signaled in the Scalability Dimension Information SEI message.
[0091] i. As an example, related layers are specified by their layer IDs.
[0092] ii. In other examples, the relevant layer is specified by the layer index.
[0093] iii. In other examples, whether an auxiliary layer applies to one or more related layers may be indicated by one or more syntactic elements for the related layers.
[0094] 1. As an example, syntactic elements may be used to indicate whether an auxiliary layer should be applied to all related layers.
[0095] 2. As an example, syntactic elements may indicate whether an auxiliary layer applies to a particular related layer.
[0096] a. As an example, one or more primary layers are indicated by syntactic elements.
[0097] i. As an example, all primary layers may be represented by syntactic elements.
[0098] ii. As an example, only primary layers whose layer index is smaller than the layer index of the auxiliary layer may be indicated by the syntactic element.
[0099] iii. As an example, only primary layers whose layer index is greater than the layer index of the auxiliary layer may be indicated by the syntactic element.
[0100] b. As an example, syntactic elements are coded as flags.
[0101] b. Alternatively, it is proposed that one or more layers associated with each auxiliary layer may be derived without explicit signal notification.
[0102] i. As an example, each auxiliary layer's associated layer may be a layer having a nuh_layer_id which is the auxiliary layer's nuh_layer_id plus N1, N2, ..., Nk, where k is an integer and Ni!=Nj if i and j (i!=j) are in the range of 1 to k.
[0103] 1. For example, k is equal to 1, and N1 is equal to 1, or 2, or -1, or -2.
[0104] 2. For example, k is greater than 1.
[0105] a. For example, k is equal to 2, so N1=1 and N2=2.
[0106] ii. As an example, the associated layers of each auxiliary layer may be layers having layer indices obtained by adding N1, N2, ..., Nk to the layer index of the auxiliary layer, where k is an integer and Ni! = Nj if i and j (i!=j) are in the range of 1 to k.
[0107] 1. For example, k is equal to 1, and N1 is equal to 1, or 2, or -1, or -2.
[0108] 2. For example, k is greater than 1.
[0109] a. For example, k is equal to 2, so N1=1 and N2=2.
[0110] c. Alternatively, instructions for the associated layers of each auxiliary layer may be explicitly signaled as one or a group of syntactic elements in a scalability dimension information SEI message.
[0111] d. Alternatively, the display of the relevant layers of the auxiliary information SEI message (e.g., depth representation information or alpha channel information) may be explicitly signaled by one or more syntactic elements within the auxiliary information SEI message.
[0112] i. For example, the auxiliary information SEI message may refer to the depth representation information SEI message or the alpha channel information SEI message.
[0113] ii. As an example, one or more syntactic elements may indicate the layer ID value of the associated layer.
[0114] 1. As an example, the layer ID indicated by a syntactic element may be required to be less than or equal to the maximum layer ID, i.e., vps_layer_id[vps_max_layers_minus1] or vps_layer_id[sdi_max_layers_minus1].
[0115] iii. As an example, one or more syntactic elements may indicate the layer index value of the associated layer.
[0116] 1. As an example, the layer index indicated by a syntactic element may be required to be less than the maximum number of layers in the bitstream (e.g., sdi_max_layers_minus1+1 or vps_max_layers_minus1+1).
[0117] iv. As an example, a signal may be provided indicating whether one or more layers are associated with auxiliary layers.
[0118] 1. As an example, a single syntactic element may be used to specify whether an auxiliary information SEI message applies to all layers.
[0119] a. For example, if auxiliary_all_layer_flag is equal to X (where X is 1 or 0), it can be specified that auxiliary information SEI messages are applied to all relevant primary layers.
[0120] 2. As an example, one or more syntactic elements may be used to specify whether an auxiliary information SEI message applies to one or more layers.
[0121] a. As an example, N syntactic elements may be used to specify whether the auxiliary information SEI message applies to N layers.
[0122] i. For example, syntactic elements may be coded as flags using 1 bit.
[0123] b. As an example, a single syntactic element may be used to specify whether an auxiliary information SEI message applies to one or more layers.
[0124] i. As an example, a syntactic element may be an exponential Golomb coding of the Kth (for example, K=0).
[0125] ii. As an example, a syntactic element equal to 5 specifies that the auxiliary information SEI message applies to the 0th and 2nd layers, but not to the 1st layer.
[0126] 1. If the syntactic element is equal to 5, the auxiliary information SEI message is applied to the (N-1)th and (N-3)th layers, but not to the (N-2)th layer.
[0127] c. The above syntactic elements may be conditionally signaled, for example, only if the auxiliary information SEI message does not apply to all layers.
[0128] e. As an example, the indication of the number of auxiliary picture layers associated with a single layer may be signaled in a bitstream.
[0129] f. As an example, the above syntactic elements may be signaled using an unsigned integer using N bits, or a fixed pattern bit sequence using N bits, or a signed integer using N bits, or truncated binary, or a signed integer K (e.g., K=0)-th exponential Golomb coding syntactic element, or an unsigned integer M (e.g., M=0)-th exponential Golomb coding syntactic element.
[0130] g. As an example, the number of associated layers for an auxiliary picture and / or the number of associated layers for an auxiliary picture can be conditionally signaled only if, for example, the auxiliary picture is included in the i-th layer of bitstreamInScope (e.g., sdi_aux_id[i]>0). bitstreamInScope (also known as bitstream scope) is defined as a sequence of AUs consisting of the first AU containing the SDI SEI message and zero or more subsequent AUs (but not subsequent AUs containing other SDI SEI messages) in decoding order.
[0131] Example 3
[0132] 3) To address issue 3, a bitstream conformance requirement is added stating that CVS without scalability dimension information SEI messages must not contain multiview or auxiliary information SEI messages.
[0133] a. Furthermore, the multiview information SEI message may refer to the multiview acquisition information SEI message.
[0134] b. Furthermore, the auxiliary information SEI message may refer to the depth representation information SEI message or the alpha channel information SEI message.
[0135] c. Alternatively, if a multiview or auxiliary information SEI message is present in the bitstream, an additional bitstream conformance requirement is added: at least one of the sdi_multiview_info_flag and sdi_auxiliary_info_flag in the scalability dimension information SEI message must be equal to 1.
[0136] Example 4
[0137] 4) To solve Problem 4, as an example, a bitstream conformance requirement is added that multi-view acquisition information SEI messages must not be scalable nested.
[0138] a. Alternatively, it is stipulated that SEI messages with payloadType equal to 179 (multi-view acquisition) must not be included in scalable nesting SEI messages.
[0139] Below are some examples of the embodiments summarized above. Each embodiment is applicable to VVC. Most of the added or modified parts are shown in bold italics, and some of the deleted parts are shown in italics. There may be other editorial changes that are not highlighted.
[0140] Each scalability dimension SEI message syntax described below includes one or more syntactic elements. Syntactic elements are, for example, one or more values, flags, variables, phrases, instructions, indices, mappings, data elements, or combinations thereof included in the scalability dimension SEI message syntax disclosed herein. In embodiments, syntactic elements may be organized into groups of values, flags, variables, phrases, instructions, indices, mappings, and / or data elements.
[0141] Embodiment 1
[0142] [Table 1]
[0143] Scalability Dimension SEI Message Semantics
[0144] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0145] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0146] [ka]
[0147] [ka]
[0148] [ka]
[0149] [ka]
[0150] [ka]
[0151] [ka]
[0152] [Table 2]
[0153] Note 1: The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 1-128 to 159 is specified by means other than the sdi_aux_id value.
[0154] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. However, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159.
[0155] Embodiment 2
[0156] [Table 3] [Table 4]
[0157] Scalability Dimension SEI Message Semantics
[0158] The scalability dimension SEI message provides scalability dimension information for each layer in bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID for each layer if auxiliary information (such as depth and alpha value) is transmitted in one or more layers within bitstreamInScope.
[0159] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0160] [ka]
[0161] [ka]
[0162] [ka]
[0163] [ka]
[0164] [ka]
[0165] [ka]
[0166] [Table 5]
[0167] The interpretation of auxiliary pictures associated with sdi_aux_id within the range of Note 1-128 to 159 is specified by means other than the sdi_aux_id value.
[0168] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0169] Embodiment 3 Scalability Dimension SEI Message Syntax
[0170] [Table 6]
[0171] Scalability Dimension SEI Message Semantics
[0172] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0173] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0174] [ka]
[0175] [ka]
[0176] [ka]
[0177] [ka]
[0178] [ka]
[0179] [ka]
[0180] [Table 7]
[0181] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0182] For bitstreams conforming to this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, however, for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0183] Embodiment 4
[0184] [Table 8]
[0185] Scalability Dimension SEI Message Semantics
[0186] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0187] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0188] [ka]
[0189] [ka]
[0190] [ka]
[0191] [ka]
[0192] [ka]
[0193] [ka]
[0194] [ka]
[0195] [ka]
[0196] [Table 9]
[0197] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0198] For bitstreams conforming to this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0199] [ka]
[0200] Embodiment 5
[0201] [Table 10]
[0202] Scalability Dimension SEI Message Semantics
[0203] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0204] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0205] [ka]
[0206] [ka]
[0207] [ka]
[0208] [ka]
[0209] [ka]
[0210] [ka]
[0211] [Table 11]
[0212] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0213] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0214] [ka]
[0215] [ka]
[0216] Embodiment 6
[0217] [Table 12] [Table 13]
[0218] Scalability Dimension SEI Message Semantics
[0219] The Scalability Dimension SEI Message is the scalability dimension information of each layer within bitstreamInScope (defined below), and provides: 1) the view ID of each layer when bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID of each layer when auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0220] bitstreamInScope is a sequence of AUs that, in decoding order, consists of the AU containing the current Scalability Dimension SEI Message and zero or more subsequent AUs (excluding subsequent AUs containing the Scalability Dimension SEI Message).
[0221] [Chemical Structure]
[0222] [Chemical Structure]
[0223] [Chemical Structure]
[0224] [Chemical Structure]
[0225]
Chem.
[0226]
Chem.
[0227]
Table 14
[0228] Note: The interpretation of the auxiliary picture related to sdi_aux_id in the range of 1 - 128 to 159 is specified by means other than the sdi_aux_id value.
[0229] For a bit stream compliant with this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0230]
Chem.
[0231]
Chem.
[0232]
Chem.
[0233] Embodiment 7
[0234]
Table 15
[0235] Scalability Dimension SEI Message Semantics
[0236] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0237] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0238] [ka]
[0239] [ka]
[0240] [ka]
[0241] [ka]
[0242] [ka]
[0243] [ka]
[0244] [Table 17]
[0245] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0246] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0247] [ka]
[0248] [ka]
[0249] [ka]
[0250] [ka]
[0251] Embodiment 8
[0252] [Table 18] [Table 19]
[0253] Scalability Dimension SEI Message Semantics
[0254] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0255] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0256] [ka]
[0257] [ka]
[0258] [ka]
[0259] [ka]
[0260] [ka]
[0261] [ka]
[0262] [Table 20]
[0263] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0264] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0265] [ka]
[0266] [ka]
[0267] [ka]
[0268] Embodiment 9
[0269] [Table 21] [Table 22]
[0270] Scalability Dimension SEI Message Semantics
[0271] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0272] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0273] [ka]
[0274] [ka]
[0275] [ka]
[0276] [ka]
[0277] [ka]
[0278] [ka]
[0279] [Table 23]
[0280] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0281] For bitstreams conforming to this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159, however, for this version of the decoder, the value of sdi_aux_id[i] shall be in the range of 0 to 255.
[0282] [ka]
[0283] [ka]
[0284] [ka]
[0285] Embodiment 10
[0286] [Table 24] [Table 25]
[0287] Scalability Dimension SEI Message Semantics
[0288] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0289] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0290] [ka]
[0291] [ka]
[0292] [ka]
[0293] [ka]
[0294] [ka]
[0295] [ka]
[0296] [Table 26]
[0297] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0298] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, however, in this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0299] [ka]
[0300] [ka]
[0301] [ka]
[0302] [ka]
[0303] Embodiment 11
[0304] [Table 27]
[0305] Scalability Dimension SEI Message Semantics
[0306] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0307] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0308] [ka]
[0309] [ka]
[0310] [ka]
[0311] [ka]
[0312] [ka]
[0313] [ka]
[0314] [Table 28]
[0315] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0316] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0317] Embodiment 12
[0318] [Table 29]
[0319] Scalability Dimension SEI Message Semantics
[0320] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream; and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0321] bitstreamInScope is a sequence of AUs consisting of the AU containing the current scalability dimension SEI message and zero or more subsequent AUs (excluding subsequent AUs containing scalability dimension SEI messages) in decryption order.
[0322] [ka]
[0323] [ka]
[0324] [ka]
[0325] [ka]
[0326] [ka]
[0327] [ka]
[0328] [ka]
[0329] [Table 30]
[0330] Note 1 - The interpretation of auxiliary pictures associated with sdi_aux_id in the range of 128 to 159 is specified by means other than the sdi_aux_id value.
[0331] For bitstreams conforming to this specification, sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] is in the range of 0 to 2 or 128 to 159, but for this version of the decoder, the value of sdi_aux_id[i] is in the range of 0 to 255.
[0332] Embodiment 13
[0333] Depth Representation Information SEI Message
[0334] [Table 31] [Table 32] [Table 33]
[0335] Depth Representation Information SEI Message Semantics
[0336] The syntax elements of the Depth Representation Information (SEI) message specify various parameters for an auxiliary picture of type AUX_DEPTH, for the purpose of processing the decoded primary and auxiliary pictures before rendering them on a 3D display, such as in view compositing. Specifically, the depth or parallax range of the depth picture is specified.
[0337] If present, the Depth Representation Information SEI message must be associated with one or more layers whose sdi_aux_id value is equal to AUX_DEPTH. The following semantics apply individually to each nuh_layer_idtargetLayerId among the nuh_layer_id values to which the Depth Representation Information SEI message applies.
[0338] If present, depth representation information (SEI) messages can be included in any access unit. If present, it is recommended to include SEI messages for random access purposes in access units where nuh_layer_id is equal to targetLayerId and the coded picture is an Intra Random Access Picture (IRAP) picture.
[0339] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is an associated first picture, that picture is a picture in the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all values of j in the range of 0-2 and 4-15.
[0340] The information shown in the SEI message is applied in decoding order to all pictures with a nuh_layer_id equal to targetLayerId, except for the next picture associated with the applicable depth representation information SEI message, starting from the access unit containing the SEI message, and then to the earlier of the two pictures: targetLayerId or the end of the CLVS with a nuh_layer_id equal to targetLayerId.
[0341] [ka]
[0342] [ka]
[0343] [ka]
[0344] [ka]
[0345] [ka]
[0346] [ka]
[0347] The variable maxVal is set to equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is a value included in or inferred from the active SPS of the layer where nuh_layer_id is equal to targetLayerId.
[0348] [Table 34] [Table 35]
[0349] [ka]
[0350] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is useful when depth_representation_type is 1 or 3.
[0351] The variable in column x of Table Y2 is derived from the variables in columns s, e, n, and v of Table Y2 as follows: - Except when the value of e is in the range of 0 to 127, x is (-1) s *2 e-31 *(1+n÷2 v It is set to be equal to ). - Otherwise (e=0), x is (-1). s *2 -(30+v) Set *n to equal to n.
[0352] Note 1 - The above specifications are the same as those described in IEC 60559:1989.
[0353] [Table 36]
[0354] The DMin and DMax values, if present, have a ViewId equal to the ViewId of the auxiliary picture and are specified in units of the luminance sample width of the coded picture.
[0355] The units for the ZNear and ZFar values are the same, if any, but are not specified.
[0356] [ka]
[0357] [ka]
[0358] Note 2 - When depth_representation_type=3, the auxiliary picture includes nonlinearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample values from a nonlinear representation to a linear representation, i.e., uniformly quantized disparity values. The shape of this conversion is defined by a line segment approximation in nonlinear disparity space from two-dimensional linear disparity. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of additional nodes are transmitted in the form of deviations from the straight curve (depth_nonlinear_representation_model[i]). These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0359] The variable DepthLUT[i] for i in the range of 0 to maxVal is specified as follows: for(k=0;k<=depth_nonlinear_representation_num_minus1+1;k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1+2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k+1)) / (depth_nonlinear_representation_num_minus1+2) dev2=depth_nonlinear_representation_model[k+1](X) x1=pos1-dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x=Max(x1,0);x<=Min(x2,maxVal);x++) DepthLUT[x]=Clip3(0,maxVal,Round(((x-x1)*(y2-y1))÷(x2-x1)+y1)) }
[0360] When depth_representation_type=3, DepthLUT[dS] for all decoded luminance sample values dS in the auxiliary picture range from 0 to maxVal represents disparity uniformly quantized in the range from 0 to maxVal.
[0361] The syntactic structure specifies the values of the elements in the depth representation information SEI message.
[0362] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen, which represent floating-point values. If the syntax structure is contained within another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced with the variable names used when the syntax structure is contained within it.
[0363] [ka]
[0364] [ka]
[0365] [ka]
[0366] [ka]
[0367] Embodiment 14 Depth Representation Information SEI Message
[0368] [Table 37] [Table 38]
[0369] [Table 39]
[0370] Depth Representation Information SEI Message Semantics
[0371] The syntax elements of the depth representation information SEI message specify various parameters for an auxiliary picture of type AUX_DEPTH, for the purpose of decoding the primary and auxiliary pictures before rendering them on a 3D display such as a view composite. Specifically, the depth or parallax range of the depth picture is specified.
[0372] If present, the Depth Representation Information SEI message must be associated with one or more layers whose sdi_aux_id value is equal to AUX_DEPTH. The following semantics apply individually to each nuh_layer_id targetLayerId among the nuh_layer_id values to which the Depth Representation Information SEI message applies.
[0373] If present, depth representation information (SEI) messages can be included in any access unit. If present, it is recommended to include SEI messages for random access purposes in access units where the coded picture with nuh_layer_id equal to targetLayerId is an IRAP picture.
[0374] In an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is an associated first picture, that picture is a picture within the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0, and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all j values in the range of 0-2 and 4-15.
[0375] The information shown in the SEI message is applied to all pictures with a nuh_layer_id equal to targetLayerId, in decoding order, from the access unit containing the SEI message, except for the next picture associated with the applicable depth representation information SEI message, whichever comes first in decoding order: targetLayerId or the end of the CLVS with a nuh_layer_id equal to targetLayerId.
[0376] [ka]
[0377] [ka]
[0378] [ka]
[0379] [ka]
[0380] [ka]
[0381] [ka]
[0382] [ka]
[0383] The variable maxVal is set to equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is the value included in, or inferred from, the active SPS of the layer where nuh_layer_id is equal to targetLayerId.
[0384] [Table 40] [Table 41]
[0385] [ka]
[0386] Note 1 - The disparity_ref_view_id exists only when d_min_flag = 1 or d_max_flag = 1, and is useful when the value of depth_representation_type is 1 or 3.
[0387] The variables in the x column of Table Y2 are derived from the variables in the s, e, n, and v columns of Table Y2 as follows: - Except when the value of e is in the range of 0 to 127, x is (-1) s * 2 e-31 * (1 + n ÷ 2 v ) and is set equal to it. - Otherwise (when e = 0), x is (-1) s * 2 -(30+v) * n and is set equal to it.
[0388] Note 1 - The above specification is the same as that described in IEC60559:1989.
[0389]
Table 42
[0390] The values of DMin and DMax, if they exist, have a ViewId equal to the ViewId of the auxiliary picture and are specified in units of the luminance sample width of the coded picture.
[0391] The units of the values of ZNear and ZFar are the same if they exist, but are not specified.
[0392]
Drawing
[0393]
Drawing
[0394] Note 2 - When depth_representation_type=3, the auxiliary picture includes nonlinearly transformed depth samples. The variable DepthLUT[i] specified below is used to convert the decoded depth sample values from a nonlinear representation to a linear representation, i.e., uniformly quantized disparity values. The shape of this conversion is defined by a line segment approximation in a two-dimensional linear-nonlinear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of additional nodes are transmitted in the form of deviations from the straight curve (depth_nonlinear_representation_model[i]). These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals corresponding to the value of nonlinear_depth_representation_num_minus1.
[0395] The variable DepthLUT[i] is specified as follows when i is in the range of 0 to maxVal: for(k=0;k<=depth_nonlinear_representation_num_minus1+1;k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1+2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k+1)) / (depth_nonlinear_representation_num_minus1+2) dev2=depth_nonlinear_representation_model[k+1](X) x1=pos1-dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x=Max(x1,0);x<=Min(x2,maxVal);x++) DepthLUT[x]=Clip3(0,maxVal,Round(((x-x1)*(y2-y1)))÷(x2-x1)+y1)) }
[0396] When depth_representation_type=3, DepthLUT[dS] for all decoded luminance sample values dS in the auxiliary picture range from 0 to maxVal represents disparity uniformly quantized in the range from 0 to maxVal.
[0397] The syntactic structure specifies the values of the elements in the depth representation information SEI message.
[0398] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen, which represent floating-point values. If the syntax structure is contained within another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced with the variable names used when the syntax structure was contained within it.
[0399] [ka]
[0400] [ka]
[0401] [ka]
[0402] [ka]
[0403] Embodiment 15 Alpha Channel Information SEI Message
[0404] [Table 43]
[0405] Alpha Channel Information SEI Message Semantics
[0406] The Alpha Channel Information SEI message provides information about the alpha channel sample values coded in an auxiliary picture of type AUX_ALPHA and one or more associated first pictures, as well as post-processing applied to the decoded alpha plane.
[0407] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, if there is an associated first picture, that picture is a picture in the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all j values from 0 to 2 and from 4 to 15.
[0408] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA persist in output order until one or more of the following conditions are true: - In output order, the next picture whose nuh_layer_id is equal to nuhLayerIdA will be output. - CLVS, including auxiliary picture picA, terminates. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer whose nuh_layer_id is equal to nuhLayerIdA is terminated.
[0409] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id to which the alpha channel information SEI message is applied.
[0410] [ka]
[0411] [ka]
[0412] Let currPic be the picture associated with the alpha channel information SEI message. The semantics of the alpha channel information SEI message persist for the current layer in output order until one or more of the following conditions are true: - A new CLVS for the current layer will begin. - The bitstream ends. - In an access unit containing an alpha channel information SEI message where nuh_layer_id is equal to targetLayerId, picture picB, where nuh_layer_id is equal to targetLayerId, is output with PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic immediately after the call to the decoding process for the picture order count of picB, respectively.
[0413] [ka]
[0414] [ka]
[0415] [ka]
[0416] [ka]
[0417] [ka]
[0418] [ka]
[0419] [ka]
[0420] Note - If both alpha_channel_incr_flag and alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by alpha_channel_clip_type_flag should be applied first, followed by the modification specified by alpha_channel_incr_flag to obtain the interpreted sample values for the luminance samples of the auxiliary picture.
[0421] Embodiment 16 Alpha Channel Information SEI Message
[0422] [Table 44]
[0423] Alpha Channel Information SEI Message Semantics
[0424] The Alpha Channel Information SEI message provides information about the alpha channel sample value and post-processing applied to the decoded alpha channel sample coded in an auxiliary picture of type AUX_ALPHA and one or more associated first pictures.
[0425] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the associated primary picture (if any) is a picture in the same access unit where scalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to scalabilityId[LayerIdxInVps[nuhLayerIdB]][j], and sdi_aux_id[nuhLayerIdB] is equal to 0 for all j values from 0 to 2 and from 4 to 15.
[0426] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA persist in output order until one or more of the following conditions are true: - In output order, the next picture whose nuh_layer_id is equal to nuhLayerIdA will be output. - CLVS, including auxiliary picture picA, terminates. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer whose nuh_layer_id is equal to nuhLayerIdA is terminated.
[0427] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id to which the alpha channel information SEI message is applied.
[0428] [ka]
[0429] Let currPic be the picture to which the alpha channel information SEI message is associated. The semantics of the alpha channel information SEI message persist for the current layer in the output order until one or more of the following conditions are true: - A new CLVS for the current layer will begin. - The bitstream ends. - For an access unit containing an alpha channel information SEI message where nuh_layer_id is equal to targetLayerId, picture picB whose nuh_layer_id is equal to targetLayerId will output PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic immediately after the decoding process call for the picture order count of picB, respectively.
[0430] [ka]
[0431] [ka]
[0432] [ka]
[0433] [ka]
[0434] [ka]
[0435] [ka]
[0436] [ka]
[0437] [ka]
[0438] [ka]
[0439] Note - If both alpha_channel_incr_flag and alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by alpha_channel_clip_type_flag should be applied first, followed by the modification specified by alpha_channel_incr_flag to obtain the interpreted sample values for the luminance samples of the auxiliary picture.
[0440] Embodiment 17 Multiview acquisition information SEI message
[0441] [Table 45] [Table 46] [Table 47]
[0442] Semantics of Multiview Acquisition Information SEI Messages
[0443] [ka]
[0444] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id values to which the multiview acquisition information SEI message is applied.
[0445] If present, the multiview acquisition information SEI message applied to the current layer must be contained within the access unit containing the IRAP picture, which is the first picture in the CLVS of the current layer. The information signaled in the SEI message is applied to the CLVS.
[0446] [ka]
[0447] [ka]
[0448] [ka]
[0449] [ka]
[0450] If a multi-view acquisition information SEI message is included in a scalable nesting SEI message, the syntax elements sn_ols_flag and sn_all_layers_flag of the scalable nesting SEI message must be equal to 0.
[0451] The variable numViewsMinus1 is derived as follows: - If the multiview acquisition information SEI message is not included in the scalable nesting SEI message, numViewsMinus1 is set to 0. - Otherwise (when the multiview acquisition information SEI message is included in the scalable nesting SEI message), numViewsMinus1 is set to equal sn_num_layers_minus1.
[0452] Some of the views included in the multi-view acquisition information SEI message do not exist.
[0453] In the following semantics, index i indicates the syntactic elements and variables that apply to the layer where nuh_layer_id is equal to NestingLayerId[i].
[0454] External camera parameters are specified according to a right-handed coordinate system where the upper-left corner of the image is the origin, i.e., (0,0) coordinates, and the other corners of the image have non-negative coordinates. With these specifications, a 3D world point wP=[xyz] is mapped to a 2D camera point cP[i]=[uv 1] for the i-th camera as follows:
[0455]
number
[0456] Here, A[i] represents the intrinsic parameter matrix of the camera, and R -1 [i] represents the inverse of the rotation matrix R[i], T[i] represents the translation vector, and s (a scalar value) is an arbitrary scaling factor chosen such that the third coordinate of cP[i] is 1. The elements of A[i], R[i], and T[i] are determined according to the syntactic elements signaled in this SEI message, as specified below.
[0457] [ka]
[0458]
change
[0459]
change
[0460]
change
[0461]
change
[0462]
change
[0463]
change
[0464]
change
[0465]
change
[0466]
change
[0467]
change
[0468] The length of a mantissa_focal_length_y[i] syntactic element is variable and is determined as follows: - If exponent_focal_length_y[i] is 0, the length will be Max(0, prec_focal_length-30). - Otherwise (exponent_focal_length_y[i] is in the range of 0 to 63, mutually exclusive), the length is Max(0, exponent_focal_length_y[i]+prec_focal_length-31).
[0469] [ka]
[0470] [ka]
[0471] [ka]
[0472] [ka]
[0473] [ka]
[0474] [ka]
[0475] [ka]
[0476] [ka]
[0477] [ka]
[0478] [ka]
[0479] The eigenmatrix A[i] of the i-th camera is given by the following equation.
[0480]
number
[0481] prec_rotation_param is 2 -prec_rotation_param Specifies the exponent of the maximum allowable truncation error for r[i][j][k] given by . The value of prec_rotation_param should be in the range of 0 to 31.
[0482] [ka]
[0483] [ka]
[0484] [ka]
[0485] [ka]
[0486] The rotation matrix R[i] for the i-th camera is expressed as follows:
[0487]
number
[0488] [ka]
[0489] [ka]
[0490] [ka]
[0491] The translation vector T[i] of the i-th camera can be expressed as follows:
[0492]
number
[0493] The relationship between camera parameter variables and their corresponding syntactic elements is specified in Table ZZ. The components of the eigenmatrix and rotation matrix, along with the translation vector, are obtained from the variables specified in Table ZZ as the variable x, calculated as follows: - If e is in the range of 0 to 63, then x is (-1). s *2 e-31 *(1+n÷2 v It is set to be equal to ). - Otherwise (e=0), x is (-1). s *2 -(30+v) Set *n to equal to n.
[0494] Note - The above specifications are similar to those described in IEC60559:1989.
[0495] [Table 48]
[0496] Embodiment 18 Depth Representation Information SEI Message
[0497] [Table 49] [Table 50]
[0498] [Table 51]
[0499] Depth Representation Information SEI Message Semantics
[0500] The syntax elements of the depth representation information SEI message specify various parameters for an auxiliary picture of type AUX_DEPTH, for the purpose of decoding the primary and auxiliary pictures before rendering them on a 3D display such as a view composite. Specifically, the depth or parallax range of the depth picture is specified.
[0501] If a Depth Representation Information SEI message exists, it must be associated with one or more layers whose sdi_aux_id value is equal to AUX_DEPTH. The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id values to which the Depth Representation Information SEI message applies.
[0502] In this case, the depth representation information SEI message can be included in any access unit. If an SEI message is present, it is recommended that it be included for random access purposes in access units where the coded picture with a nuh_layer_id equal to targetLayerId is an IRAP picture.
[0503] [ka]
[0504] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is an associated first picture, that picture is a picture in the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all j values from 0 to 2 and from 4 to 15.
[0505] The information shown in the SEI message is applied from the access unit containing the SEI message to all pictures with a nuh_layer_id equal to targetLayerId, whichever comes first in the decoding order: until the next picture is excluded in the decoding order, or until the end of the CLVS for a nuh_layer_id equal to targetLayerId, associated with the depth representation information SEI message applied to targetLayerId.
[0506] [ka]
[0507] [ka]
[0508] [ka]
[0509] [ka]
[0510] [ka]
[0511] The variable maxVal is set to equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is the value included in or inferred from the active SPS of the layer where nuh_layer_id is equal to targetLayerId.
[0512] [Table 52] [Table 53]
[0513] [ka]
[0514] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is valid when the value of depth_representation_type is equal to 1 and 3.
[0515] The variable in column x of Table Y2 is derived from the variables in columns s, e, n, and v of Table Y2 as follows: - If the value of e is in the range of 0 to 127 (mutually exclusive), then x is (-1). s *2 e-31 *(1+n÷2 v It is set to be equal to ). - Otherwise (if e is equal to 0), x is (-1). s *2 -(30+v) Set *n to equal to n.
[0516] Note 1 - The above specifications are the same as those described in IEC 60559:1989.
[0517] [Table 54]
[0518] The DMin and DMax values are specified in units of the luminance sample width of the coded picture having a ViewId equal to the ViewId of the auxiliary picture, if present.
[0519] The units for ZNear and ZFar values are the same if they exist, but are not specified.
[0520] [ka]
[0521] [ka]
[0522] Note 2 - When depth_representation_type is equal to 3, the auxiliary picture includes nonlinearly transformed depth samples. The variable DepthLUT[i] specified below is used to transform the decoded depth sample values from a nonlinear representation to a linear representation, i.e., uniformly quantized disparity values. The shape of this transformation is defined by a line segment approximation in a two-dimensional linear disparity space-nonlinear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of additional nodes are transmitted in the form of deviations from the straight curve (depth_nonlinear_representation_model[i]). These deviations are uniformly distributed along the entire range from 0 to maxVal, at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0523] The variable DepthLUT[i], where i ranges from 0 to maxVal, is specified as follows: for(k=0;k<=depth_nonlinear_representation_num_minus1+1;k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1+2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k+1)) / (depth_nonlinear_representation_num_minus1+2) dev2=depth_nonlinear_representation_model[k+1](X) x1=pos1-dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x=Max(x1,0);x<=Min(x2,maxVal);x++) DepthLUT[x]=Clip3(0,maxVal,Round(((x-x1)*(y2-y1)))÷x2-x1)+y1)) }
[0524] If depth_representation_type is equal to 3, the DepthLUT[dS] for all decoded luminance sample values dS in the auxiliary picture range from 0 to maxVal represents disparity uniformly quantized in the range from 0 to maxVal.
[0525] The syntactic structure specifies the values of the elements in the depth representation information SEI message.
[0526] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen, which represent floating-point values. If the syntax structure is contained within another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced with the variable names used when the syntax structure was contained within it.
[0527] [ka]
[0528] [ka]
[0529] [ka]
[0530] [ka]
[0531] Embodiment 19 Depth Representation Information SEI Message
[0532] [Table 55] [Table 56]
[0533] [Table 57]
[0534] Depth Representation Information SEI Message Semantics
[0535] [ka]
[0536] [ka]
[0537] [ka]
[0538] [ka]
[0539] If present, depth representation information (SEI) messages can be included in any access unit. In access units where the coded picture with nuh_layer_id equal to targetLayerId is an IRAP picture, it is recommended to include SEI messages for random access purposes if they exist.
[0540] For an auxiliary picture where sdi_aux_id[targetLayerId] is equal to AUX_DEPTH, if there is an associated first picture, that picture is a picture in the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[targetLayerId]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all j values from 0 to 2 and from 4 to 15.
[0541] The information indicated in the SEI message is applied to the earlier of the following pictures in the decoding order, from the access unit containing the SEI message to all pictures with a nuh_layer_id equal to targetLayerId, excluding the next picture related to the depth representation information SEI message applied to targetLayerId, or until the end of the CLVS for the nuh_layer_id equal to targetLayerId.
[0542] [ka]
[0543] [ka]
[0544] [ka]
[0545] [ka]
[0546] [ka]
[0547] The variable maxVal is set to equal to (1 << (8 + sps_bitdepth_minus8)) - 1, where sps_bitdepth_minus8 is a value included in or inferred from the active SPS of the layer where nuh_layer_id is equal to targetLayerId.
[0548] [Table 58] [Table 59]
[0549] [ka]
[0550] Note 1 - disparity_ref_view_id exists only when d_min_flag is 1 or d_max_flag is 1, and is valid when the value of depth_representation_type is equal to 1 and 3.
[0551] The variable in column x of Table Y2 is derived from the variables in columns s, e, n, and v of Table Y2 as follows: - If the value of e is in the range of 0 to 127 (mutually exclusive), then x is (-1). s *2 e-31 *(1+n÷2 v It is set to be equal to ). - Otherwise (e=0), x is (-1). s *2 -(30+v) Set *n to equal to n.
[0552] Note 1 - The above specifications are similar to those described in IEC 60559:1989.
[0553] [Table 60]
[0554] The DMin and DMax values are specified in units of the luminance sample width of the coded picture having a ViewId equal to the ViewId of the auxiliary picture, if present.
[0555] The units for ZNear and ZFar values are the same if they exist, but are not specified.
[0556] [ka]
[0557] [ka]
[0558] Note 2 - When depth_representation_type is equal to 3, the auxiliary picture includes nonlinearly transformed depth samples. The variable DepthLUT[i] specified below is used to transform the decoded depth sample values from a nonlinear representation to a linear representation, i.e., uniformly quantized disparity values. The shape of this transformation is defined by a line segment approximation in a two-dimensional linear disparity space-nonlinear disparity space. The first node (0,0) and the last node (maxVal,maxVal) of the curve are predefined. The positions of additional nodes are transmitted in the form of deviations from the straight curve (depth_nonlinear_representation_model[i]). These deviations are uniformly distributed along the entire range from 0 to maxVal at intervals that depend on the value of nonlinear_depth_representation_num_minus1.
[0559] The variable DepthLUT[i], where i ranges from 0 to maxVal, is specified as follows: for(k=0;k<=depth_nonlinear_representation_num_minus1+1;k++){ pos1=(maxVal*k) / (depth_nonlinear_representation_num_minus1+2) dev1=depth_nonlinear_representation_model[k] pos2=(maxVal*(k+1)) / (depth_nonlinear_representation_num_minus1+2) dev2=depth_nonlinear_representation_model[k+1](X) x1=pos1-dev1 y1 = pos1 + dev1 x2 = pos2 - dev2 y2 = pos2 + dev2 for(x=Max(x1,0);x<=Min(x2,maxVal);x++) DepthLUT[x]=Clip3(0,maxVal,Round(((x-x1)*(y2-y1))÷x2-x1)+y1)) }
[0560] If depth_representation_type is equal to 3, the DepthLUT[dS] for all decoded luminance sample values dS in the auxiliary picture range from 0 to maxVal represents disparity uniformly quantized in the range from 0 to maxVal.
[0561] The syntactic structure specifies the values of the elements in the depth representation information SEI message.
[0562] The syntax structure sets the values of the variables OutSign, OutExp, OutMantissa, and OutManLen, which represent floating-point values. If the syntax structure is contained within another syntax structure, the variable names OutSign, OutExp, OutMantissa, and OutManLen are interpreted as being replaced with the variable names used when the syntax structure was contained within it.
[0563] [ka]
[0564] [ka]
[0565] [ka]
[0566] [ka]
[0567] Embodiment 20 Alpha Channel Information SEI Message
[0568] [Table 61]
[0569] Alpha Channel Information SEI Message Semantics
[0570] The Alpha Channel Information SEI message provides information about the alpha channel sample values coded in an auxiliary picture of type AUX_ALPHA and one or more associated first pictures, as well as post-processing applied to the decoded alpha plane.
[0571] For an auxiliary picture where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, if there is an associated first picture, that picture is a picture in the same access unit where sdi_aux_id[nuhLayerIdB] is equal to 0 and ScalabilityId[LayerIdxInVps[nuhLayerIdA]][j] is equal to ScalabilityId[LayerIdxInVps[nuhLayerIdB]][j] for all j values from 0 to 2 and from 4 to 15.
[0572] When an access unit contains an auxiliary picture picA where nuh_layer_id is equal to nuhLayerIdA and sdi_aux_id[nuhLayerIdA] is equal to AUX_ALPHA, the alpha channel sample values of picA persist in output order until one or more of the following conditions are true: - In output order, the next picture whose nuh_layer_id is equal to nuhLayerIdA will be output. - CLVS, including auxiliary picture picA, terminates. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer whose nuh_layer_id is equal to nuhLayerIdA is terminated.
[0573] [ka]
[0574] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id to which the alpha channel information SEI message is applied.
[0575] [ka]
[0576] Let currPic be the picture to which the alpha channel information SEI message is associated. The semantics of the alpha channel information SEI message persist for the current layer in the output order until one or more of the following conditions are true: - A new CLVS for the current layer will begin. - The bitstream ends. - For an access unit containing an alpha channel information SEI message where nuh_layer_id is equal to targetLayerId, picture picB whose nuh_layer_id is equal to targetLayerId will output PicOrderCnt(picB) greater than PicOrderCnt(currPic), where PicOrderCnt(picB) and PicOrderCnt(currPic) are the PicOrderCntVal values of picB and currPic immediately after the decoding process for the picture order count of picB, respectively.
[0577] [ka]
[0578] [ka]
[0579] [ka]
[0580] [ka]
[0581] [ka]
[0582] [ka]
[0583] [ka]
[0584] Note - If both alpha_channel_incr_flag and alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by alpha_channel_clip_type_flag should be applied first, followed by the modification specified by alpha_channel_incr_flag to obtain the interpreted sample values for the luminance samples of the auxiliary picture.
[0585] Embodiment 21 Alpha Channel Information SEI Message
[0586] [Table 62]
[0587] Alpha Channel Information SEI Message Semantics
[0588] [ka]
[0589] [ka]
[0590] [ka]
[0591] [ka]
[0592] [ka]
[0593] In output order, the next picture whose nuh_layer_id is equal to nuhLayerIdA will be output. - CLVS, including auxiliary picture picA, terminates. - The bitstream ends. - The CLVS of the associated primary layer of the auxiliary picture layer whose nuh_layer_id is equal to nuhLayerIdA is terminated.
[0594] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id to which the alpha channel information SEI message is applied.
[0595] [ka]
[0596] [ka]
[0597] [ka]
[0598] [ka]
[0599] [ka]
[0600] [ka]
[0601] [ka]
[0602] [ka]
[0603] [ka]
[0604] Note - If both alpha_channel_incr_flag and alpha_channel_clip_type_flag are equal to 1, the clipping operation specified by alpha_channel_clip_type_flag should be applied first, followed by the modification specified by alpha_channel_incr_flag to obtain the interpreted sample values for the luminance samples of the auxiliary picture.
[0605] Embodiment 22 Scalability Dimensional Information (SDI) SEI Message
[0606] [Table 63]
[0607] Scalability Dimension SEI Message Semantics
[0608] The scalability dimension SEI message provides scalability dimension information for each layer within bitstreamInScope (defined below), including: 1) the view ID for each layer if bitstreamInScope is a multi-view bitstream, and 2) the auxiliary ID for each layer if auxiliary information (such as depth or alpha value) is transmitted in one or more layers within bitstreamInScope.
[0609] bitstreamInScope is a sequence of zero or more AUs that, in decryption order, include the AU containing the current scalability dimension SEI message, but do not include subsequent AUs containing scalability dimension SEI messages.
[0610] [ka]
[0611] [ka]
[0612] [ka]
[0613] [ka]
[0614] [ka]
[0615] [ka]
[0616] [ka]
[0617] [Table 64]
[0618] Note 3-128 to 159: The interpretation of auxiliary pictures associated with sdi_aux_id is specified by means other than the sdi_aux_id value.
[0619] For a bitstream conforming to this specification, sdi_aux_id[i] shall be in the range of 0 to 2 or 128 to 159. The value of sdi_aux_id[i] must be in the range of 0 to 2 or 128 to 159, however, in this version of this specification, the decoder must allow sdi_aux_id[[i]] values in the range of 0 to 255.
[0620] Multiview acquisition information SEI message
[0621] [Table 65] [Table 66] [Table 67]
[0622] Semantics of Multiview Acquisition Information SEI Messages
[0623] [ka]
[0624] [ka]
[0625] [ka]
[0626] [ka]
[0627] [ka]
[0628] The following semantics apply individually to the targetLayerId of each nuh_layer_id among the nuh_layer_id to which the multiview acquisition information SEI message is applied.
[0629] If present, the multiview acquisition information SEI message applied to the current layer must be contained within the access unit containing the IRAP picture, which is the first picture of the CLVS for the current layer. The information signaled in the SEI message is applied to the CLVS.
[0630] [ka]
[0631] [ka]
[0632] Some of the views included in the multi-view acquisition information SEI message do not exist.
[0633] In the following semantics, index i refers to the syntactic elements and variables applied to the layer where nuh_layer_id is equal to NestingLayerId[i].
[0634] The camera's external parameters are specified according to a right-handed coordinate system where the upper-left corner of the image is the origin, i.e., (0,0) coordinates, and the other corners of the image have non-negative coordinates. With these specifications, a point wP=[xyz] in the 3D world is mapped to a point cP[i]=[uv 1] in the 2D camera for the i-th camera as follows:
[0635]
number
[0636] Here, A[i] represents the matrix of the camera's intrinsic parameters, and R -1 [i] represents the inverse of the rotation matrix R[i], T[i] represents the translation vector, and s (a scalar value) is an arbitrary scaling factor chosen such that the third coordinate of cP[i] is 1. The elements of A[i], R[i], and T[i] are determined according to the syntactic elements signaled in this SEI message, as specified below.
[0637] [ka]
[0638] [ka]
[0639] [ka]
[0640] [ka]
[0641] [ka]
[0642] [ka]
[0643] [ka]
[0644] [ka]
[0645] [ka]
[0646] [ka]
[0647] [ka]
[0648] [ka]
[0649] The length of a mantissa_focal_length_y[i] syntactic element is variable and is determined as follows: If exponent_focal_length_y[i] is 0, the length will be Max(0, prec_focal_length-30). - Otherwise (exponent_focal_length_y[i] is in the range of 0 to 63, mutually exclusive), the length is Max(0, exponent_focal_length_y[i]+prec_focal_length-31).
[0650] [ka]
[0651] [ka]
[0652] [ka]
[0653] [ka]
[0654] [ka]
[0655] [ka]
[0656] [ka]
[0657] [ka]
[0658] [ka]
[0659] [ka]
[0660] The eigenmatrix A[i] of the i-th camera is expressed by the following equation:
[0661]
number
[0662] prec_rotation_param is 2 -prec_rotation_param Specifies the exponent of the maximum allowable truncation error for r[i][j][k] given by . The value of prec_rotation_param should be in the range of 0 to 31.
[0663] [ka]
[0664] [ka]
[0665] [ka]
[0666] [ka]
[0667] The rotation matrix R[i] for the i-th camera is expressed as follows:
[0668]
number
[0669] [ka]
[0670] [ka]
[0671] [ka]
[0672] The translation vector T[i] of the i-th camera is expressed as follows:
[0673]
number
[0674] The relationship between camera parameter variables and their corresponding syntactic elements is specified in Table ZZ. The components of the eigenmatrix and rotation matrix, along with the translation vector, are obtained from the variables specified in Table ZZ as the variable x, calculated as follows: - If e is in the range of 0 to 63, then x is (-1) s *2 e-31 *(1+n÷2 v It is set to be equal to ). - Otherwise (if e is equal to 0), x is (-1) s *2 -(30+v) Set *n to equal to n.
[0675] Note - The above specifications are similar to those described in IEC60559:1989.
[0676] [Table 68]
[0677] Figure 4 is a block diagram showing an exemplary video processing system 400 in which various technologies disclosed herein may be implemented. Various embodiments may include some or all of the components of the video processing system 400. The video processing system 400 may include an input 402 for receiving video content. The video content may be received in raw or uncompressed format, for example, as 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. The input 402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet and passive optical networks (PON), and wireless interfaces such as Wi-Fi and cellular interfaces.
[0678] The video processing system 400 may include a coding component 404 capable of performing various coding or encoding methods described herein. The coding component 404 may reduce the average bitrate of the video from input 402 to the output of the coding component 404 to produce a coded representation of the video. Thus, coding techniques may also be referred to as video compression techniques or video transcoding techniques. The output of the coding component 404 is either stored or transmitted over a communication connection, as represented by component 406. The stored or transmitted bitstream (or coded) representation of the video received at input 402 may be used by component 408 to generate pixel values or a viewable video to be sent to the display interface 410. The process of generating a user-viewable video from the bitstream representation may be referred to as video restoration. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that coding tools or operations are used in encoders, and corresponding decoding tools or operations that reverse the coding result are performed in decoders.
[0679] Examples of peripheral bus interfaces and display interfaces include USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface) (registered trademark), and DisplayPort. Examples of storage interfaces include SATA (serial advanced technology attachment), PCI (Peripheral Component Interconnect), and IDE (Integrated Drive Electronics) interfaces. The technologies described herein can be implemented in a variety of electronic devices capable of performing digital data processing and / or video display, such as mobile phones, laptops, and smartphones.
[0680] Figure 5 is a block diagram of the image processing device 500. The device 500 can be used to carry out one or more methods described herein. The device 500 can be embodied in smartphones, tablets, computers, IoT (Internet of Things) receivers, etc. The device 500 may include one or more processors 502, one or more memories 504, and image processing hardware 506 (also known as image processing circuits). One or more processors 502 may be configured to carry out one or more methods described herein. The memories (storage devices) 504 may be used to store data and coding used to carry out the methods and techniques described herein. The image processing hardware 506 may be used to implement some of the techniques described herein in hardware circuits. In some embodiments, the hardware 506 may be partially or completely located on the processor 502, for example, a graphics processor, etc.
[0681] Figure 6 is a block diagram showing an exemplary video coding system 600 that may utilize the technology of this disclosure. As shown in Figure 6, the video coding system 600 may include a source device 610 and a destination device 620. The source device 610 generates coded video data, which may be referred to as a video coding device. The destination device 620 may decode the coded video data generated by the source device 610, which may be referred to as a video decoding device.
[0682] The source device 610 may include a video source 612, a video encoder 614, and an input / output (I / O) interface 616.
[0683] The video source 612 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may consist of one or more pictures. The video encoder 614 encodes the video data from the video source 612 to produce a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data may include a sequence parameter set, a picture parameter set, and other syntactic structures. The I / O interface 616 may include a modulator / demodulator (modem) and / or transmitter. The coded video data may be transmitted directly to the destination device 620 via the I / O interface 616 over the network 630. The coded video data may also be stored in a storage medium / server 640 for access by the destination device 620.
[0684] The destination device 620 may include an I / O interface 626, a video decoder 624, and a display device 622.
[0685] The I / O interface 626 may include a receiver and / or a modem. The I / O interface 626 may acquire encoded video data from the source device 610 or storage medium / server 640. The video decoder 624 may decode the encoded video data. The display device 622 may display the decoded video data to the user. The display device 622 may be integrated with the destination device 620, or it may be external to the destination device 620 and configured to connect to an external display device.
[0686] The video encoder 614 and video decoder 624 can operate in accordance with video compression standards such as the HEVC (High Efficiency Video Coding) standard, the VVC (Versatile Video Coding) standard, and other current and / or further standards.
[0687] Figure 7 is a block diagram showing an example of a video encoder 700, which may be the video encoder 614 in the video coding system 600 shown in Figure 6.
[0688] The video encoder 700 can be configured to perform any or all of the techniques of this disclosure. In the example in Figure 7, the video encoder 700 includes several functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 700. In some examples, the processor may be configured to perform any or all of the techniques described in this disclosure.
[0689] The functional components of the video encoder 700 may include a prediction unit 702 which may include a splitting unit 701, a mode selection unit 703, a motion estimation unit 704, a motion compensation unit 705, and an intra-prediction unit 706, a residual generation unit 707, a conversion unit 708, a quantization unit 709, an inverse quantization unit 710, an inverse conversion unit 711, a reconstruction unit 712, a buffer 713, and an entropy coding unit 714.
[0690] In other examples, the video encoder 700 may include more, fewer, or different functional components. In one example, the prediction unit 702 may include an intrablock copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture on which the current video block is located.
[0691] Furthermore, some components, such as the motion estimation unit 704 and the motion compensation unit 705, may be highly integrated, but in the example of Figure 7, they are shown separately for illustrative purposes.
[0692] The splitting unit 701 can divide the picture into one or more video blocks. The video encoder 614 and video decoder 624 in Figure 6 can support various video block sizes.
[0693] The mode selection unit 703 may, for example, select either an intra- or inter-coding mode based on the error result, and provide the resulting intra- or inter-coded blocks to the residual generation unit 707, which generates residual block data, and the reconstruction unit 712, which reconstructs the encoded blocks for use as a reference picture. In some examples, the mode selection unit 703 may select a combination of intra-prediction and inter-prediction (CIIP) modes, where the prediction is based on the inter-prediction signal and the intra-prediction signal. The mode selection unit 703 may also select the resolution of the motion vectors for the blocks (e.g., sub-pixel or integer pixel precision) in the case of inter-prediction.
[0694] To perform interpretation for the current video block, motion compensation unit 704 may generate motion information for the current video block by comparing one or more reference frames from buffer 713 with the current video block. Motion compensation unit 705 may determine the predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 713 other than the picture associated with the current video block.
[0695] The motion estimation unit 704 and the motion compensation unit 705 can perform different operations on the current video block depending, for example, whether the current video block is an I-slice, a P-slice, or a B-slice. An I-slice (or I-frame) is the least compressible but does not require other video frames for decoding. An S-slice (or P-frame) can use data from the previous frame for decompression and is more compressible than an I-frame. A B-slice (or B-frame) can use both the previous and next frames for data referencing to obtain the highest data compression ratio.
[0696] In some examples, the motion estimation unit 704 may perform a unidirectional prediction for the current video block, and the motion estimation unit 704 may look up a reference picture in list 0 or list 1 for a reference video block for the current video block. The motion estimation unit 704 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 704 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. The motion compensation unit 705 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0697] In other examples, the motion estimation unit 704 may perform bidirectional prediction for the current video block, and may search for reference pictures in List 0 for reference video blocks for the current video block, and may search for reference pictures in List 1 for other reference video blocks for the current video block. The motion estimation unit 704 may generate reference indices indicating the reference pictures in List 0 and List 1 that include the reference video blocks, and motion vectors indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 704 may output the reference indices and the motion vectors of the current video block as motion information for the current video block. The motion compensation unit 705 may generate predicted video blocks for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0698] In some cases, the motion estimation unit 704 can output a full set of motion information for the decoder's decoding process.
[0699] In some cases, the motion estimation unit 704 does not need to output the full set of motion information for the current video. Rather, the motion estimation unit 704 may refer to the motion information of another video block and signal the motion information of the current video block. For example, the motion estimation unit 704 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0700] As an example, the motion estimation unit 704 can indicate a value to the video decoder 624 in a syntactic structure related to the current video block that the current video block has the same motion information as another video block.
[0701] In another example, the motion estimation unit 704 may identify another video block and a motion vector difference (MVD) in a syntactic structure related to the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 624 can use the difference between the motion vector of the indicated video block and the motion vector to determine the motion vector of the current video block.
[0702] As described above, the video encoder 614 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 614 are advanced motion vector prediction (AMVP) and merge mode signaling.
[0703] The intra-prediction unit 706 may perform intra-prediction on the current video block. When the intra-prediction unit 706 performs intra-prediction on the current video block, it may generate prediction data for the current video block based on decoded samples from other video blocks within the same picture. The prediction data for the current video block may include the predicted video block and various syntactic elements.
[0704] The residual generation unit 707 may generate residual data for the current video block by subtracting (for example, indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the sample of the current video block.
[0705] In other cases, residual data for the current video block may not be available; for example, in skip mode, the residual generation unit 707 may not perform a subtraction operation.
[0706] The conversion unit 708 can generate one or more conversion coefficient video blocks for the current video block by applying one or more transformations to the residual video blocks associated with the current video block.
[0707] After the conversion unit 708 generates a conversion coefficient video block related to the current video block, the quantization unit 709 can quantize the conversion coefficient video block related to the current video block based on one or more quantization parameter (QP) values related to the current video block.
[0708] The inverse quantization unit 710 and the inverse transformation unit 711 may apply inverse quantization and inverse transformation to the transformation coefficient image block, respectively, to reconstruct a residual image block from the transformation coefficient image block. The reconstruction unit 712 can generate a reconstructed image block associated with the current block for storage in the buffer 713 by adding the reconstructed residual image block to the corresponding sample from one or more predicted image blocks generated by the prediction unit 702.
[0709] After the reconstruction unit 712 reconstructs the video block, loop filtering may be performed to reduce video block artifacts within the video block.
[0710] The entropy coding unit 714 may receive data from other functional components of the video encoder 700. When the entropy coding unit 714 receives data, it can perform one or more entropy coding operations to generate entropy coded data and output a bitstream containing the entropy coded data.
[0711] Figure 8 is a block diagram showing an example of a video decoder 800, which may be the video decoder 624 in the video coding system 600 shown in Figure 6.
[0712] The video decoder 800 may be configured to perform some or all of the technologies described herein. In the example shown in Figure 8, the video decoder 800 includes several functional components. The technologies described herein may be shared among the components of the video decoder 800. In some examples, a processor may be configured to perform some or all of the technologies described herein.
[0713] In the example shown in Figure 8, the video decoder 800 includes an entropy decoding unit 801, a motion compensation unit 802, an intra-prediction unit 803, an inverse quantization unit 804, an inverse transformation unit 805, a reconstruction unit 806, and a buffer 807. In some examples, the video decoder 800 can perform a decoding path that is generally the reverse of the encoding path described with respect to the video encoder 614 (Figure 6).
[0714] The entropy decoding unit 801 may acquire an encoded bitstream. The encoded bitstream may contain entropy-encoded video data (e.g., blocks of encoded video data). The entropy decoding unit 801 may decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 802 may determine motion information, including motion vectors, motion vector precision, a reference picture list index, and other motion information. The motion compensation unit 802 may determine such information, for example, by performing AMVP and merge mode signal notifications.
[0715] The motion compensation unit 802 generates motion-compensated blocks and may perform interpolation based on interpolation filters if possible. Identifiers of interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0716] The motion compensation unit 802 may calculate interpolated values for sub-integer pixels of a reference block using an interpolation filter used by the video encoder 614 during the encoding of the video block. The motion compensation unit 802 may determine the interpolation filter used by the video encoder 614 according to the received syntax information and generate a predicted block using the interpolation filter.
[0717] The motion compensation unit 802 may use some of the syntax information to determine the size of the blocks used to encode the frames (one or more) and / or slices (one or more) of the encoded video sequence, partition information describing how each macroblock of the picture in the encoded video sequence is divided, a mode indicating how each partition is encoded, one or more reference frames (and a list of reference frames) for each inter-encoded block, and other information for decoding the encoded video sequence.
[0718] The intra-prediction unit 803 can form prediction blocks from spatially adjacent blocks, for example, using the intra-prediction mode received in a bitstream. The inverse quantization unit 804 inverse quantizes, i.e., dequantizes, the quantized image block coefficients provided in the bitstream and decoded by the entropy decoding unit 801. The inverse transform unit 805 applies the inverse transform.
[0719] The reconstruction unit 806 can combine the residual blocks with the corresponding predicted blocks generated by the motion compensation unit 802 or the intra-prediction unit 803 to form decoded blocks. If necessary, a deblocking filter can also be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are stored in the buffer 807, which provides reference blocks for subsequent motion compensation / intra-prediction and generates decoded video for display on a display device.
[0720] Figure 9 shows a method 900 for coding video data according to an embodiment of the present disclosure. Method 900 may be performed by a coding device (e.g., an encoder) having a processor and memory. Method 900 may be performed when determining which primary layer is associated with an auxiliary layer when auxiliary information is present in the bitstream.
[0721] In block 902, the coding device uses Scalability Dimension Information (SDI) Auxiliary Extended Information (SEI) messages to indicate which primary layer is associated with an auxiliary layer, if auxiliary information is present in the bitstream. In the embodiment, a primary layer is associated with an auxiliary layer if it maps to an auxiliary layer, uses information from an auxiliary layer, or is associated with an auxiliary layer.
[0722] An SDI SEI message is a type of SEI message, such as the SEI message in bitstream 300 in Figure 3. An SEI message, including an SDI SEI message, can transmit any of the syntactic elements disclosed herein.
[0723] If sdi_aux_id[i]=0, the i-th layer is called the primary layer. Otherwise, the i-th layer is called the auxiliary layer. If sdi_aux_id[i]=1, the i-th layer is also called the alpha auxiliary layer. If sdi_aux_id[i]=2, the i-th layer is also called the depth auxiliary layer.
[0724] In block 904, the coding device performs a conversion between the video media file and the bitstream based on the SDI SEI message.
[0725] When performed in an encoder, the conversion includes receiving a media file (e.g., a video unit) and encoding the SEI messages into a bitstream. When performed in a decoder, the conversion includes receiving a bitstream containing the SEI messages and decoding the SEI messages in the bitstream to generate a video media file.
[0726] In this embodiment, one or more syntactic elements of the SDI SEI message indicate which primary layer is associated with the auxiliary layer, if auxiliary information is present in the bitstream.
[0727] In this embodiment, the auxiliary layer has a layer identifier (ID) designated as sdi_aux_id[i], where i is an integer corresponding to the auxiliary layer (e.g., 1, 2, 3, etc.).
[0728] In the embodiment, if auxiliary information is present in the bitstream, a layer index is included in the SDI SEI message to indicate which primary layer is associated with the auxiliary layer. In the embodiment, each layer index contains an entry or value that associates a primary layer with an auxiliary layer.
[0729] In the embodiment, one or more syntactic elements for the primary layer indicate whether the auxiliary layer applies to one or more primary layers.
[0730] In the embodiment, the syntactic element indicates whether the auxiliary layer is applied from a primary layer to a particular primary layer. In the embodiment, the syntactic element indicates whether the auxiliary layer is applied to one or more primary layers. In the embodiment, for example, the auxiliary layer is applied to a primary layer if the primary layer uses or benefits from the information contained in the auxiliary layer.
[0731] In this embodiment, the auxiliary layer is one of several auxiliary layers in the bitstream, and when auxiliary information is present in the bitstream, one or more syntactic elements are included in the SDI SEI message to indicate which primary layer is associated with each of the several auxiliary layers.
[0732] In this embodiment, an indication of the number of primary layers associated with the auxiliary pictures of the auxiliary layer is signaled in a bitstream.
[0733] In this embodiment, the instruction for the number of primary layers is specified as sdi_num_associated_primary_layers_minus1.
[0734] In this embodiment, sdi_num_associated_primary_layers_minus1 is signaled as a 6-bit unsigned integer. For example, the unsigned integer is an integer (e.g., the whole number) that is not associated with a sign (e.g., positive or negative).
[0735] In embodiments, an indication of the number of primary layers associated with an auxiliary layer, or associated with an auxiliary picture of an auxiliary layer, is conditionally notified in the bitstream. In embodiments, conditional notification means signaling specific information only when a condition is met.
[0736] In an embodiment, the bitstream includes a bitstream in scope, and the conditional notification includes signaling an indication of the number of primary layers only if the i-th layer of the bitstream in scope includes an auxiliary picture.
[0737] In this embodiment, if the layer identifier (ID) specified in sdi_aux_id[i] is greater than 0, the i-th layer in the bitstream within the scope contains an auxiliary picture.
[0738] In the embodiment, the bitstream includes a bitstream in scope, which is a sequence of access units (AUs) consisting of zero or more subsequent AUs, in decoding order, that do not include a first AU containing an SDI SEI message followed by a subsequent AU containing another SDI SEI message.
[0739] In the embodiment, if auxiliary information is present in the bitstream, or if the bitstream includes a bitstream within a scope and the bitstream within the scope is a multi-view bitstream, the SDI SEI message includes an auxiliary identifier (ID) for each layer. In the embodiment, a multi-layer bitstream is a bitstream that includes multiple layers, as shown in Figure 1, for example.
[0740] In one embodiment, the i-th layer is referred to as the primary layer if the layer identifier (ID) specified as sdi_aux_id[i] is equal to zero; otherwise, the i-th layer is referred to as the auxiliary layer.
[0741] In this embodiment, if the layer identifier (ID) specified as sdi_aux_id[i] is equal to 1, the i-th layer is referred to as the alpha auxiliary layer, and if the layer ID specified as sdi_aux_id[i] is equal to 2, the i-th layer is referred to as the depth auxiliary layer.
[0742] In embodiments, Method 900 may utilize or incorporate one or more features or steps of other methods disclosed herein.
[0743] Next, a list of preferred solutions in several embodiments is presented.
[0744] The following solutions illustrate an example of an embodiment of the technology discussed in this disclosure (e.g., Example 1).
[0745] 1. A video processing method comprising the steps of: performing a conversion between video and a video bitstream; ensuring the bitstream conforms to a format rule; and specifying that the format rule indicates a length obtained by subtracting L from the length of the syntactic element of a view identifier, where L is an integer.
[0746] 2. The method according to claim 1, wherein the syntactic element is coded as an unsigned integer using N bits.
[0747] 3. The method according to any one of claims 1 to 2, wherein L is a positive integer.
[0748] 4. The method according to claim 1, wherein L=0 and syntactic elements are prohibited from having zero values.
[0749] 5. A video processing method comprising the step of performing a conversion between a video having a plurality of layers and a bitstream of the video, wherein the bitstream conforms to a format rule, and the format rule specifies that the bitstream includes auxiliary layers associated with one or more associated layers of the video.
[0750] 6. The method according to claim 5, wherein the format rule further specifies whether or how the bitstream includes one or more syntactic elements indicating a relationship between the auxiliary layer and the one or more related layers, and the one or more syntactic elements are included in a scalability dimension auxiliary extended information syntactic structure.
[0751] 7. The method according to claim 6, wherein the format rule specifies that the one or more associated layers are indicated by corresponding layer identifiers (IDs).
[0752] 8. The method according to claim 6, wherein the format rule specifies that the one or more associated layers are indicated by their corresponding layer index.
[0753] 9. The method according to any one of claims 5 to 8, wherein the format rule specifies that the bitstream includes one or more syntactic elements indicating whether the auxiliary layer is applicable to one or more associated layers.
[0754] 10. The method according to claim 9, wherein the one or more syntactic elements include a syntactic element indicating that the auxiliary layer is applicable to all of the one or more related layers.
[0755] 11. The method according to claim 9, wherein the formatting rules specify that each related layer includes a syntactic element indicating whether the auxiliary layer is applicable to the corresponding related layer.
[0756] 12. The method according to claim 11, wherein the syntactic element indicates all primary layers related to the auxiliary layer.
[0757] 13. The method according to claim 11, wherein the syntactic element represents all primary layers relating to the auxiliary layer and having a layer index smaller than the layer index of the auxiliary layer.
[0758] 14. The method according to claim 11, wherein the syntactic element represents all primary layers relating to the auxiliary layer and having a layer index greater than the layer index of the auxiliary layer.
[0759] 15. The method according to any one of claims 11 to 14, wherein the syntactic element is a flag.
[0760] 16. The method according to claim 6, wherein the formatting rules specify that the bitstream does not include any explicit syntactic elements indicating the applicability of the auxiliary layer to the one or more associated layers, and that the applicability is derived during the transformation.
[0761] 17. The method according to claim 16, wherein the format rule specifies that the associated layers of the auxiliary layer have a layer ID that is the auxiliary layer's layer ID plus N1, N2...Nk, where k is an integer and for i=1,...k, the two Ni are not equal to each other.
[0762] 18. The method according to claim 17, wherein k=1 and N1 is one of 1, -1, 2 or -2.
[0763] 19. The method according to claim 17, wherein k is greater than 1.
[0764] The method according to claim 19, wherein k is equal to 2, N1=1, and N2=2.
[0765] 21. The method according to claim 5, wherein the formatting rules further specify that the bitstream omits one or more syntactic elements indicating a relationship between the auxiliary layer and one or more related layers, the relationship being derived based on predetermined rules.
[0766] 22. The method according to claim 5, wherein the formatting rules further specify that the bitstream includes one or more syntactic elements indicating a relationship between the auxiliary layer and one or more related layers, and the one or more syntactic elements are included in the auxiliary information auxiliary extended information syntactic structure.
[0767] 23. The method according to any one of claims 5 to 22, wherein the formatting rule specifies that the bitstream includes a syntactic element indicating the number of related layers of auxiliary pictures of the layer.
[0768] 24. The method according to any one of claims 5 to 22, wherein the format rule specifies that if the conditions are met, the bitstream will contain a syntactic element indicating the number of layers or auxiliary pictures related to the auxiliary pictures of the layers related to the auxiliary pictures.
[0769] 25. The method according to claim 24, wherein the condition includes that the i-th layer in the bitstream InScope includes an auxiliary picture.
[0770] 26. A video processing method comprising performing a conversion between a video including a plurality of video layers and a bitstream of the video, wherein the bitstream conforms to a format rule, and the format rule specifies that the bitstream includes multiview auxiliary extended information (SEI) messages or auxiliary information SEI messages, depending on whether the encoded video sequence of the bitstream includes scalability dimension information SEI messages.
[0771] 27. The method according to claim 26, wherein the format rule specifies that the multiview information SEI message refers to the multiview acquisition information SEI message.
[0772] 28. The method according to any one of claims 26 to 27, wherein the format rule specifies that the auxiliary information SEI message refers to a depth representation information SEI message or an alpha channel information SEI message.
[0773] 29. A video processing method comprising the steps of: performing a conversion between a video including a plurality of video layers and a bitstream of the video; and the bitstream conforming to a format rule, wherein the format rule specifies that, depending on the presence of multiview information or auxiliary information (SEI) messages in the bitstream, at least one of a first flag indicating the presence of multiview information or a second flag indicating the presence of auxiliary information in a scalability dimension information SEI message is equal to 1.
[0774] 30. A video processing method comprising the steps of: performing a conversion between a video including a plurality of video layers and a bitstream of the video; and the steps of: the bitstream conforming to a format rule, wherein the format rule specifies that multiview acquisition information auxiliary extension information messages included in the bitstream are not scalable nested or are included in scalable nesting auxiliary extension information messages.
[0775] 31. The method according to any one of claims 1 to 30, wherein the conversion includes generating the video from the bitstream or generating the bitstream from the video.
[0776] 32. A method for storing a bitstream in a computer-readable medium, comprising the steps of generating a bitstream according to one or more of the methods described in claims 1 to 31, and storing the bitstream in a computer-readable medium.
[0777] 33. A computer-readable medium having a bitstream of video, wherein the bitstream, when processed by a processor of a video decoder, causes the video decoder to generate the video, and the bitstream is generated according to one or more of the methods of claims 1 to 31.
[0778] 34. A video decoding device comprising a processor configured to perform one or more of the methods described in any one of claims 1 to 31.
[0779] 35. A video encoding device comprising a processor configured to perform one or more of the methods described in any one of claims 1 to 31.
[0780] 36. A computer program product having stored computer code, wherein the code, when executed by a processor, causes the processor to carry out the method according to any one of claims 1 to 31.
[0781] 37. A computer-readable medium on which a bitstream compliant with a bitstream format generated according to any one of claims 1 to 31 is recorded.
[0782] 38. A bitstream generated in accordance with the methods, apparatus, methods or systems disclosed herein.
[0783] The following documents may contain additional details related to the technology disclosed herein:
[0784] [1] ITU-T and ISO / IEC, "High efficiency video coding", Rec. ITU-T H.265 | ISO / IEC 23008-2 (current version).
[0785] [2] J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, J. Boyce, “Algorithm description of Joint Exploration Test Model 7 (JEM7),” JVET-G1001, Aug. 2017.
[0786] [3]Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Coding”, 2020.
[0787] [4]B. Bross, J. Chen, S. Liu, Y.-K. Wang (editors), “Versatile Video Coding (Draft 10),” JVET-S2001.
[0788] [5]Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, 「Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams」, 2020.
[0789] [6]J. Boyce, V. Drugeon, G. Sullivan, Y.-K. Wang (editors), 「Versatile supplemental enhancement information messages for coded video bitstreams (Draft 5),」 JVET-S2007.
[0790] Other solutions, examples, embodiments, modules, and functional operations disclosed herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or one or more combinations thereof, including the structures disclosed herein and their structural equivalents. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing device or for controlling the operation of a data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of a material that makes machine-readable propagating signals effective, or one or more combinations thereof. The term “data processing device” encompasses all devices and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the device may include code that constitutes the execution environment of the computer program, such as code that makes up processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, such as a mechanically generated electrical signal, optical signal, or electromagnetic wave signal, which is produced to encode information for transmission to a suitable receiving device.
[0791] Computer programs (also known as programs, software, software applications, scripts, or coding) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form as standalone programs, modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs do not necessarily correspond to files in a file system. A program can be part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program, or multiple coordinated files (e.g., a file containing one or more modules, subprograms, or parts of code). Computer programs can be deployed to run on a single computer, or on multiple computers located at one site, or distributed across multiple sites and interconnected by a communication network.
[0792] The processes and logic flows described herein can be executed by one or more programmable processors that run one or more computer programs that perform a function by acting on input data and producing outputs. The processes and logic flows can also be executed by special-purpose logic circuits, such as FPGAs (field programmable gate arrays) or ASICs (application-specific integrated circuits), and devices can also be implemented using such circuits.
[0793] Processors suitable for executing computer programs include, as an example, general-purpose and specialized microprocessors, and any one or more processors in any type of digital computer. Generally, a processor receives instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer is operablely coupled to receive data from, transfer data to, or transfer data to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, optical disks, etc. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and compact disks. The processor and memory can be supplemented by or integrated into specialized logic circuits.
[0794] While this patent document contains many specific details, these should not be interpreted as limitations on the scope of any subject matter or claimable matters, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Specific features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable partial combination in multiple embodiments. Furthermore, even if features are described above as acting in a particular combination and are initially claimed as such, one or more features may be removed from the claimed combination, and the claimed combination may be directed towards a subcombination or variation of a subcombination.
[0795] Similarly, although the operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in a specific order or sequence shown, or that all illustrated operations be performed, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0796] Only a few implementations and examples are described, and other implementations, extensions, and modifications are possible based on those described and illustrated in this patent document.
Claims
1. A method for processing video data, Regarding the conversion between the video and the bitstream of the said video, it is determined that the current auxiliary layer is applied to at least one associated primary layer. Based on the above decision, the conversion is performed, The at least one associated primary layer has a first syntactic element equal to a first value, and the current auxiliary layer has a first syntactic element greater than the first value. The first syntactic element is included in the Scalability Dimension Information (SDI) Supplementary Extension Information (SEI) message, At least one syntactic element indicating at least one associated primary layer of each auxiliary layer is included in the SDI SEI message, method.
2. A second syntactic element of the at least one syntactic element representing the at least one associated primary layer of each auxiliary layer is explicitly signaled as one or a group of syntactic elements in the SDI SEI message. The method according to claim 1.
3. The second syntactic element of the SDI SEI message indicates the layer index of the at least one associated primary layer of each auxiliary layer. The method according to claim 2.
4. The third syntactic element of the at least one syntactic element indicates the number of the at least one associated primary layers of each auxiliary layer. The method according to claim 1.
5. The third syntactic element is conditionally included in the bitstream, In response that the value of the first syntactic element is greater than 0, the third syntactic element is included in the bitstream. The method according to claim 4.
6. The second syntactic element described above is coded as an unsigned integer using N bits, where N is an integer. The method according to claim 2.
7. N = 6. The method according to claim 6.
8. The third syntactic element described above is coded as an unsigned integer using M bits, where M is an integer. The method according to claim 4.
9. M = 6. The method according to claim 8.
10. A first syntactic element equal to zero indicates that the current layer in the bitstream does not contain an auxiliary picture. The first syntactic element greater than zero indicates the type of auxiliary picture in the current layer within the bitstream. The method according to claim 1.
11. If the value of the first syntactic element is equal to zero, the current layer is referred to as the primary layer; otherwise, the current layer is referred to as the auxiliary layer. If the value of the first syntactic element is equal to 1, the current layer is referred to as the alpha auxiliary layer. If the value of the first syntactic element is equal to 2, the current layer is referred to as the depth auxiliary layer. The method according to claim 1.
12. The conversion includes encoding the video into a bitstream. The method according to any one of claims 1 to 11.
13. The conversion includes decoding the video from the bitstream. The method according to any one of claims 1 to 11.
14. A device for processing video data, comprising a processor and a non-transient memory in which instructions are stored, When the aforementioned instruction is executed by the processor, the processor will: In the conversion between video and the bitstream of said video, the current auxiliary layer is determined to be applied to at least one associated primary layer. Based on the above decision, the conversion is performed, The at least one associated primary layer has a first syntactic element equal to a first value, and the current auxiliary layer has a first syntactic element greater than the first value. The first syntactic element is included in the Scalability Dimension Information (SDI) Supplementary Extension Information (SEI) message, At least one syntactic element indicating at least one associated primary layer of each auxiliary layer is included in the SDI SEI message, Device.
15. A non-transient computer-readable storage medium that stores instructions to be executed by a processor, In the conversion between video and the bitstream of said video, the current auxiliary layer is determined to be applied to at least one associated primary layer. Based on the above decision, the conversion is performed, The at least one associated primary layer has a first syntactic element equal to a first value, and the current auxiliary layer has a first syntactic element greater than the first value. The first syntactic element is included in the Scalability Dimension Information (SDI) Supplementary Extension Information (SEI) message, At least one syntactic element indicating at least one associated primary layer of each auxiliary layer is included in the SDI SEI message, A non-transient computer-readable storage medium.
16. A method for storing a video bitstream, The aforementioned method, Determine that the current auxiliary layer is applied to at least one associated primary layer, Based on the above decision, the bitstream of the video is generated, The bitstream is stored on a non-transient computer-readable recording medium. The at least one associated primary layer has a first syntactic element equal to a first value, and the current auxiliary layer has a first syntactic element greater than the first value. The first syntactic element is included in the Scalability Dimension Information (SDI) Supplementary Extension Information (SEI) message, At least one syntactic element indicating at least one associated primary layer of each auxiliary layer is included in the SDI SEI message, method.