Methods, systems, and computer programs for supporting mixed NAL unit type in coded picture

JP2025041860A5Active Publication Date: 2025-05-12TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024228442
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-10-22
Filing Date
2024-12-25
Publication Date
2025-05-12
Estimated Expiration
2040-12-16

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide methods and systems for decoding video streams.SOLUTION: A method includes receiving a first network abstraction layer (NAL) unit of a first slice of a coded picture and a second VCL NAL unit of a second slice of the coded picture, the first VCL NAL unit having a first VCL NAL unit type and the second VCL NAL unit having a second VCL NAL unit type that is different from the first VCL NAL unit type. The method also includes decoding the coded picture, the decoding including determining a picture type of the coded picture based on the first VCL NAL unit type of the first VCL NAL unit and the second VCL NAL unit type of the second VCL NAL unit, or based on an indicator, received by the at least one processor, indicating that the coded picture includes mixed VCL NAL unit types.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 956,254, filed January 1, 2020, and U.S. Patent Application No. 17 / 077,035, filed October 22, 2020, both of which are incorporated herein in their entireties.

[0002] TECHNICAL FIELD Embodiments of the present disclosure relate to video encoding and decoding, and more particularly, to supporting mixed Network Abstraction (NAL) unit types for coded pictures. [Background technology]

[0003] The Generic Video Coding (VVC) draft specification JVET-P2001 (incorporated in its entirety) (editorially updated by JVET-Q0041) supports a mixed network abstraction layer (NAL) unit type feature, which allows having one or more slice NAL units with NAL unit type equal to Intra Random Access Point (IRAP) or Clean Random Access (CRA) and one or more slice NAL units with NAL unit type equal to non-IRAP. This feature can be used to merge two different bitstreams into one or to support different random access periods for each local region (subpicture). Currently, the following syntax and semantics are defined to support the feature:

[0004] An example picture parameter set raw byte sequence payload (RBSP) syntax is provided in Table 1 below. Table 1 [Table 1]

[0005] The syntax element mixed_nalu_types_in_pic_flag equal to 1 specifies that each picture that references a picture parameter set (PPS) has more than one video coding layer (VCL) NAL unit, that no VCL NAL units have the same value of nal_unit_type, and that the picture is not an IRAP picture. The syntax element mixed_nalu_types_in_pic_flag equal to 0 specifies that each picture that references a PPS has one or more VCL NAL units, and that the VCL NAL units of each picture that references a PPS have the same value of nal_unit_type.

[0006] If the syntax element no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of the syntax element mixed_nalu_types_in_pic_flag shall be equal to 0.

[0007] According to the current VVC specification, NAL unit type codes and NAL unit type classes are defined as shown in Table 2 below. Table 2 [Table 2] JPEG2025041860000004.jpg110161

[0008] For each slice with value nalUnitTypeA of nal_unit_type in the range IDR_W_RADL to CRA_NUT inclusive in picture picA (picA also contains one or more slices with other values ​​of nal_unit_type), the following applies:

[0009] (A) The slice belongs to subpicA, for which the value of the corresponding syntax element subpic_treated_as_pic_flag[ i ] is equal to 1.

[0010] (B) The slice shall not belong to a subpicture of picA that contains a VCL NAL unit with syntax element nal_unit_type not equal to nalUnitTypeA.

[0011] (C) For all subsequent PUs in the coding layer video sequence (CLVS) in decoding order, neither RefPicList[0] nor RefPicList[1] of a slice of subpicA shall contain any picture in picA that precedes it in decoding order as an active entry.

[0012] For any particular picture's VCL NAL units, the following applies:

[0013] If the syntax element mixed_nalu_types_in_pic_flag is equal to 0, the value of the syntax element nal_unit_type shall be the same for all coded slice NAL units of a picture. A picture or PU is referred to as having the same NAL unit type as the coded slice NAL units of the picture or PU.

[0014] Otherwise (syntax element mixed_nalu_types_in_pic_flag is equal to 1), one or more VCL NAL units shall all have a specific value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive, and all other VCL NAL units shall have a specific value of nal_unit_type in the range of TRAIL_NUT to RSV_VCL_6, inclusive, or equal to GDR_NUT. Summary of the Invention

[0015] The current design of the mixed VCL NAL unit types described in the Background section above has several potential problems.

[0016] In some cases, the picture type of a picture can be ambiguous if the picture is composed of mixed VCL NAL unit types.

[0017] In some cases, temporal identifier (e.g., TemporalId) constraints may conflict when NAL unit types are mixed in the same PU (picture).

[0018] For example, the current VVC specification has the following constraints on TemporalIds: If the syntax element nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_12, inclusive, the syntax element TemporalId shall be equal to 0. If the syntax element nal_unit_type is equal to STSA_NUT, the syntax element TemporalId shall not be equal to 0.

[0019] Optionally, if the syntax element mixed_nalu_types_in_pic_flag is signaled in a PPS, at least two PPS NAL units shall be referenced by a slice NAL unit in CLVS, and if a subpicture is extracted, the associated PPS shall be rewritten by changing the value of the syntax element mixed_nalu_types_in_pic_flag.

[0020] In some cases, the current design may not support the coexistence of random-access decodable leading (RADL) / random-access skipped leading (RASL) NAL units with trail pictures (PUs) within a picture.

[0021] In some cases, the syntax element mixed_nalu_types_in_pic_flag may be inconsistent when a picture in a layer references another picture in a different layer.

[0022] SUMMARY OF THE DISCLOSURE Embodiments of the present disclosure may address one or more of the problems set forth above and / or other problems.

[0023] According to one or more embodiments, a method is provided that is performed by at least one processor, the method including: receiving a first video coding layer (VCL) network abstraction layer (NAL) unit of a first slice of a coded picture and a second VCL NAL unit of a second slice of the coded picture, the first VCL NAL unit having a first VCL NAL unit type and the second VCL NAL unit having a second VCL NAL unit type that is different from the first VCL NAL unit type; and decoding the coded picture, the decoding including determining a picture type of the coded picture based on the first VCL NAL unit type of the first VCL NAL unit and the second VCL NAL unit type of the second VCL NAL unit or based on an indicator received by the at least one processor, the indicator indicating that the coded picture includes a different VCL NAL unit type.

[0024] According to an embodiment, the determining step includes determining that the coded picture is a trailing picture based on a first VCL NAL unit type indicating that the first VCL NAL unit includes a trailing picture coded slice and a second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0025] According to an embodiment, the determining step includes determining that the coded picture is a random-access decodable reading (RADL) picture based on a first VCL NAL unit type indicating that the first VCL NAL unit includes a RADL picture coded slice and a second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0026] According to an embodiment, the determining step includes determining that the coded picture is a step-wise temporal sub-layer access (STSA) picture based on a first VCL NAL unit type indicating that the first VCL NAL unit includes an STSA picture coded slice and a second VCL NAL unit type indicating that the second VCL NAL unit does not include an instantaneous decoding refresh (IDR) picture coded slice.

[0027] According to an embodiment, the determining step includes determining that the coded picture is a trailing picture based on a first VCL NAL unit type indicating that the first VCL NAL unit includes a step-wise temporal sublayer access (STSA) picture coded slice and a second VCL NAL unit type indicating that the second VCL NAL unit does not include a clean random access (CRA) picture coded slice.

[0028] According to an embodiment, the determining step includes determining that the coded picture is a trailing picture based on the first VCL NAL unit type indicating that the first VCL NAL unit includes a gradual decoding refresh (GDR) picture coded slice and the second VCL NAL unit type indicating that the second VCL NAL unit does not include an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0029] According to an embodiment, the indicator is a flag and the determining step includes determining that the coded picture is a trailing picture based on the flag indicating that the coded picture includes a mixed VCL NAL unit type.

[0030] According to an embodiment, the indicator is a flag, and the step of decoding the coded picture further includes a step of determining that a temporal ID of the coded picture is 0 based on the flag indicating that the coded picture contains a mixed VCL NAL unit type.

[0031] According to an embodiment, the indicator is a flag and the method further comprises receiving the flag in a picture header or a slice header.

[0032] According to an embodiment, the indicator is a flag and the coded picture is in a first layer, and the method further includes receiving the flag; and determining that an additional coded picture in a second layer, which is a reference layer for the first layer, includes a mixed VCL NAL unit type based on the flag indicating that the coded picture includes a mixed VCL NAL unit type.

[0033] According to one or more embodiments, a system is provided that includes a memory configured to store a computer program; and at least one processor configured to receive at least one coded video stream, access the computer program code, and operate as instructed by the computer code. The computer program code includes a decoding code configured to cause the at least one processor to decode a coded picture from the at least one coded video stream, the decoding code including a decision code configured to cause the at least one processor to determine a picture type of the coded picture based on a first Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit type of a first slice of the coded picture and a second VCL NAL unit type of a second slice of the coded picture, or based on an indicator received by the at least one processor indicating that the coded picture includes mixed VCL NAL unit types, the first VCL NAL unit type being different from the second VCL NAL unit type.

[0034] According to an embodiment, the decision code is configured to cause the at least one processor to determine that the coded picture is a trailing picture based on the first VCL NAL unit type indicating that the first VCL NAL unit includes a trailing picture coded slice and the second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0035] According to an embodiment, the decision code is configured to cause at least one processor to determine that the coded picture is a random-access decodable reading (RADL) picture based on a first VCL NAL unit type indicating that the first VCL NAL unit includes a RADL picture coded slice and a second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0036] According to an embodiment, the decision code is configured to cause the at least one processor to determine that the coded picture is a Step-Wise Temporal Sublayer Access (STSA) picture based on the first VCL NAL unit type indicating that the first VCL NAL unit includes an STSA picture coded slice and the second VCL NAL unit type indicating that the second VCL NAL unit does not include an Instant Decoding Refresh (IDR) picture coded slice.

[0037] According to an embodiment, the decision code is configured to cause the at least one processor to determine that the coded picture is a trailing picture based on the first VCL NAL unit type indicating that the first VCL NAL unit includes a step-wise temporal sublayer access (STSA) picture coded slice and the second VCL NAL unit type indicating that the second VCL NAL unit does not include a clean random access (CRA) picture coded slice.

[0038] According to an embodiment, the decision code is configured to cause the at least one processor to determine that the coded picture is a trailing picture based on the first VCL NAL unit type indicating that the first VCL NAL unit includes a gradual decoding refresh (GDR) picture coded slice and the second VCL NAL unit type indicating that the second VCL NAL unit does not include an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice.

[0039] According to an embodiment, the indicator is a flag and the decision code is configured to cause the at least one processor to determine that the coded picture is a trailing picture based on the flag indicating that the coded picture includes mixed VCL NAL unit types.

[0040] According to an embodiment, the indicator is a flag, and the decision code is further configured to cause the at least one processor to determine that the temporal ID of the coded picture is 0 based on the flag indicating that the coded picture includes mixed VCL NAL unit types.

[0041] According to an embodiment, the indicator is a flag, and the at least one processor is configured to receive the flag in a picture header or a slice header.

[0042] According to one or more embodiments, a non-transitory computer-readable medium is provided that stores computer instructions that, when executed by at least one processor, cause the at least one processor to decode a coded picture from at least one coded video stream, the decoding including determining a picture type of the coded picture based on a first video coding layer (VCL) network abstraction layer (NAL) unit type of a first slice of the coded picture and a second VCL NAL unit type of a second slice of the coded picture, or based on an indicator received by the at least one processor that indicates the coded picture includes mixed VCL NAL unit types, where the first VCL NAL unit type is different from the second VCL NAL unit type. [Brief description of the drawings]

[0043] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0044] [Figure 1] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.

[0045] [Diagram 2] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.

[0046] [Diagram 3] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment;

[0047] [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment;

[0048] [Diagram 5] FIG. 2 is a block diagram of a NAL unit according to an embodiment.

[0049] [Figure 6] FIG. 2 is a block diagram of a decoder according to an embodiment.

[0050] [Figure 7] FIG. 1 is a diagram of a computer system suitable for implementing embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0051] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). For one-way transmission of data, a first terminal (110) may code video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the coded video data of the other terminal from the network (150), decode the coded data, and display the recovered video data. One-way data transmission is common in media serving applications and the like.

[0052] 1 illustrates a second pair of terminals (130, 140) provided to support bidirectional transmission of coded video, such as may occur during a video conference. For the bidirectional transmission of data, each terminal (130, 140) can code video data captured at a local location for transmission over the network (150) to the other terminal. Each terminal (130, 140) can also receive coded video data transmitted by the other terminal, can decode the coded data, and can display the recovered video data on a local display device.

[0053] In FIG. 1, the terminals (110-140) may be depicted as servers, personal computers, smartphones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that transmit coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (150) may not be important to the operation of the present disclosure, unless described below.

[0054] Figure 2 illustrates the arrangement of video encoders and decoders in a streaming environment as an example application of the disclosed subject matter, which may be similarly applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.

[0055] As shown in FIG. 2, the streaming system (200) can include a video source (201) and a capture system (213) that can include an encoder (203). The video source (201) can be, for example, a digital camera and can be configured to generate an uncompressed video sample stream (202). The uncompressed video sample stream (202) provides a high data volume when compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof and can enable or achieve aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (204) can include a lower data volume when compared to the sample stream and can be stored at the streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to retrieve a video bitstream (209), which may be a copy (207) of the encoded video bitstream (204).

[0056] In an embodiment, the streaming server (205) may function as a media aware network element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to tailor potentially different bitstreams to one or more streaming clients (206). In an embodiment, a MANE may be provided separately from the streaming server (205) in the streaming system (200).

[0057] The streaming client (206) may include a video decoder (210) and a display (212). The video decoder (210) may, for example, decode a video bitstream (209), which may be an incoming copy of the encoded video bitstream (204), and generate a running video sample stream (211) that may be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) may be encoded according to a particular video coding / compression standard. Examples of these standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard informally known as Versatile Video Coding (VVC) is under development. The disclosed embodiments may be used in the context of VVC.

[0058] FIG. 3 illustrates an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure.

[0059] The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensated prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (357). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The video decoder (210) may also be implemented, in part or in whole, in software running on one or more CPUs with associated memory.

[0060] In this and other embodiments, the receiver (310) can receive one or more coded video sequences to be decoded by the decoder (210), one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (312), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (310) can receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which can each be forwarded using an entity (not shown). The receiver (310) can separate the coded video sequences from the other data. To address network jitter, a buffer memory (315) can be coupled between the receiver (310) and the entropy decoder / parser (320) (hereafter referred to as the "parser"). If the receiver (310) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer (315) may not be needed or may be small. For use in best-effort packet networks such as the Internet, the buffer (315) may be required and may be relatively large and adaptively sized.

[0061] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from the entropy coded video sequence. These symbol categories include, for example, information used to manage the operation of the decoder (210) and information that may control a rendering device, such as a display (212), which may be coupled to the decoder, as shown in FIG. 2. The rendering device control information may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (320) may parse / entropy decode the coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context effects, etc. The parser (320) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on the at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (320) can also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0062] The parser (320) can perform entropy decoding / parsing operations on the video sequence received from the buffer (315) to create symbols (321).

[0063] The reconstruction of the symbols (321) may involve a number of different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture, intra-picture, inter-block, intra-block) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (320) from the coded video sequence. The flow of such subgroup control information between the parser (320) and subsequent units is not depicted for clarity.

[0064] Beyond the functional blocks already mentioned, the decoder 210 can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is adequate.

[0065] One unit may be a scalar / inverse transform unit (351), which receives quantized transform coefficients as well as control information, including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc., as symbols (321) from the parser (320). The scalar / inverse transform unit (351) may output blocks containing sample values ​​that may be input to an aggregator (355).

[0066] In some cases, the output samples of the scalar / inverse transform (351) may be associated with intra-coded blocks, i.e. blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by an intra picture prediction unit (352). In some cases, the intra picture prediction unit (352) generates blocks of the same size and shape of the block being reconstructed using surrounding already reconstructed information retrieved from the current (partially reconstructed) picture from the current picture memory (358). The aggregator (355) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra prediction unit (352) to the output sample information as provided by the scalar / inverse transform unit (351).

[0067] In other cases, the output samples of the scalar / inverse transform unit (351) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensation prediction unit (353) may access the reference picture memory (357) to retrieve samples used for prediction. After motion compensating the retrieved samples according to the symbols (321) associated with the block, these samples may be added by the aggregator (355) to the output of the scalar / inverse transform unit (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (357) from which the motion compensation unit retrieves the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion compensation prediction unit (353) in the form of a symbol (321), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​taken from a reference picture memory (357), motion vector prediction mechanisms, etc., when sub-sample accurate motion vectors are used.

[0068] The output samples of the aggregator (355) may be subjected to various loop filtering techniques in a loop filter unit (356). The video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but may also be responsive to meta-information obtained during the decoding of a previous part (in decoding order) of the coded picture or coded video sequence, and may also be responsive to previously reconstructed loop filtered sample values.

[0069] The output of the loop filter unit (356) may be a sample stream that can be output to a rendering device, such as a display (212), or that can be stored in a reference picture memory for use in future inter-picture prediction.

[0070] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and memory for the new current picture can be reallocated before beginning reconstruction of a future coded picture.

[0071] The video decoder (210) may perform decoding operations according to a given video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to a syntax defined by the video compression technique or standard used, in the sense that the coded video sequence conforms to the syntax of the video compression technique or standard as defined in the video compression technique document or standard, and in particular in a profile document therein. Also, for compliance with some video compression techniques or standards, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. The limits set by the level may be further limited, in some cases, by metadata for HRD buffer management and Hypothetical Reference Decoder (HRD) specifications signaled in the coded video sequence.

[0072] In an embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0073] FIG. 4 illustrates an example functional block diagram of a video encoder (203) associated with a video source (201) in accordance with an embodiment of the present disclosure.

[0074] The video encoder (203) may include, for example, an encoder including a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).

[0075] The encoder (203) can receive video samples from a video source (201) (not part of the encoder) capable of capturing video images to be coded by the encoder (203).

[0076] The video source (201) may provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...) and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following discussion focuses on examples.

[0077] According to an embodiment, the encoder (203) can code and compress pictures of a source video sequence into a coded video sequence (443) in real time or under any other time constraint required by the application. Imposing an appropriate coding rate is one function of the controller (450). The controller can also control and be functionally coupled to other functional units as described below, the couplings of which are not depicted for clarity. Parameters set by the controller can include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. One skilled in the art can easily identify other functions of the controller (450) as they may be relevant to a video encoder (203) optimized for a particular system design.

[0078] Some video encoders operate in what one skilled in the art would easily recognize as a "coding loop". As an oversimplified explanation, the coding loop can consist of an encoding part with a source coder (430) (responsible for generating symbols based on the input picture to be coded and reference pictures) and a (local) decoder (433) built into the encoder (203) that reconstructs the symbols to create sample data that a (remote) decoder will also create if the video compression technique provides lossless compression between the symbols and the coded video bitstream. The reconstructed sample stream can be input to a reference picture memory (434). Since the decoding of the symbol stream produces bit-exact results independent of the location of the decoder (local or remote), the contents of the reference picture memory are also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values ​​as the reference picture samples that the decoder would "see" if it were to use prediction during decoding. This basic principle of reference picture synchrony (if synchrony cannot be maintained, e.g., due to channel errors, drift will result) is well known to those skilled in the art.

[0079] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder (210), which has already been described in detail in relation to Figure 3. However, since symbols are available and the encoding / decoding of symbols by the entropy coder (445) and parser (320) for the coded video sequence may be lossless, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer (315) and parser (320), may not be fully implemented in the local decoder (433).

[0080] An insight that can be made at this point is that any decoder technique present in the decoder, with the exception of analysis / entropy decoding, must also be present in the corresponding encoder, in substantially the same functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder technique can be omitted, as it may be the inverse of the decoder technique described generically. Only in certain areas is a more detailed description required, which is provided below.

[0081] As part of its operation, the source coder (430) can perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (432) codes differences between pixel blocks of the input frame and pixel blocks of the reference frames that can be selected as predictive references for the input frame.

[0082] The local video decoder (433) can decode the coded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (430). The operation of the coding engine (432) can advantageously be a lossless process. If the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can be a replica of the source video sequence, typically with some errors. The local video decoder (433) can replicate the decoding process that can be performed by the video decoder on the reference frame and cause the reconstructed reference frame to be stored in the reference picture memory (434). In this way, the encoder (203) can store a copy of the locally reconstructed reference frame that has a common content with the reconstructed reference frame obtained by the far-end video decoder (in the absence of transmission errors).

[0083] The predictor (435) can perform a prediction search for the coding engine (432). That is, for a new frame to be coded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata (e.g., reference picture motion vectors, block shapes, etc.) that may serve as suitable prediction references for the new picture. The predictor (435) can operate block-by-block on the samples to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (435), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (434).

[0084] The controller (450) can manage the coding operations of the video coder (430), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0085] The output of all the aforementioned functional units may be subjected to entropy coding in an entropy coder (445), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.

[0086] The transmitter (440) can buffer the coded video sequence as produced by the entropy coder (445) and prepare it for transmission over a communication channel (460), which may be a hardware / software link to a storage device that will store the coded video data. The transmitter (440) can merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0087] The controller (450) can manage the operation of the encoder (203). During coding, the controller (450) can assign a particular coded picture type to each coded picture, which can affect the coding technique that can be applied to the individual picture. For example, pictures can often be designated as intra pictures (I pictures), predicted pictures (P pictures), or bidirectionally predicted pictures (B pictures).

[0088] An intra picture (I-picture) may be one that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoding Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.

[0089] A predictive picture (P-picture) may be one that can be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0090] Bidirectionally predicted pictures (B-pictures) may be those that can be encoded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a block.

[0091] A source picture is usually spatially divided into several sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 blocks of samples each) and can be coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the respective picture of the block. For example, blocks of I pictures can be non-predictively coded or they can be predictively coded (spatial or intra-predicted) with reference to already coded blocks of the same picture. Pixel blocks of P pictures can be non-predictively coded with spatial prediction with reference to one previously coded reference picture or with temporal prediction. Blocks of B pictures can be non-predictively coded with spatial prediction with reference to one or two previously coded reference pictures or with temporal prediction.

[0092] The video coder (203) is capable of performing coding operations according to a given video coding technique or standard, such as ITU-T Rec. H.265. During operation, the video coder (203) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0093] In an embodiment, the transmitter (440) can transmit additional data along with the encoded video. The video coder (430) can include such data as part of the coded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0094] Embodiments of the present disclosure may modify the current VVC specification and may implement the NAL unit type codes and NAL unit type classes defined in Table 2 above.

[0095] An "Intra Random Access Point picture" (or "IRAP picture") may be a picture that does not reference any other picture for inter prediction in the decoding process and may be a Clean Random Access picture (CRA) or an Immediate Decoding Refresh (IDR) picture. The first picture of the bitstream in decoding order may be an IRAP or a Gradual Decoding Refresh (GDR) picture. An IRAP picture and all subsequent non-IRAP pictures in decoding order in a Coded Video Sequence (CVS) can be correctly decoded without performing the decoding process of any picture preceding the IRAP picture in decoding order, provided that the necessary parameter sets are available if they need to be referenced.

[0096] A "trailing picture" may be a non-IRAP picture that follows the associated IRAP picture in output order and is not a Step-Wise Temporal Sub-Layer Access (STSA) picture.

[0097] A "Step-Wise Temporal Sublayer Access picture" (or "STSA picture") may be a picture that does not use a picture with the same TemporalId as the STSA picture for inter-prediction reference. A picture following the STSA picture in decoding order with the same TemporalId as the STSA picture may not use a picture preceding the STSA picture in decoding order with the same TemporalId as the STSA picture for inter-prediction reference. A STSA picture may allow an up-switch at the STSA picture from the sublayer immediately below it to the sublayer that contains the STSA picture. A STSA picture may have a TemporalId greater than 0.

[0098] A "Random Access Skip Leading Picture" (or "RASL picture") may be a leading picture of its associated CRA picture. If the associated CRA picture has NoIncorrectPicOutputFlag equal to 1, the RASL picture may not be output and may not be decoded correctly because it may contain references to pictures that do not exist in the bitstream. RASL pictures may not be used as reference pictures in the decoding process of non-RASL pictures. If present, every RASL picture may precede, in decoding order, all trailing pictures of the same associated CRA picture.

[0099] A "random access decodable leading picture" (or "RADL picture") may be a leading picture that is not used as a reference picture for the decoding process of the trailing pictures of the same associated IRAP picture. If present, all RADL pictures may precede, in decoding order, all trailing pictures of the same associated IRAP picture.

[0100] An "instant decoding refresh picture" (or "IDR picture") may be a picture that has no associated leading picture present in the bitstream (e.g., nal_unit_type equals IDR_N_LP), or a picture that has no associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream (e.g., nal_unit_type equals IDR_W_RADL).

[0101] A "Clean Random Access Picture" (or "CRA picture") may be a picture that does not reference any other pictures for inter prediction in its decoding process and may be the first picture in the bitstream in decoding order or may appear later in the bitstream. A CRA picture may have associated RADL or RASL pictures. If a CRA picture has NoIncorrectPicOutputFlag equal to 1, the associated RASL pictures may not be output by the decoder since they may not be decodable because they may contain references to pictures that are not present in the bitstream.

[0102] According to one or more embodiments, if the syntax element mixed_nalu_types_in_pic_flag of the PPS referenced by a coded picture is equal to 1, the picture type of the coded picture is determined (e.g., by a decoder) as follows:

[0103] (A) If the nal_unit_type of a NAL unit of a picture is equal to TRAIL_NUT and the nal_unit_type of another NAL unit of the picture is in the range from IDR_W_RADL to CRA_NUT, the picture is determined as a trailing picture.

[0104] (B) If the nal_unit_type of a NAL unit of a picture is equal to RADL_NUT and the nal_unit_type of another NAL unit of the picture is within the range of IDR_W_RADL to CRA_NUT, the picture is determined as a RADL picture.

[0105] (C) If the nal_unit_type of a NAL unit of a picture is equal to STSA_NUT and the nal_unit_type of another NAL unit of the picture is IDR_W_RADL or IDR_N_LP, the picture is determined as an STSA picture.

[0106] (D) If the nal_unit_type of a NAL unit of a picture is equal to STSA_NUT and the nal_unit_type of another NAL unit of the picture is CRA_NUT, the picture is determined as a trailing picture.

[0107] (E) If the nal_unit_type of a NAL unit of a picture is equal to GDR_NUT and the nal_unit_type of another NAL unit of the picture is in the range from IDR_W_RADL to CRA_NUT, the picture is determined as a trailing picture.

[0108] According to one or more embodiments, if the syntax element mixed_nalu_types_in_pic_flag of the PPS referenced by the coded picture is equal to 1, the picture type of the coded picture is determined (eg, by a decoder) as a trailing picture.

[0109] The above aspects can provide a solution to "Problem 1" described in the Summary section above.

[0110] According to one or more embodiments, mixing of STSA NAL units with IRAP NAL units may not be allowed.

[0111] For example, for any picture VCL NAL unit, it is possible to implement the following:

[0112] If mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same (e.g., can be determined to be the same) for all coded slice NAL units of a picture. A picture or PU is referred to as having the same NAL unit type as the coded slice NAL units of the picture or PU.

[0113] Otherwise (mixed_nalu_types_in_pic_flag is equal to 1), one or more VCL NAL units shall all have (and may be determined to all have) a particular value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive, and all other VCL NAL units shall all have (and may be determined to all have) a particular value of nal_unit_type in the range of RADL_NUT to RSV_VCL_6, inclusive, or equal to GDR_NUT or TRAIL_NUT.

[0114] According to an embodiment, an encoder may be configured to apply the above to prohibit mixing of STSA NAL units with IRAP NAL units, and according to an embodiment, a decoder may be configured to determine a value of the NAL unit type based on the above.

[0115] In accordance with one or more embodiments, the TemporalId constraint on STSA_NUT in the current VVC specification draft JVET-P2001 may be removed.

[0116] That is, embodiments of the present disclosure may not implement a constraint that, for example, if nal_unit_type is equal to STSA_NUT, then TemporalId shall not be equal to 0. However, embodiments may still implement a constraint that if nal_unit_type is within the range of IDR_W_RADL to RSV_IRAP_12, inclusive, then TemporalId shall be equal to 0 (e.g., may be determined to be equal).

[0117] According to one or more embodiments, a constraint may be implemented that the TemporalId of a picture having mixed_nalu_types_in_pic_flag equal to 1 shall be equal to 0. For example, an encoder or decoder of this disclosure may determine that the temporal ID of a picture is 0 based on the flag mixed_nalu_types_in_pic_flag being equal to 1.

[0118] The above aspects can provide a solution to "Problem 2" described in the Summary section above.

[0119] According to one or more embodiments, the syntax element mixed_nalu_types_in_pic_flag may be provided in a picture header or a slice header instead of in the PPS. An example of the syntax element mixed_nalu_types_in_pic_flag in a picture header is provided in Table 3 below. Table 3 [Table 3]

[0120] The syntax element mixed_nalu_types_in_pic_flag equal to 1 may specify that each picture associated with the PH has more than one VCL NAL unit, that the VCL NAL units do not have the same value of nal_unit_type, and that the picture is not an IRAP picture. The syntax element mixed_nalu_types_in_pic_flag equal to 0 may specify that each picture associated with the PH has one or more VCL NAL units, and that the VCL NAL units of each picture associated with the PH have the same value of nal_unit_type.

[0121] If the syntax element no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, then the value of mixed_nalu_types_in_pic_flag shall be (eg, can be determined to be) equal to 0.

[0122] According to one or more embodiments, the syntax element mixed_nalu_types_in_pic_flag may be provided in the picture header or slice header along with the current flags in the SPS.

[0123] An example of an SPS with the current flag (sps_mixed_nalu_types_present_flag) is provided in Table 4 below. Table 4 [Table 4]

[0124] An example of a picture header with the syntax element mixed_nalu_types_in_pic_flag is provided in Table 5 below. Table 5 [Table 5]

[0125] The syntax element sps_mixed_nalu_types_present_flag equal to 1 may specify that zero or more pictures that reference an SPS have multiple VCL NAL units, that no VCL NAL units have the same value of nal_unit_type, and that the picture is not an IRAP picture. The syntax element sps_mixed_nalu_types_present_flag equal to 0 may specify that each picture that references an SPS has one or more VCL NAL units, and that the VCL NAL units of each picture that references a PPS have the same value of nal_unit_type.

[0126] If the syntax element no_mixed_nalu_types_in_pic_constraint_flag is equal to 1, the value of the syntax element sps_mixed_nalu_types_present_flag shall be equal to 0 (e.g., it can be determined).

[0127] The syntax element mixed_nalu_types_in_pic_flag equal to 1 may specify that each picture associated with the PH has multiple VCL NAL units, that the VCL NAL units do not have the same value of nal_unit_type, and that the picture is not an IRAP picture. The syntax element mixed_nalu_types_in_pic_flag equal to 0 may specify that each picture associated with the PH has one or more VCL NAL units, and that the VCL NAL units of each picture associated with the PH have the same value of nal_unit_type. If not present, the value of mixed_nalu_types_in_pic_flag may be inferred (e.g., by a decoder) to be equal to 0.

[0128] The above aspects can provide a solution to "Problem 3" described in the Summary section above.

[0129] In accordance with one or more embodiments, the syntax element flag mixed_nalu_types_in_pic_flag may be replaced with the indicator mixed_nalu_types_in_pic_idc.

[0130] An example of a picture parameter set with the syntax element mixed_nalu_types_in_pic_idc is provided in Table 6 below. Table 6 [Table 6]

[0131] The syntax element mixed_nalu_types_in_pic_idc equal to 1 or 2 may specify that each picture that references a PPS has multiple VCL NAL units, that the VCL NAL units do not have the same value of nal_unit_type, and that the picture is not an IRAP picture. The syntax element mixed_nalu_types_in_pic_idc equal to 0 may specify that each picture that references a PPS has one or more VCL NAL units, and that the VCL NAL units of each picture that references a PPS have the same value of nal_unit_type. Other values ​​of the syntax element mixed_nalu_types_in_pic_idc may be reserved for future use by ITU-T|ISO / IEC.

[0132] If the syntax element no_mixed_nalu_types_in_pic_constraint_idc is equal to 1, then the value of mixed_nalu_types_in_pic_idc shall be equal to 0 (eg, can be determined by the decoder).

[0133] For each slice with a value of nal_unit_type nalUnitTypeA in the range IDR_W_RADL to CRA_NUT inclusive, in a picture picA that also contains one or more slices with another value of nal_unit_type (i.e. the value of mixed_nalu_types_in_pic_idc for picture picA is equal to 1), it is possible to implement the following:

[0134] (A) Slice belongs to (eg, can be determined to belong to) subpicA, whose corresponding subpic_treated_as_pic_flag[i] has a value equal to 1.

[0135] (B) The slice shall not belong (eg, it may be determined that it does not belong) to a subpicture of picA that contains a VCL NAL unit with a nal_unit_type not equal to nalUnitTypeA.

[0136] (C) For all subsequent PUs in the CLVS in decoding order, neither RefPicList[0] nor RefPicList[1] of the slices in subpicA shall contain any picture that precedes picA in decoding order in the active entries.

[0137] RefPicList[0] may be the reference picture list used for inter prediction of a P slice or the first reference picture list used for inter prediction of a B slice, and RefPicList[1] may be the second reference picture list used for inter prediction of a B slice.

[0138] For any particular picture's VCL NAL unit, it is possible to implement the following:

[0139] (A) If the syntax element mixed_nalu_types_in_pic_idc is equal to 1, then one or more VCL NAL units shall all have (e.g., be capable of being determined to have) a particular value of nal_unit_type in the range of IDR_W_RADL to CRA_NUT, inclusive, and all other VCL NAL units shall all have (e.g., be capable of being determined to have) a particular value of nal_unit_type in the range of TRAIL_NUT to RSV_VCL_6, inclusive, or equal to GDR_NUT.

[0140] (B) If the syntax element mixed_nalu_types_in_pic_idc is equal to 2, then one or more VCL NAL units all have (e.g., can be determined to have) a particular value of nal_unit_type equal to RASL_NUT or RADL_NUT, inclusive, or equal to GDR_NUT, and all other VCL NAL units all have (e.g., can be determined to have) a particular value of nal_unit_type within the range of TRAIL_NUT to RSV_VCL_6, inclusive, or equal to GDR_NUT, where nal_unit_type is different from the other nal_unit_types.

[0141] (C) Otherwise (mixed_nalu_types_in_pic_idc is equal to 0), the value of nal_unit_type shall be identical (e.g., can be determined to be the same) for all coded slice NAL units of a picture. A picture or PU is referred to as having the same NAL unit type as the coded slice NAL units of the picture or PU.

[0142] If mixed_nalu_types_in_pic_idc of the PPS referenced by a coded picture is equal to 1 or 2, the picture is determined (e.g., by the decoder) as a trailing picture.

[0143] The above aspects can provide a solution to "Problem 4" described in the Summary section above.

[0144] According to one or more embodiments, if the syntax element mixed_nalu_types_in_pic_flag of a picture of layer A is equal to 1, then mixed_nalu_types_in_pic_flag of a picture of layer B, which is a reference layer for layer A, shall be equal to 1 for the same AU (e.g., it may be determined that they are equal).

[0145] The above aspects can provide a solution to "Problem 5" described in the Summary section above.

[0146] In accordance with one or more embodiments, one or more coded video data bitstreams, and syntax structures and elements therein (such as the VCL NAL units and parameter sets described above), may be received by a decoder of this disclosure to decode the received video data. A decoder of this disclosure may decode coded pictures of a video based on VCL NAL units of coded pictures having mixed VCL NAL unit types (e.g., VCL NAL unit (500) shown in FIG. 5) in accordance with embodiments of this disclosure.

[0147] For example, referring to FIG. 6, the decoder (600) may include decoding code (610) configured to cause at least one processor of the decoder (600) to decode coded pictures based on VCL NAL units. According to one or more embodiments, the decoding code (610) can include a decision code (620) that can be configured to: (a) determine or constrain a NAL unit type of one or more VCL NAL units of a coded picture based on a NAL unit type or indicator (e.g., flag) of another one or more VCL NAL units of the coded picture; (b) determine or constrain a picture type of a coded picture based on a NAL unit type or indicator (e.g., flag) of one or more VCL NAL units of the coded picture; (c) determine or constrain a TemporalID of a coded picture based on a VCL NAL unit type or indicator (e.g., flag) of one or more VCL NAL units of the coded picture; and / or (d) determine or constrain a TemporalID of a coded picture based on a VCL NAL unit type or indicator (e.g., flag) of one or more VCL NAL units of the coded picture having mixed VCL NAL unit types, as described in embodiments of the present disclosure. The decoder is configured to cause at least one processor of the decoder (600) to determine or constrain an indicator (e.g., a flag) indicating whether a NAL unit is present based on another received or determined indicator (e.g., a flag).

[0148] The embodiments of the present disclosure may be used separately or in any order in combination. Furthermore, each of the methods, encoders, and decoders of the present disclosure may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0149] The techniques described above may be implemented as computer software using computer readable instructions physically stored on one or more computer readable media. For example, Figure 7 illustrates a computer system (900) suitable for implementing embodiments of the disclosed subject matter.

[0150] Computer software may be coded using any suitable machine code or computer language that may be affected by mechanisms such as assembly, compilation, linking, etc. to generate code including instructions that may be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., either directly or through interpretation, microcode execution, etc.

[0151] The instructions may be executed in various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things (IoT) devices, etc.

[0152] 7 for computer system (900) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement related to any one or combination of components illustrated in the exemplary embodiment of computer system (900).

[0153] The computer system (900) may also include certain human interface input devices that may be responsive to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0154] The input human interface devices may include one or more of a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove, a joystick (905), a microphone (906), a scanner (907), and a camera (908) (although only one of each is depicted).

[0155] The computer system (900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (910), data gloves, or joystick (905), although it is also possible that there are haptic feedback devices that do not function as input devices). For example, such devices can include audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which are capable of outputting two-dimensional visual output or three or more dimensional output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0156] The computer system (900) may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (920) with CD / DVD or similar media (921), thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media such as tape and floppy disks (not shown), specialized ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0157] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0158] The computer system (900) may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, terrestrial TV), vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter (e.g., a USB port on the computer system (900)) that connects to a specific general-purpose data port or peripheral bus (949); other networks are generally integrated into the core of the computer system (900) by connection to a system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system, a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communications can be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to a particular CANbus device), or bidirectional, such as to other computer systems using local or wide area digital networks. Such communications can include communications to cloud computing environments (955). Specific protocols and protocol stacks can be used with each of these networks and network interfaces, as described above.

[0159] The aforementioned human interface devices, human accessible storage, and network interfaces (954) may be attached to the core (940) of the computer system (900).

[0160] The core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), specific programmable processing units in the form of field programmable gate areas (FPGAs) (943), hardware accelerators for specific tasks (944), etc. These devices may be connected via a system bus (948) along with read only memory (ROM) (945), random access memory (RAM) (946), internal mass storage such as internal non-user accessible hard drives, SSDs, and the like (947). In some computer systems, the system bus (948) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (948) or via a peripheral bus (949). Peripheral bus architectures include PCI, USB, etc. A graphics adapter 950 may be included in core 940 .

[0161] The CPU (941), GPU (942), FPGA (943), and accelerator (944) may combine to execute certain instructions that may constitute the computer code described above. The computer code may be stored in ROM (945) or RAM (946). Temporary data may be stored in RAM (946), while persistent data may be stored in, for example, internal mass storage (947). The use of cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage (947), ROM (945), RAM (946), etc., may enable fast storage and retrieval from any memory device.

[0162] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0163] By way of example and not limitation, a computer system having the architecture (900), and in particular the core (940), can provide functionality derived from a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as specific storage of the core (940) of a non-transitory nature, such as the core internal mass storage (947) or ROM (945). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (940). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (940), and in particular the processor therein (including a CPU, GPU, FPGA, etc.), to execute certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM (946) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality from hardwired or otherwise embedded logic in circuitry (e.g., accelerator (944)) that may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may include logic, and vice versa, as appropriate. References to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) that store software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0164] While this disclosure describes some exemplary, non-limiting embodiments, there are modifications, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0165] (Appendix 1) A method executed by at least one processor, comprising: receiving a first video coding layer (VCL) network abstraction layer (NAL) unit of a first slice of a coded picture and a second VCL NAL unit of a second slice of the coded picture, the first VCL NAL unit having a first VCL NAL unit type and the second VCL NAL unit having a second VCL NAL unit type different from the first VCL NAL unit type; decoding the coded picture, the decoding step including determining a picture type of the coded picture based on the first VCL NAL unit type of the first VCL NAL unit and the second VCL NAL unit type of the second VCL NAL unit or based on an indicator received by the at least one processor that indicates that the coded picture contains mixed VCL NAL unit types; The method includes: (Appendix 2) The determining step includes: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a trailing picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 2. The method of claim 1, comprising determining based on: (Appendix 3) The determining step includes: the coded picture being a random access decodable reading (RADL) picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a RADL picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 2. The method of claim 1, comprising determining based on: (Appendix 4) The determining step includes: the coded picture being a step-wise temporal sub-layer access (STSA) picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes an STSA picture coded slice; and the second VCL NAL unit type indicates that the second VCL NAL unit does not contain an immediate decoding refresh (IDR) picture coded slice; and 2. The method of claim 1, comprising determining based on: (Appendix 5) The determining step includes: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a step-wise temporal sublayer access (STSA) picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit does not contain a clean random access (CRA) picture coded slice; 2. The method of claim 1, comprising determining based on: (Appendix 6) The determining step includes: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a gradual decoding refresh (GDR) picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit does not contain an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 2. The method of claim 1, comprising determining based on: (Appendix 7) the indicator is a flag; and The determining step includes: the coded picture being a trailing picture; the flag indicating that the coded picture contains mixed VCL NAL unit types; 2. The method of claim 1, comprising determining based on: (Appendix 8) the indicator is a flag; and The step of decoding the coded picture further comprises: The temporal ID of the coded picture is 0. the flag indicating that the coded picture contains mixed VCL NAL unit types; 2. The method of claim 1, further comprising determining based on: (Appendix 9) the indicator is a flag; and 2. The method of claim 1, further comprising receiving the flag in a picture header or a slice header. (Appendix 10) the indicator is a flag and the coded picture is in a first layer; The method further comprises: receiving the flag; an additional coded picture in a second layer that is a reference layer of the first layer includes a mixed VCL NAL unit type; determining, based on the flag indicating, that the coded picture includes a mixed VCL NAL unit type; The method of claim 1, comprising: (Appendix 11) A system comprising: and a memory configured to store a computer program; at least one processor configured to receive at least one coded video stream, to access said computer program code, and to operate as directed by said computer code; the computer program code comprising: decoding code configured to cause the at least one processor to decode coded pictures from the at least one coded video stream, the decoding code comprising: based on a first video coding layer (VCL) network abstraction layer (NAL) unit type of a first VCL NAL unit of a first slice of the coded picture and a second VCL NAL unit type of a second VCL NAL unit of a second slice of the coded picture; or based on an indicator received by the at least one processor indicating that the coded picture includes a mixed VCL NAL unit type, a decoding code including a determining code configured to cause the at least one processor to determine a picture type of the coded picture; the first VCL NAL unit type is different from the second VCL NAL unit type. (Appendix 12) The decision code is: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a trailing picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 13) The decision code is: the coded picture being a random access decodable reading (RADL) picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a RADL picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit includes an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 14) The decision code is: the coded picture being a step-wise temporal sub-layer access (STSA) picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes an STSA picture coded slice; and the second VCL NAL unit type indicates that the second VCL NAL unit does not contain an immediate decoding refresh (IDR) picture coded slice; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 15) The decision code is: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a step-wise temporal sublayer access (STSA) picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit does not contain a clean random access (CRA) picture coded slice; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 16) The decision code is: the coded picture being a trailing picture; the first VCL NAL unit type indicates that the first VCL NAL unit includes a gradual decoding refresh (GDR) picture coded slice; and the second VCL NAL unit type indicating that the second VCL NAL unit does not contain an immediate decoding refresh (IDR) picture coded slice or a clean random access (CRA) picture coded slice; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 17) the indicator is a flag; and The decision code is: the coded picture being a trailing picture; the flag indicating that the coded picture contains mixed VCL NAL unit types; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 18) the indicator is a flag; and The decision code is: The temporal ID of the coded picture is 0. the flag indicating that the coded picture contains mixed VCL NAL unit types; 12. The system of claim 11, configured to cause the at least one processor to make a decision based on: (Appendix 19) the indicator is a flag; and 12. The system of claim 11, wherein the at least one processor is configured to receive the flag in a picture header or a slice header. (Appendix 20) A non-transitory computer-readable medium storing computer instructions that, when executed by at least one processor, cause the at least one processor to decode coded pictures from at least one coded video stream; The decoding step comprises: based on a first video coding layer (VCL) network abstraction layer (NAL) unit type of a first VCL NAL unit of a first slice of the coded picture and a second VCL NAL unit type of a second VCL NAL unit of a second slice of the coded picture; or based on an indicator received by the at least one processor indicating that the coded picture includes a mixed VCL NAL unit type, determining a picture type of the coded picture; The first VCL NAL unit type is different from the second VCL NAL unit type.

Claims

1. 1. A method executed by at least one processor, comprising: receiving video data including a first video coding layer (VCL) network abstraction layer (NAL) unit of a first slice of a coded picture, a second VCL NAL unit of a second slice of the coded picture, and a mixed flag; determining a type of a first VCL NAL unit of the first slice and a type of a second VCL NAL unit of the second slice based on a value of the mixed flag, and decoding the coded picture; wherein a value of the mixed flag equal to a predetermined value indicates that a type of the first VCL NAL unit is an intra random access point (IRAP) picture and that a type of the second VCL NAL unit is a non-IRAP picture that is not a step-wise temporal sub-layer access (STSA) picture.

2. 2. The method of claim 1, wherein a value of the mixed flag equal to another predefined value indicates that the type of the first VCL NAL unit is the same as the type of the second VCL NAL unit.

3. 2. The method of claim 1, wherein if the value of the mixed flag is equal to a predetermined value, the type of the first VCL NAL unit indicates an immediate decoding refresh (IDR) picture or a clean random access (CRA) picture.

4. 2. The method of claim 1, wherein if the value of the mixed flag is equal to a predetermined value, the type of the second VCL NAL unit indicates a random-access decodable leading (RADL) picture.

5. 2. The method of claim 1, wherein if the value of the mixed flag is equal to a predetermined value, the type of the second VCL NAL unit indicates a trailing (TRAIL) picture or a gradual decoding refresh (GDR) picture.

6. The method of claim 1 , wherein the mixing flag is included in a picture header or a slice header.

7. 2. The method of claim 1, wherein a type of a first VCL NAL unit of the first slice is indicated by a first "nal_unit_type" syntax element and a type of a second VCL NAL unit of the second slice is indicated by a second "nal_unit_type" syntax element; The method, wherein the mixed flag is indicated by a "mixed_nalu_types_in_pic_flag" syntax element and the predetermined value is 1.

8. A computer program product configured to cause said at least one processor to perform a method according to any one of claims 1 to 7.

9. 1. A method executed by at least one processor, comprising: generating a first video coding layer (VCL) network abstraction layer (NAL) unit for a first slice of a picture included in the video data and a second VCL NAL unit for a second slice of the picture included in the video data; setting a value of a mixed flag to a predetermined value if the type of the first VCL NAL unit indicates an intra random access point (IRAP) picture and the type of the second VCL NAL unit indicates a non-IRAP picture that is not a step-wise temporal sub-layer access (STSA) picture, and setting the value of the mixed flag to another predetermined value if the type of the first VCL NAL unit is the same as the type of the second VCL NAL unit; signaling the mixing flag to a decoder that decodes the video data; A method comprising:

10. 1. A method executed by at least one processor, comprising: generating a first video coding layer (VCL) network abstraction layer (NAL) unit for a first slice of a picture included in the video data and a second VCL NAL unit for a second slice of the picture included in the video data; setting a value of a mixed flag to a predetermined value if the type of the first VCL NAL unit indicates an intra random access point (IRAP) picture and the type of the second VCL NAL unit indicates a non-IRAP picture that is not a step-wise temporal sub-layer access (STSA) picture, and setting the value of the mixed flag to another predetermined value if the type of the first VCL NAL unit is the same as the type of the second VCL NAL unit; transmitting the video data including the first VCL NAL unit, the second VCL NAL unit, and the mixed flag as a coded bitstream to a decoder; A method comprising: