Method for identification of random access point and picture types

JP2024164078A5Pending Publication Date: 2025-05-22TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024135554
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2024-08-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing video coding standards lack efficient mechanisms for identifying random access points and picture types in high-level syntax structures, such as NAL unit headers, which complicates error recovery and bitstream manipulation in video decoding.

Method used

Reconfiguring the network abstraction layer (NAL) unit by determining if it is an intra random access picture (IRAP) and decoding it as either an Instant Decoder Refresh (IDR), Broken Link Access (BLA), or Clean Random Access (CRA) based on the previous NAL unit's status, allowing for improved identification and decoding of random access points.

Benefits of technology

Enhances the efficiency of video decoding by reducing the need for redundant coding in NAL unit type identification, improving error recovery and bitstream manipulation, and optimizing decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for coding and decoding a current Network Abstraction Layer (NAL) unit for a video bit stream.SOLUTION: The method includes: determining that a current NAL unit to be an Intra Random Access Picture (IRAP) NAL unit (910); determining whether a previous NAL unit decoded immediately before the current NAL unit indicates an end of a coded video sequence (CVS); based on determining that the previous NAL unit indicates the end of the CVS (Yes in 920), decoding the current NAL unit as one from among an Instantaneous Decoder Refresh (IDR) NAL unit or a Broken Link Access (BLA) NAL unit (930); and based on determining that the previous NAL unit does not indicate the end of the CVS (No in 920), decoding the current NAL unit as a Clean Random Access (CRA) NAL unit, and reconstructing the decoded current NAL unit (940).SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The disclosed technical subject matter relates to video encoding and decoding, and more particularly to including image references in high level syntax structures such as fixed length codepoints and network abstraction layer unit headers.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. §119 to U.S. Provisional Patent Application No. 62 / 786,306, filed in the U.S. Patent and Trademark Office on December 28, 2018, and U.S. Provisional Patent Application No. 16 / 541,693, filed in the U.S. Patent and Trademark Office on August 15, 2019, the disclosures of which are incorporated herein by reference in their entireties. [Background technology]

[0003] Examples of video encoding and decoding using inter-picture prediction with motion compensation have been known for decades. Uncompressed digital video can be composed of a sequence of images, each having a spatial dimension of, for example, 1920x1080 luminance samples and associated chrominance samples. The sequence of images can have a fixed or variable picture rate, also known as a frame rate, for example, 60 pictures / second or 60Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a frame rate of 60Hz) at 8 bits / sample requires a bandwidth approaching 1.5 Gigabits / second (Gbit / s). One hour of such video requires more than 600GB of storage.

[0004] One objective of video encoding and decoding may be the reduction of redundancy in the input video signal through compression. Compression helps to reduce the aforementioned bandwidth or storage requirements, possibly by more than one order of magnitude. Both lossless and lossy compression, as well as combinations thereof, may be utilized. Lossless compression refers to techniques that allow an exact copy of the original signal to be reconstructed from a compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion is application dependent, for example, a user of a given consumer streaming application may tolerate a larger distortion than a user of a television contribution application. The achievable compression ratio may reflect that a larger allowable / tolerable distortion may result in a higher compression ratio.

[0005] Video encoders and decoders may utilize techniques from a number of broad categories, including, for example, motion compensation, transform, quantization, and entropy coding, some of which are introduced below.

[0006] The example of segmenting a coded video bitstream into packets for transport over packet networks has been in use for decades. In the early days, video coding standards and technologies were mostly optimized for bit-oriented transport and bitstreams were defined. Packetization occurred at the system layer interface, specified for example in the Real-Time Transport Protocol (RTP) payload format. With the advent of Internet connectivity suitable for the mass use of video over the Internet, video coding standards reflected its prominent use cases through the conceptual differentiation of the video coding layer (VCL) and the network abstraction layer (NAL). The NAL unit was introduced in H.264 in 2003 and, with only minor modifications, has been maintained in given video coding standards and technologies since then.

[0007] A NAL unit can often be considered as the smallest entity on which a decoder can operate without necessarily decoding all preceding NAL units of a coded video sequence. NAL units enable certain error recovery techniques, as well as certain bitstream manipulation techniques, including bitstream pruning by Media Aware Network elements (MANE), such as Selective Forwarding Units (SFUs) or Multipoint Control Units (MCUs).

[0008] Figure 1 shows relevant parts of the syntax diagram of the NAL unit header according to H.264 (101) and H.265 (102), in both cases without their respective extensions. In both cases, forbidden_zero_bit is a zero bit used for start code emulation prevention in a given system layer environment. The nal_unit_type syntax element indicates the type of data the NAL unit carries, which can be, for example, one of a given slice type, a parameter set type, a Supplementary Enhancement Information (SEI) message, etc. The H.265 NAL unit header further includes nuh_layer_id and nuh_temporal_id_plus1, which indicate the spatial / SNR and temporal layer of the coded image to which the NAL unit belongs.

[0009] It can be seen that NAL unit headers contain only easily parseable fixed-length codewords that do not have any parsing dependency on other data in the bitstream, e.g., other NAL unit headers, parameter sets, etc. Because NAL unit headers are the first octet in a NAL unit, the MANE can easily extract, parse, and act on them. Other high-level syntax elements, e.g., slice or tile headers, in contrast, are less accessible to the MANE because they may require maintaining a parameter set context and / or processing of variable-length or arithmetically encoded codepoints.

[0010] Furthermore, it can be seen that the NAL unit header shown in FIG. 1 does not contain any information that would allow the NAL unit to be associated with a coded image that is composed of multiple NAL units (e.g., multiple tiles or slices, at least some of which are packetized in individual NAL units).

[0011] A given transport technology, such as RTP (RFC 3550), MPEG-system standard, ISO file format, etc., often includes certain information in the form of timing information, such as presentation time, in the example of MPEG and ISO file formats, or capture time, in the example of RTP, that is easily accessible by the MANE and can help to associate each transport unit with a coded picture. However, the semantics of these information may differ between transport / storage technologies and may not have a direct relationship to the picture structure used in the video coding. Thus, these information are at best heuristic and may also not be particularly well suited to identify whether NAL units in a NAL unit stream belong to the same coded picture or not. Summary of the Invention

[0012] In one embodiment, a method for reconstructing a current network abstraction layer (NAL) unit for video decoding using at least one processor is provided, the method including the steps of determining that the current NAL unit is an intra random access picture (IRAP) NAL unit, determining whether a previous NAL unit decoded immediately before the current NAL unit indicates an end of a coded video sequence (CVSA), decoding the current NAL unit as one of an instantaneous decoder refresh (IDR) NAL unit or a broken link access (BLA) NAL unit based on the determination that the previous NAL unit indicates an end of a CVS, decoding the current NAL unit as a clean random access (CRA) NAL unit based on the determination that the previous NAL unit does not indicate an end of a CVS, and reconstructing the decoded current NAL unit.

[0013] In one embodiment, an apparatus is provided for reconstructing current Network Abstraction Layer (NAL) units for video decoding, the apparatus including at least one memory configured to store program code and at least one processor configured to read the program code and operate as directed by the program code. The program code includes a first decision code configured to cause the at least one processor to determine that a current NAL unit is an intra random access picture (IRAP) NAL unit; a second decision code configured to cause the at least one processor to determine whether a previous NAL unit decoded immediately before the current NAL unit indicates an end of a coded video sequence (CVS); a first decoding code configured to cause the at least one processor to decode the current NAL unit as one of an instantaneous decoder refresh (IDR) NAL unit or a broken link access (BLA) NAL unit based on a determination that the previous NAL unit indicates the end of a CVS; a second decoding code configured to cause the at least one processor to decode the current NAL unit as a clean random access (CRA) NAL unit based on a determination that the previous NAL unit does not indicate the end of a CVS; and a reconstruction code configured to cause the processor to reconstruct the decoded current NAL unit.

[0014] In one embodiment, a non-transitory computer-readable medium having stored thereon one or more instructions that, when executed by one or more processors of an apparatus for reconstructing a current network abstraction layer (NAL) unit for video decoding, cause the one or more processors to determine that the current NAL unit is an intra random access picture (IRAP) NAL unit, determine whether a previous NAL unit decoded immediately before the current NAL unit indicates an end of a coded video sequence (CVS), decode the current NAL unit as one of an instantaneous decoder refresh (IDR) NAL unit or a broken link access (BLA) NAL unit based on the determination that the previous NAL unit indicates the end of a CVS, decode the current NAL unit as a clean random access (CRA) NAL unit based on the determination that the previous NAL unit does not indicate the end of a CVS, and reconstruct the decoded current NAL unit. [Brief description of the drawings]

[0015] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of a NAL unit header according to H.264 and H.265. [Diagram 2] FIG. 2 is a simplified block diagram of a communication system according to one embodiment. [Diagram 3] FIG. 3 is a simplified block diagram of a communication system according to one embodiment. [Figure 4] FIG. 4 is a simplified block diagram schematic of a decoder according to one embodiment. [Diagram 5] FIG. 5 is a simplified block diagram schematic of an encoder according to one embodiment. [Figure 6]FIG. 6 is a schematic diagram of a NAL unit header according to one embodiment. [Figure 7] FIG. 7 is a schematic diagram of an image of a random access point and a leading image according to one embodiment. [Figure 8] FIG. 8 is a schematic diagram of syntax elements in a parameter set according to one embodiment. [Figure 9] FIG. 9 is a flow diagram of an exemplary process for reconstructing current Network Abstraction Layer (NAL) units for video decoding according to one embodiment. [Figure 10] FIG. 10 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] Problem to be solved Existing video coding syntax lacks an easily identifiable / parsable syntax element that identifies the leading picture type in high-level syntax structures such as random access points, random access picture types, and NAL unit headers.

[0017] FIG. 2 shows a simplified block diagram of a communication system (200) according to one embodiment. The system (200) may include at least two terminals (210-220) interconnected via a network (250). For unidirectional transmission of data, a first terminal (210) may encode video data at a local location for transmission to the other terminal (220) via the network (250). The second terminal (220) may receive the other terminal's encoded video data from the network (250), decode the encoded data, and display the recovered video data. Unidirectional data transmission may be common in media provisioning applications, etc.

[0018] 2 illustrates a second pair of terminals (230, 240) provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For the bidirectional transmission of data, each terminal (230, 240) may encode video data captured at a local location for transmission to the other terminal over the network (250). Each terminal (230, 240) may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the recovered video data on a local display device.

[0019] In FIG. 2, the terminals (210-240) may be depicted as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network (250) represents any number of networks that convey encoded video data between the terminals (210-240), including, for example, wired and / or wireless communication networks. The communication network (250) may exchange data in circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of network (250) is not important to the operation of the present disclosure, unless otherwise described below.

[0020] 3 shows the arrangement of video encoders and decoders in a streaming environment as one example of an application for the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0021] The streaming system may include a capture subsystem (313), such as a video source (301), e.g., a digital camera, that generates an uncompressed video sample stream (302). The sample stream (302) is shown as a thick line to emphasize its large data volume when compared to an encoded video bitstream, and may be processed by an encoder (303) coupled to the camera (301). The encoder (303) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (304) is shown as a thin line to emphasize its smaller data volume when compared to the sample stream, and may be stored in a streaming server (305) for future use. One or more streaming clients (306, 308) may access the streaming server (305) to retrieve copies (307, 309) of the encoded video bitstream (304). The client (306) may include a video decoder (310) that decodes an input copy of the encoded video bitstream (307) and generates an output video sample stream (311) that can be rendered on a display (312) or other rendering device (not shown). In some streaming systems, the video bitstreams (304, 307, 309) may be encoded according to a predetermined video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. A developing video encoding standard is informally known as Versatile Video Coding, or VVC. The disclosed subject matter may be used in the context of VVC.

[0022] FIG. 4 may be a functional block diagram of a video decoder (310) according to one embodiment of the disclosure.

[0023] The receiver (410) can receive one or more codec video sequences to be decoded by the decoder (310), one coded video sequence at a time, in the same or another embodiment, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (412), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (410) can receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which can be transferred using respective entities (not shown). The receiver (410) can separate the coded video sequences from the other data. To try to eliminate network jitter, a buffer memory (415) can be coupled between the receiver (410) and the entropy decoder / parser (420) (hereinafter the “parser”). If the receiver 410 is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network, then the buffer 415 may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer 415 may be needed but may be relatively large and may advantageously be of adaptive size.

[0024] The video decoder (310) may include a parser (420) for reconstructing symbols (421) from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder (310) and potentially information to control a rendering device, such as a display (312) that is not an integral part of the decoder but may be connected to the decoder, as shown in FIG. 3. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (420) may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence may follow any video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (420) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0025] The parser (420) can perform entropy decoding / parsing operations on the video sequence received from the buffer (415) to generate symbols (421).

[0026] The reconstruction of the symbols (421) may involve a number of different units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.) and other factors. Which units participate and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (420). The flow of such subgroup control information between the parser (420) and the following units is not shown for clarity.

[0027] In addition to the above-mentioned functional blocks, the decoder 310 may be conceptually subdivided into several functional units, as described below. In a practical implementation operating under commercial constraints, many of these units may closely interact with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed technical matter, a conceptual subdivision into the following functional units is appropriate:

[0028] The first unit is a scalar / inverse transform unit (451), which receives quantized transform coefficients as symbols (421) from the parser (420) as well as control information including the transform used, block size, quantization coefficients, quantization scaling matrices, etc. It can output blocks containing sample values ​​that can be input to an aggregator (455).

[0029] In some cases, the output samples of the scalar / inverse transform (451) may belong to intra-coded blocks, i.e. blocks that do not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed part of the current image. Such prediction information may be provided by an intra-image prediction unit (452), which in some cases uses surrounding already reconstructed information fetched from the current (partially reconstructed) image (456) to generate a block of the same size and shape as the block being reconstructed. The aggregator (455) optionally adds, sample by sample, the prediction information generated by the intra-prediction unit (452) to the output sample information provided by the scalar / inverse transform unit (451).

[0030] In other cases, the output samples of the scalar / inverse transform unit (451) may belong to an inter-coded and potentially motion compensated block. In such a case, the motion compensation prediction unit (453) may access a reference picture memory (457) to fetch samples used for prediction. After motion compensation of the fetched samples according to the symbols (421) belonging to the block, these samples may be added by an aggregator (455) to the output of the scalar / inverse transform unit (in this case, called residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory form from which the motion compensation unit fetches the prediction samples may be controlled by a motion vector available to the motion compensation unit, for example, in the form of a symbol (421) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0031] The output samples of the aggregator (455) may be subject to various loop filtering techniques in a loop filter unit (456). Video compression techniques may include in-loop filter techniques that are controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (456) as symbols (421) from the parser (420), but may also be responsive to meta-information obtained during decoding of previous portions of the coded image or coded video sequence (in decoding order), as well as to previously reconstructed and loop filtered sample values.

[0032] The output of the loop filter unit (456) may be a sample stream that may be output to the rendering device (312) as well as stored in a reference picture memory for use in future inter-picture prediction.

[0033] Once a given coded image is fully reconstructed, it can be used as a reference image for future prediction. Once a coded image is fully reconstructed and identified as a reference image (e.g., by the parser (420)), the current reference image (456) becomes part of the reference image buffer (457) and a fresh current image memory can be reallocated before starting the reconstruction of the next coded image.

[0034] The video decoder 420 may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as ITU-T Rec. H.265. The encoded video sequence may comply with the syntax defined by the video compression technique or standard being used, in the sense of following the syntax of the video compression technique or standard as defined in the video compression technique document or standard, and in particular in the profile documents therein. Also required for compliance is that the complexity of the encoded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited, in some cases, through a Hypothetical Reference Decoder (HRD) specification and HRD buffer metadata signaled with the encoded video sequence.

[0035] In one embodiment, the receiver (410) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (420) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0036] FIG. 5 may be a functional block diagram of a video encoder (303) according to one embodiment of the present disclosure.

[0037] The encoder (303) can receive video samples from a video source (301) (not part of the encoder) that can capture video images to be encoded by the encoder (303).

[0038] The video source (301) may provide a source video sequence to be encoded by the encoder (303) in the form of a digital video sample stream, which may be in any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (303) may be a camera capturing local image information as a video sequence. The video data may be provided as a number of individual images that convey motion when viewed in sequence. The images themselves may be organized as a spatial array of pixels, where each pixel may contain one or more samples, depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0039] According to one embodiment, the encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (543) in real-time or under any other time constraint required by the application. Enforcing an appropriate encoding rate is one function of the controller (550). The controller controls other functional units, as described below, and is operatively coupled to these units. Coupling is not shown for clarity. Parameters set by the controller may include rate control related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. One skilled in the art can easily identify other functions of the controller (550) as they may be relevant to a video encoder (303) optimized for a given system design.

[0040] Some video encoders operate in what one skilled in the art would readily recognize as a "coding loop." As an oversimplified explanation, the coding loop consists of the coding portion of the encoder (530) (hereafter the "source coder") (responsible for generating symbols based on the input picture to be encoded and the reference picture), and a (local) decoder (533) embedded within the encoder (303) that reconstructs the symbols to generate sample data that the (remote) decoder will also generate (so that in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to a reference picture memory (534). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the contents of the reference picture buffer are also bit-perfect between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values ​​as the reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, e.g., due to channel errors) is well known to those skilled in the art.

[0041] The operation of the "local" decoder (533) may be the same as the "remote" decoder (310), as already described in detail above in connection with FIG. 4. However, with brief reference again to FIG. 4, because symbols are available and the encoding / decoding of symbols for the encoded video sequence by the entropy coder (545) and parser (420) may be lossless, the entropy decoding portion of the decoder (310), including the channel (412), receiver (410), buffer (415), and parser (420), may not be implemented entirely within the local decoder (533).

[0042] An observation that can be made at this point is that any decoder technique, except for analysis / entropy decoding, that is present in the decoder must also be present in a substantially identical functional form in the corresponding encoder. For this reason, the technical matters disclosed will focus on the operation of the decoder. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques that have been generically described. Only in certain areas is a more detailed explanation necessary, and is provided below.

[0043] As part of its operation, the source coder (530) may perform motion compensated predictive coding, which predictively codes an input frame with respect to one or more previously coded-frames from a video sequence designated as “reference frames.” In this manner, the coding engine (532) codes differences between pixel blocks of an input frame and pixel blocks of reference frames that may be selected as prediction references for the input frame.

[0044] The local video decoder (533) can decode the encoded video data of frames that may be designated as reference frames based on the symbols generated by the source coder (530). The operation of the encoding engine (532) can advantageously be a lossy process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence can be a replica of the source video sequence, typically with some errors. The local video decoder (533) replicates the decoding process performed by the video decoder on the reference frames and can cause the reconstructed reference frames to be stored in a reference image cache (534). In this way, the encoder (303) can locally store (without transmission errors) copies of reconstructed reference frames that have common content as reconstructed reference frames obtained by a far-end video decoder.

[0045] The predictor (535) may perform a prediction search for the coding engine (532). That is, for a new frame to be coded, the predictor (535) may search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or predefined metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (535) may operate on a sample block-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (535), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (534).

[0046] The controller (550) can manage the encoding operations of the video coder (530), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0047] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (545), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0048] The sender (540) can buffer the encoded video sequence as it is generated by the entropy coder (545) and prepare it for transmission over a communication channel (560), which may be a hardware / software link to a storage device that stores the encoded video data. The sender (540) can merge the encoded video data from the video coder (530) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0049] A controller (550) can manage the operation of the encoder (303). During encoding, the controller (550) can assign a predefined encoding image type to each encoded image, which can affect the encoding technique that can be applied to the respective image. For example, a picture can often be assigned as one of the following frame types:

[0050] An Intra picture (I-picture) may be one that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow different types of Intra pictures, including, for example, Independent Decoder Refresh Pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective applications and functions.

[0051] A predictive picture (P picture) may be one that can be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0052] A Bi-directionally Predictive Picture (B-picture) may be one that can be encoded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use two or more reference pictures and associated metadata for the reconstruction of a single block.

[0053] A source image is generally spatially subdivided into a number of sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the image of the block each. For example, blocks of I-pictures may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same image (spatial or intra-prediction). Pixel blocks of P-pictures may be coded non-predictively, via spatial prediction, or via temporal prediction, with reference to one previously coded reference image. Blocks of B-pictures may be coded non-predictively, via spatial prediction, or via temporal prediction, with reference to one or two previously coded reference images.

[0054] The video coder (303) may perform encoding operations in accordance with a predefined video encoding technique or standard, such as ITU-T Rec. H.265. In its operation, the video coder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. The encoded video data may therefore conform to a syntax specified by the video encoding technique or standard being used.

[0055] In one embodiment, the transmitter (540) can transmit additional data along with the coded video. The video coder (530) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Video Usability Information (VUI) parameter set fragments, etc.

[0056] FIG. 7 illustrates various types of random access mechanisms and their associated picture types. In particular, a Clean Random Access picture (701), a Random Access Skipped Leading picture (702), and a Random Access Decodable Leading picture (703) are shown. Possible predictive relationships from these various random access picture types to non-random access pictures are also shown. As an example, RASL pictures (701) may be labeled as B15, B16, and B17, RADL pictures 703 may be labeled as B18 and B19, and CRA pictures 701 may be labeled as I20 and B21. In this example, the designation B may indicate a B picture that contains motion compensated difference information associated with a previously decoded picture, and the designation I may indicate an I picture that is coded independently of other pictures. The numbers following this designation can indicate the order of the pictures.

[0057] An Instantaneous Decoding Refresh (IDR) picture (not shown) may be used for random access. Pictures following the IDR in decoding order may not reference previously decoded pictures of the IDR picture. This restriction may result in lower coding efficiency.

[0058] A Clean Random Access (CRA) picture (701) may be used to randomly access a reference picture decoded before the CRA picture by allowing a leading picture following the CRA picture in decoding order but before the CRA picture in output order. CRA may provide better coding efficiency than IDR because it may exploit an open GOP structure for inter prediction. A Broken Link Access (BLA) picture (not shown) may be a CRA picture with a broken link. This is a position in the bitstream where some pictures following in decoding order are indicated as likely to contain significant visual artifacts due to unspecified operations performed in generating the bitstream.

[0059] A random access, skip leading (RASL) picture (702) may not be decoded correctly when random access from this CRA picture occurs because some relevant reference pictures may not be present in the bitstream. A random access, decodable leading (RADL) picture (703) may be decoded correctly when random access from this CRA picture occurs because the RADL picture does not need to use the previous reference pictures of the CRA picture as references.

[0060] As mentioned above, a given video coding technology or standard includes at least five different types of random access mechanisms and associated picture types. According to a given related technology, such as H.265, each of these picture types may require a numbering space in the NAL unit type field, which is described below. The numbering space in this field is valuable and expensive in terms of coding efficiency. Therefore, from the viewpoint of coding efficiency, it is desirable to reduce the use of the numbering space.

[0061] Referring to FIG. 6, an example of a NAL unit header (601) is shown in one embodiment. A predefined value of the NAL unit type (602) may indicate that the current picture contained within the NAL unit is identified as a random access point picture or a leading picture. The NAL unit type value may be coded, for example, as an unsigned integer of a predefined length, here 5 bits (603). A single combination of these 5 bits may be represented as IRAP_NUT (IRAP Nal Unit Type) in a video compression technology or standard. The numbering space required in the NAL unit type field may thereby be reduced, for example, from 5 entries (as in H.265) to, for example, 2 entries: one IDR and the other IRAP_NUT.

[0062] Referring to FIG. 8, in one embodiment, a flag RA-derivation-flag in an appropriate high-level syntax structure, such as a sequence parameter set, can indicate whether the Intra Random Access Point (IRAP) type (IDR, CRA, or BLA) or the leading picture type (RASL or RADL) is indicated by explicit signaling in the NAL unit header, tile group header, or slice header, or is implicitly derived from other information, including a coded video sequence start indication or picture order count, as described below. As one example, the flag (802) can be included in a decoding parameter set (801) or a sequence parameter set. When the ra_derivation-flag indicates an implicit derivation, the NUT value can be reused for other purposes.

[0063] In one embodiment, a flag leading_pic_type_derivation_flag (803) in an appropriate high-level syntax structure, such as a parameter set, can indicate whether the leading picture type (RASL or RADL) is explicitly identified by a flag that may be present in an appropriate high-level syntax structure associated with the picture or portion thereof, such as a slice header or tile group header, or whether it is implicitly identified by an associated IRAP type, Picture Order Count (POC) value, or Reference Picture Set (RPS) information.

[0064] Described below are certain derivation / implication mechanisms that may be used when one, or possibly both, of the above flag signals are enabled, or when the above flags are not present and the flag's value is implicitly inferred to be equal to a certain value.

[0065] In one embodiment, if the NAL unit decoded before the IRAP NAL unit identifies the end of the coded video sequence (e.g., by an EOS NAL unit), then the IRAP picture (indicated by a NAL unit with NUT equal to IRAP_NUT) may be treated as an IDR or BLA picture. If the NAL unit decoded before the IRAP NAL unit is not equivalent to the end of the coded video sequence, then the IRAP picture may be treated as a CRA picture. This means that there is no need to waste code points to differentiate between CRA and IDR / BLA, since the difference can be implied from the end of the CVS.

[0066] In one embodiment, the value of the NAL unit type (e.g., IRAP_NUT) may indicate that the current picture contained within the NAL unit is identified as a leading picture. A leading picture may be decoded after a random access picture, but may be displayed before a random access picture (not necessarily new).

[0067] In one embodiment, the leading picture type (RASL or RADL) may be identified implicitly by interpreting the associated IRAP type. For example, if the associated IRAP type is CRA or BLA, the leading picture may be processed as a RASL picture, but if the associated IRAP type is IDR, the leading picture may be processed as a RADL. This may require the decoder to maintain state, and in particular, which IRAP pictures have been previously decoded.

[0068] In one embodiment, the leading picture type (RASL or RADL) may be indicated by an additional flag that may be present in the NAL unit header, slice header, tile group header, parameter set, or any other suitable high-level syntax structure. It has been determined that the identification of leading pictures and their types is necessary or useful in the following cases: (a) the bitstream starts with a CRA picture; (b) the CRA picture is randomly accessed, i.e., the decoding starts with a CRA picture; (c) a sequence of pictures starting with a CRA picture is spliced ​​after another sequence of coded pictures, and an EOS NAL unit is inserted between these two sequences; and (d) a file writer takes as input a bitstream containing a CRA picture and creates an ISOBMFF file with sync sample indications, SAP sample grouping, and / or is_leading indication.

[0069] In one embodiment, no leading picture type is assigned in the numbering space of the NUT to indicate the leading picture type. Instead, the leading picture type may be derived from the POC value between the associated IRAP picture and the next picture in decoding order. Specifically, a non-IRAP picture that is decoded after an IRAP picture and has a POC value smaller than the POC value of the IRAP picture may be identified as a leading picture. A non-IRAP picture that is decoded after an IRAP picture and has a POC value larger than the POC value of the IRAP picture may be identified as a trailing picture.

[0070] In one embodiment, if the decoded NAL unit before the IRAP NAL unit is equivalent to the end of the coded video sequence, the decoder internal state CvsStartFlag may be set equal to 1. When CvsStartFlag is equal to 1, the IRAP picture may be processed as an IDR or BLA picture, and all leading pictures associated with the IRAP picture may be processed as IDR or BLA pictures. When CvsStartFlag is equal to 1, if the NAL unit type of the IRAP NAL unit is equal to CRA, HandleCraAsCvsStartFlag may be set equal to 1.

[0071] In one embodiment, if some external means are available to change the decoder state without decoding the bitstream to set the variable HandleCraAsCvsStartFlag to the value of the current picture, the variable HandleCraAsCvsStartFlag may be set equal to the value provided by the external means, and the variable NoIncorrectPicOutputFlag (which may be referred to as NoRaslOutputFlag) may be set equal to HandleCraAsCvsStartFlag.

[0072] In one embodiment, when the variable HandleCraAsCvsStartFlag is equal to 1, all leading pictures associated with the CRA picture may be discarded without decoding.

[0073] In one embodiment, when the variable NoIncorrectPicOutputFlag is equal to 1, all RASL pictures associated with the CRA picture may be discarded without decoding.

[0074] 9 is a flowchart of an example process 900 for reconstructing a current network abstraction layer (NAL) unit for video decoding. In some implementations, one or more process blocks of FIG. 9 may be performed by the decoder 310. In some implementations, one or more process blocks of FIG. 9 may be performed by another device or group of devices, such as the encoder 303, separate from or including the decoder 310.

[0075] As shown in FIG. 9, process 900 may include determining that a current NAL unit is an intra random access picture (IRAP) NAL unit (up to block E10).

[0076] As further shown in FIG. 9, process 900 may include determining whether a previous NAL unit decoded immediately before the current NAL unit indicates the end of a coded video sequence (CVS) (block 920).

[0077] As further shown in FIG. 9, process 900 may include decoding the current NAL unit as one of an instantaneous decoding refresh (IDR) NAL unit or a broken link access (BLA) NAL unit based on determining that the previous NAL unit indicates the end of the CVS, and reconstructing the decoded current NAL unit (block 930).

[0078] As further shown in FIG. 9, process 900 may include decoding the current NAL unit as a clean random access (CRA) NAL unit based on determining that the previous NAL unit does not indicate the end of the CVS, and reconstructing the decoded current NAL unit (block 940).

[0079] In one embodiment, process 900 may further include determining whether to decode the current NAL unit as an IDR NAL unit or a BLA NAL unit based on a header of the current NAL unit.

[0080] In one embodiment, the NAL unit type of the current NAL unit may be an IRAP NAL unit type.

[0081] In one embodiment, the NAL unit type of the previous NAL unit may be a CVS End of Stream (EOS) NAL unit type.

[0082] In one embodiment, process 900 may further include setting a first flag based on determining that the previous NAL unit indicates the end of the CVS.

[0083] In one embodiment, the process 900 may further include treating the leading picture associated with the current NAL unit as a random access decodable leading (RADL) picture based on the first flag being set.

[0084] In one embodiment, process 900 may further include setting a second flag based on the NAL unit type of the current NAL unit being a CRA NAL unit type while the first flag is set.

[0085] In one embodiment, the process 900 may further include setting a second flag based on the external signal, and setting a third flag corresponding to the second flag.

[0086] In one embodiment, process 900 may further include discarding a random access skip reading (RASL) picture associated with the current NAL unit without decoding the RASL picture based on a third flag being set.

[0087] In one embodiment, the process 900 may further include discarding the leading picture associated with the current NAL unit without decoding the leading picture based on the second flag being set.

[0088] 9 illustrates example blocks of process 900, in some implementations process 900 may include additional, fewer, different, or differently arranged blocks than those illustrated in FIG 9. Additionally or alternatively, two or more of the blocks of process 900 may be performed in parallel.

[0089] Additionally, the proposed methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0090] The techniques for image referencing in network abstraction unit headers described above may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, FIG. 10 illustrates a computer system 1000 suitable for implementing certain embodiments of the disclosed subject matter.

[0091] Computer software may be encoded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code containing instructions. The instructions may be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc. directly, or through interpretation, microcode execution, etc.

[0092] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things, etc.

[0093] 10 for computer system 1000 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement related to any one or combination of components illustrated in the exemplary embodiment of computer system 1000.

[0094] The computer system 1000 may include certain human interface input devices that may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that do not necessarily directly involve conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic images).

[0095] The input human interface devices may include one or more of a keyboard 1001, a mouse 1002, a trackpad 1003, a touch screen 1010, a data glove 1004, a joystick 1005, a microphone 1006, a scanner 1007, and a camera 1008 (only one of each is depicted).

[0096] Computer system 1000 may also include certain human interface output devices that can stimulate one or more of the senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 1010, data gloves 1204, or joystick 1005, although haptic feedback devices that do not function as input devices may also be present), audio output devices (e.g., speakers 1009, headphones (not shown)), visual output devices (screens 1010 including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light emitting diode (OLED) screens, each with or without touchscreen input capability, each with or without haptic feedback capability - some of which may output two-dimensional visual output or three or more dimensional output through means such as stereoscopic output, such as virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0097] Computer system 1000 may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 1020 using media such as CD / DVD 1021, thumb drives 1022, removable hard drives or solid state drives 1023, legacy magnetic media such as tapes and floppy disks, specialized ROM / ASIC / PLD based devices such as security dongles (not shown), and the like.

[0098] Those skilled in the art should also understand that the term "computer readable media" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitory signals.

[0099] The computer system 1000 may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks may include Ethernet, wireless LAN, cellular networks including Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), long-term evolution (LTE), etc., wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial including CANBus, etc. A given network generally requires an external network interface adapter that is attached to a given general purpose data port or peripheral bus (1049). Some may be connected to the peripheral bus 1049 using a network interface 1054, such as a Universal Serial Bus (USB) port of the computer system 1000, while others are typically built into the core of the computer system 1000 by attachment to a system bus described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). As an example, a network 1055 may be connected to the peripheral bus 1049 using a network interface 1054. Using any of these networks, the computer system 1000 may communicate with other entities. Such communications may be uni-directional, receive-only (e.g., broadcast television), uni-directional transmit-only (e.g., a CAN bus to a given CAN bus device), or bi-directional, for example, to other computer systems using local or wide area digital networks. Predetermined protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0100] The human interface devices, human accessible storage devices, and network interfaces mentioned above may be attached to the core 1040 of the computer system 1000 .

[0101] The core 1040 may include one or more central processing units (CPUs) 1041, graphics processing units (GPUs) 1042, specialized programmable processing units in the form of field programmable gate arrays (FPGAs) 1043, hardware accelerators for certain tasks 1044, etc. These devices may be connected through a system bus 1248 along with read only memory (ROM) 1045, random access memory (RAM) 1046, internal mass storage devices 1047 such as internal non-user accessible hard drives, solid state drives (SSDs), etc. In some computer systems, the system bus 1248 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 1248 or through a peripheral bus 1049. Peripheral bus architectures include peripheral component interconnect (PCI), USB, etc.

[0102] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 can execute predetermined instructions that, in combination, may constitute the computer code described above. The computer code may be stored in ROM 1045 or RAM 1046. Transient data may also be stored in RAM 1046, while permanent data may be stored, for example, in an internal mass storage device 1047. Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU 1041, GPU 1042, mass storage device 1047, ROM 1045, RAM 1046, etc.

[0103] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The media and computer code can be specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0104] As an example, and not by way of limitation, a computer system having architecture 1000, and specifically core 1040, can provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage devices as described above, as well as storage devices associated with core 1040 that are non-transitory in nature, such as core internal mass storage device 1047 or ROM 1045. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by core 1040. Computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause core 1040, and specifically the processors therein (including CPU, GPU, FPGA, etc.) to execute particular processes or particular portions of particular processes described herein. This includes defining data structures stored in RAM 1046 and modifying such data structures according to a process defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1044). Circuitry may operate in place of, or in conjunction with, software to perform particular processes or particular portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. This disclosure encompasses any appropriate combination of hardware and software.

[0105] While this disclosure has described several exemplary embodiments, there exist modifications, substitutions, and various substitute equivalents, which are within the scope of this disclosure. Thus, it will be appreciated that those skilled in the art will be able to devise many systems and methods that embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure, even if not explicitly shown or described herein.

Claims

1. A method of video encoding in an encoder, comprising: generating a video bitstream; storing the generated video bitstream; Including, The step of generating a video bitstream comprises: determining that the current Network Abstraction Layer (NAL) unit is not an Intra Random Access Point (IRAP) NAL unit; determining that a current picture included in the current NAL unit is a leading picture; setting a first flag indicating whether the type of the leading picture is explicitly signaled in the video bitstream; encoding a type of the leading picture based on the first flag; encoding the current NAL unit based on a type of the leading picture; A method comprising:

2. The current picture is determined to be the leading picture based on a Picture Order Count (POC) value of the current picture and a POC value of an IRAP picture associated with the current picture. The method of claim 1.

3. The current picture is determined to be the leading picture based on the POC value of the current picture being less than the POC value of the IRAP picture. The method of claim 2.

4. The step of generating the video bitstream further comprises: obtaining a second flag indicating the type of the leading picture based on the first flag indicating that the type of the leading picture is explicitly signaled in the video bitstream; determining a type of the leading picture based on the second flag; Including, The type of the leading picture is at least one of a random access skip leading (RASL) type and a random access decodable leading (RADL) type; The method of claim 1.

5. The step of generating the video bitstream further comprises: determining a type of the leading picture based on at least one of a type of an IRAP picture associated with the current picture, a Picture Order Count (POC) value of the current picture, or Reference Picture Set (RPS) information based on the first flag indicating that the type of the leading picture is not explicitly signaled in the video bitstream; The method of claim 1.

6. Based on whether the type of the IRAP picture is a clean random access (CRA) type or a broken link access (BLA) type, the type of the leading picture is determined to be a random access skip leading (RASL) type. The method according to claim 5.

7. Based on the type of the IRAP picture being an Instantaneous Decoder Refresh (IDR) type, it is determined that the type of the leading picture is a Random Access Decodable Leading (RADL) type. The method according to claim 5.

8. An apparatus for video encoding, comprising: at least one memory storing program code including a plurality of computer instructions; at least one processor that reads and executes the program code; The computer instructions, when executed by the at least one processor, cause a computer to perform the method of any one of claims 1 to 7. Device.

9. A method for causing a computer to carry out the method according to any one of claims 1 to 7. Computer program.

10. A method of video encoding in an encoder, comprising the steps of: generating a video bitstream; storing the generated video bitstream; Including, The step of generating a video bitstream comprises: determining that the current NAL unit is an IRAP NAL unit; determining whether a previous NAL unit coded immediately before the current NAL unit indicates an end of a coded video sequence (CVS); encoding the current NAL unit as one of an IDR NAL unit or a BLA NAL unit based on a determination that the previous NAL unit indicates the end of the CVS; encoding the current NAL unit as a CRA NAL unit based on a determination that the previous NAL unit does not indicate the end of the CVS; A method comprising:

11. The step of generating the video bitstream further comprises: determining whether to encode the current NAL unit as the IDR NAL unit or the BLA NAL unit, and encoding a header of the current NAL unit. The method of claim 10.

12. The NAL unit type of the current NAL unit is an IRAP NAL unit type. The method of claim 10.

13. The NAL unit type of the previous NAL unit is a CVS end-of-stream (EOS) NAL unit type. The method of claim 10.

14. The step of generating the video bitstream further comprises: based on a determination that the previous NAL unit indicates the end of the CVS, setting a first flag; The method of claim 10.

15. The step of generating the video bitstream further comprises: Based on the first flag being set, processing a leading picture associated with the current NAL unit as a RADL picture image; The method of claim 14.

16. An apparatus for video encoding, comprising: at least one memory storing program code including a plurality of computer instructions; at least one processor that reads and executes the program code; The computer instructions, when executed by the at least one processor, cause a computer to perform the method of any one of claims 10 to 15. Device.

17. A method for causing a computer to carry out the method according to any one of claims 10 to 15. Computer program.