Signaling desired NAL units
By using the required container NAL units to mark key SEI messages and other NAL units during the video encoding and decoding process, the problem that SEI messages may be discarded in the prior art is solved, and the guarantee of video decoding quality and consistency is achieved.
Patent Information
- Application Number
- CN202480004463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-04
- Filing Date
- 2024-09-05
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively process auxiliary enhancement information (SEI messages) during video encoding and decoding, especially in the case of network congestion or computing resource limitations, resulting in the possibility of SEI messages being discarded, affecting the quality and consistency of video decoding.
By introducing the required container NAL units into the network abstraction layer (NAL) unit stream, certain SEI messages and other NAL units are clearly marked as not to be discarded, and corresponding processing mechanisms are implemented in the decoder and intermediate box to ensure that critical SEI messages are correctly processed and delivered during video decoding.
The effective processing and delivery of key SEI messages during the video encoding and decoding process is realized, ensuring the quality and consistency of video decoding, especially in the case of network congestion or computing resource limitations.
Smart Images

Figure CN120077664A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 536,694, filed on September 5, 2023, and U.S. Application No. 18 / 824,367, filed on September 4, 2024. The entire contents of the above - mentioned priority applications are incorporated herein by reference. Technical Field
[0002] The disclosed subject matter relates to video encoding and decoding, and more particularly to mechanisms for signaling the required processing of SEI messages in a system. Background Art
[0003] The use of inter - picture prediction with motion compensation for video encoding and decoding has been around for decades. Uncompressed digital video can consist of a sequence of pictures, each having a spatial size of, for example, luminance samples of 1920×1080 and associated chrominance samples. The sequence of pictures can have a fixed or variable picture rate (also informally called the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has high bit - rate requirements. For example, a 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.
[0004] One goal of video encoding and decoding can be to reduce redundancy in the input video signal through compression. Compression can help reduce the above - mentioned bandwidth or storage - space requirements, which can be reduced by two orders of magnitude or more in some cases. Lossless compression, lossy compression, and combinations thereof can be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application. For example, users of some consumer streaming applications can tolerate higher distortion than users of television - like applications. The achievable compression ratio can reflect that higher admissible / acceptable distortion can result in a higher compression ratio.
[0005] Video encoders and video decoders can utilize a wide variety of broad classes of techniques, including, for example, motion compensation, transformation, quantization, and entropy coding, some of which will be described below.
[0006] Some video coding specifications and standards (including ITU-T H.266 v2) include Supplementary Enhancement Information (SEI messages), and these specifications and standards are incorporated herein by reference in their entirety. In these specifications or standards, by definition, SEI messages are not required for decoding luma or chroma sample data. Some system techniques may discard such SEI messages under certain conditions. Summary of the Invention
[0007] The disclosed subject matter relates to video encoding and decoding, and more particularly to mechanisms for signaling the required processing of SEI messages by a system.
[0008] According to one aspect of the present disclosure, a method of video decoding is performed by at least one processor of a decoder, the method comprising: receiving (i) one or more encoded pictures and (ii) a Network Abstraction Layer (NAL) unit stream, the NAL unit stream including at least one first NAL unit of a first type; interpreting the first NAL unit; and decoding at least one of the one or more encoded pictures based on the interpretation of the first NAL unit; wherein the decoder is notified that the first NAL unit cannot be discarded by the decoder by at least one of: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicating a second NAL unit type that is a NAL unit in the NAL unit stream that is not required for the decoder to decode chroma or luma samples; the first NAL unit being preceded by a required container NAL unit of a third NAL unit type; the first NAL unit being preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit being encapsulated by a required container NAL unit; and the first NAL unit being preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit.
[0009] According to one aspect of the present disclosure, a method for video encoding is performed by at least one processor of an encoder, the method comprising: generating a Network Abstraction Layer unit (NAL unit) stream, the NAL unit stream including at least one first NAL unit of a first type; and encoding one or more encoded pictures according to the first NAL unit; wherein the NAL unit stream indicates that the first NAL unit cannot be discarded by at least one of the following: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicates a second NAL unit type, the second NAL unit type being an NAL unit in the NAL unit stream that is not required for processing chrominance or luminance samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; the first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit is encapsulated by a required container NAL unit; and the first NAL unit is preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit.
[0010] According to one aspect of the present disclosure, a method performed by at least one processor includes: receiving an NAL unit stream, the NAL unit stream including at least one first NAL unit of a first type; wherein the decoder is notified that the first NAL unit cannot be discarded by the decoder by at least one of the following: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicates a second NAL unit type, the second NAL unit type being an NAL unit in the NAL unit stream that is not required for the decoder to decode chrominance or luminance samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; the first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit is encapsulated by a required container NAL unit; and the first NAL unit is preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0012] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0013] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.
[0014] Figure 3 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.
[0015] Figure 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.
[0016] Figure 5 is a schematic diagram of a NAL unit and an SEI header according to an embodiment.
[0017] Figure 6 is a schematic diagram of the required presence of signaling an SEI message via profile_id according to an embodiment.
[0018] Figure 7 is a schematic diagram of the required presence of signaling an SEI message or other NAL units via a required container NAL unit according to an embodiment.
[0019] Figure 8 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION
[0020] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0021] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation to the exact forms disclosed. Modifications and variations are possible in light of the foregoing disclosure, or may be obtained from practice of the embodiments. Additionally, one or more features or components of one embodiment may be incorporated into or combined with those of another embodiment (or one or more features of another embodiment). Further, in the flowcharts and operation descriptions provided below, it will be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be switched.
[0022] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit these implementations. Accordingly, the operations and behavior of the systems and / or methods are described herein without reference to specific software code. It should be understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.
[0023] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure that may be achieved. In fact, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may depend directly on only one claim, the disclosure that may be achieved includes combinations of each dependent claim with every other claim in the claim set.
[0024] Unless explicitly stated, no element, act, or instruction used herein should be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be interchanged with "one or more." If only one item is intended, the term "one" or similar language is used. Additionally, as used herein, the terms "having," "including," etc. are intended to be open-ended terms. Additionally, unless otherwise explicitly specified, the term "based on" is intended to mean "at least partially based on." Additionally, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0025] Throughout the specification, references to "one embodiment," "an embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the present solution. Thus, the phrases "in one embodiment," "in an embodiment," and similar language throughout the specification may, but do not necessarily, all refer to the same embodiment.
[0026] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. Based on the description herein, those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages that may be recognized in certain embodiments may not be present in all embodiments of the present disclosure.
[0027] SEI messages do not need to be processed and decoded by a luma / chroma sample decoder and are thus sometimes considered "optional" in system standards. Therefore, some system standards recommend removing SEI messages in scenarios such as network congestion. However, recent advances in video decoding in certain areas, including machine-oriented video coding and neural network-based guided post-filters, enable the system to forward relevant SEI messages even in the presence of network congestion. Therefore, a mechanism is needed to inform network-based middleboxes that certain SEI messages are required from the perspective of the receiving system, even though they are not required from the perspective of luma / chroma sample decoding.
[0028] Figure 1 FIG. Figure 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110-120) interconnected via a network (150). For unidirectional transmission of data, a first terminal (110) may encode video data at a local location for transmission via the network (150) to another terminal (120). The second terminal (120) may receive the encoded video data of the other terminal from the network (150), decode the encoded data, and display the recovered video data. Unidirectional data transmission is relatively common in applications such as media services.
[0029] Figure 1 FIG.
[0029] shows a second pair of terminals (130, 140) provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional transmission of data, each terminal (130, 140) may encode video data captured at a local location for transmission via the network (150) to the other terminal. Each terminal (130, 140) may also receive the encoded video data sent by the other terminal, decode the encoded data, and display the recovered video data on a local display device.
[0030] In Figure 1 FIG. Figure 1 and FIG.
[0029] , the terminals (110-140) may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (150) represents any number of networks that convey encoded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present disclosure, unless otherwise explained hereinafter, the architecture and topology of the network (150) may be immaterial to the operations disclosed herein. The network (150) may include a Media Aware Network Element (MANE) (160), which may be included, for example, in the transmission path between terminal (130) and terminal (140). The purpose of the MANE may be to selectively forward portions of media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users. Such a MANE may be able to parse and react to a limited portion of the media being conveyed through the network, such as syntax elements related to video coding techniques or the network abstraction layer of a standard.
[0031] As an example of an application of the disclosed subject matter, Figure 2 illustrates the placement of a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0032] A streaming system may include a capture subsystem (213), which may include a video source (201). The video source is, for example, a digital camera, which is used to create an uncompressed video sample stream (202). This sample stream (202) is depicted as a thick line to emphasize the high data volume as compared to an encoded video bitstream, and may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video bitstream (204) is depicted as a thin line to emphasize the lower data volume as compared to the sample stream (202), and may be stored on a streaming server (205) for future use. One or more streaming clients (206, 208) may access the streaming server (205) to retrieve a copy (207, 209) of the encoded video bitstream (204). The client (206) may include a video decoder (210), which decodes an incoming copy of the encoded video bitstream (207) and creates an outgoing video sample stream (211) that can be presented on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendations H.265 and H.266. This application may be used in the context of the VVC standard.
[0033] Figure 3 may be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0034] A receiver (310) may receive one or more coded video sequences to be decoded by a decoder (210); in the same or another embodiment, one coded video sequence is received at a time, where the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device storing the coded video data. The receiver (310) may receive the coded video data as well as other data, e.g., coded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver (310) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (310) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). While it may not be necessary to configure the buffer (315), or the buffer may be made smaller, when the receiver (310) receives data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network. Of course, a buffer (315) may be needed for use on a service packet network such as the Internet, and the buffer may be relatively large and may advantageously have an adaptive size.
[0035] A video decoder (210) may include a parser (320) to reconstruct symbols (321) from an entropy-coded video sequence. The classes of these symbols include information for managing the operation of the decoder (210), and potential information for controlling a display device such as a display screen (212), which is not part of the decoder but may be coupled to the decoder, such as Figure 2As shown. The control information for the display device may be a parameter set segment (not labeled) of Supplementary Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (320) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be carried out according to video coding techniques or standards and can follow principles well-known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (320) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Unit (CU), blocks, Transform Unit (TU), Prediction Unit (PU), and so on. The entropy decoder / parser can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0036] The parser (320) can perform entropy decoding / parsing operations on the video sequence received from the buffer (315) to create symbols (321).
[0037] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (321) may involve multiple different units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser (320) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (320) and multiple units below are not described.
[0038] In addition to the functional blocks already mentioned, the decoder (210) can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0039] The first unit is a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantized transform coefficients as symbols (321) and control information from the parser (320), including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit can output a block including sample values, and the sample values can be input into an aggregator (355).
[0040] In some cases, the output samples of the scaler / inverse transform unit (351) can belong to intra-coded blocks; that is: blocks that do not use predictive information from previously reconstructed pictures, but can use predictive information from the previously reconstructed part of the current picture. Such predictive information can be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates a block with the same size and shape as the block being reconstructed using the surrounding reconstructed information extracted from the current (partially reconstructed) picture (356). In some cases, the aggregator (355) adds the predictive information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) based on each sample.
[0041] In other cases, the output samples of the scaler / inverse transform unit (351) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (353) can access the reference picture memory (357) to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbols (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation unit obtaining the prediction samples from the address in the reference picture memory can be controlled by a motion vector, and the motion vector is in the form of the symbol (321) for use by the motion compensation unit, and the symbol includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.
[0042] The output samples of the aggregator (355) can be adopted by various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream, and the parameters can be used for the loop filter unit (356) as symbols (321) from the parser (320). However, video compression techniques can also respond to meta-information obtained during decoding of the previously (in decoding order) part of the encoded picture or encoded video sequence, and respond to previously reconstructed and loop-filtered sample values.
[0043] The output of the loop filter unit (356) can be a sample stream, which can be output to the display device (212) and stored in the reference picture memory (356) for subsequent inter-picture prediction.
[0044] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by a parser (320)), the current reference picture (356) can become part of the reference picture buffer (357), and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0045] The video decoder 320 can perform decoding operations according to a predetermined video compression technique described, for example, in the ITU-T H.266 standard. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard (such as the syntax specified in the video compression technique document or standard and specifically in the profile therein). For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0046] In an embodiment, the receiver (310) can receive additional (redundant) data together with the encoded video. This additional data can be part of the encoded video sequence. The additional data can be used by the video decoder (320) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0047] Figure 4 It can be a functional block diagram of a video encoder (203) according to an embodiment of the present disclosure.
[0048] The encoder (203) can receive video samples from a video source (201) (not part of the encoder), which can capture video images to be encoded by the encoder.
[0049] A video source (201) may provide a source video sequence in the form of a stream of digital video samples to be encoded by an encoder (203). The stream of digital video samples may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCb, RGB, etc.) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (201) may be a storage device storing previously prepared video. In a video conferencing system, the video source (203) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial arrays of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can readily understand the relationship between pixels and samples. The following focuses on describing samples.
[0050] According to an embodiment, the encoder (203) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Enforcing an appropriate encoding speed is a function of the controller (450). The controller controls other functional units as described below and is functionally coupled to these units. For simplicity, the couplings are not labeled in the figure. Parameters set by the controller may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller (450) that may be involved in optimizing the video encoder (203) for a particular system design.
[0051] Some video encoders operate in a "coding loop" that is readily recognizable to those skilled in the art. As a simple description, the coding loop can consist of an encoding portion of an encoder (430) (hereinafter referred to as the "source encoder") that is responsible for creating symbols based on an input picture to be encoded and reference pictures, and a (local) decoder (433) embedded in the encoder (203). The decoder reconstructs the symbols in the same way as a (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input into a reference picture memory (434). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content of the reference picture buffer is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is well known to those skilled in the art.
[0052] The operation of the "local" decoder (433) can be the same as, for example, the "remote" decoder (210) described in detail above in connection with Figure 3 However, briefly referring additionally to Figure 3 , when the symbols are available and the entropy encoder (445) and the parser (320) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding portion of the decoder (210), including the channel (312), the receiver (310), the buffer (315), and the parser (320), may not be fully implementable in the local decoder (433).
[0053] It can be observed at this time that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, this application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.
[0054] As part of its operation, the source encoder (430) can perform motion compensation predictive coding. Referencing one or more previously encoded frames in the video sequence designated as "reference frames", this motion compensation predictive coding performs predictive coding on the input frame. In this way, the coding engine (432) encodes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame.
[0055] The local video decoder (433) can decode the encoded video data of the frame that can be specified as a reference frame based on the symbols created by the source encoder (430). The operation of the encoding engine (432) can be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (433) repeats the decoding process that can be performed by the video decoder on the reference frame and can store the reconstructed reference frame in the reference picture cache (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame to be obtained by the remote video decoder (without transmission errors).
[0056] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as a reference picture motion vector, block shape, etc. The predictor (435) can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (435), it can be determined that the input picture can have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (434).
[0057] The controller (450) can manage the encoding operations of the video encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0058] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The entropy encoder performs lossless compression on the symbols generated by various functional units according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc., so as to convert the symbols into an encoded video sequence.
[0059] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission through the communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0060] The controller (450) may manage the operation of the encoder (203). During encoding, the controller (450) may assign a certain type of encoded picture to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may generally be assigned to any of the following frame types:
[0061] Intra pictures (I pictures), which may be pictures that can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.
[0062] Predictive pictures (P pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0063] Bi - predictive pictures (B pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0064] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block - by - block. These blocks can be prediction - encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non - prediction - encoded, or the block can be prediction - encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non - prediction - encoded with reference to a previously encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture can be non - prediction - encoded with reference to one or two previously encoded reference pictures either through spatial prediction or through temporal prediction.
[0065] The video encoder (203) may perform encoding operations according to a predetermined video encoding technique or a standard such as ITU - T H.266. In operation, the video encoder (203) may perform various compression operations, including prediction - encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0066] In an embodiment, the transmitter (440) may transmit additional data when transmitting the encoded video. The video encoder (430) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set fragments, etc.
[0067] The compressed video may be enhanced in the video bitstream by auxiliary enhancement information, for example, in the form of SEI messages or VUI. The video coding standard may include a specification section for SEI and VUI. The SEI and VUI information may also be specified in a separate specification that can be referenced by the video coding specification.
[0068] Reference Figure 5 , shows an exemplary layout of a Coded Video Sequence (CVS) according to H.266. The encoded video sequence is subdivided into Network Abstraction Layer units (NAL units). An exemplary NAL unit (501) may include a NAL unit header (502), which in turn includes the following 16 bits: The forbidden_zero_bit (503) and nuh_reserved_zero_bit (504) may not be used by H.266 and may be zero in the NAL unit. Conform to the H.266 standard. Three bits of nuh_layer_id (505) may indicate the layer (spatial, SNR, or multi-view enhancement) to which the NAL unit belongs. Five bits of nuh_nal_unit_type define the type of the NAL unit. In H.266 (04 / 2022), 22 NAL unit type values are defined for the NAL unit types defined in H.266, 6 NAL unit types are reserved, and 4 NAL unit type values are unspecified and may be used by specifications other than H.266. Finally, three bits nuh_temporal_id_plus1 (506) of the NAL unit header indicate the temporal layer to which the NAL unit belongs.
[0069] The encoded picture may contain one or more Video Coding Layer (VCL) NAL units and zero or more non-VCL NAL units. The VCL NAL units may contain the encoded data that conceptually belongs to the video coding layer as described above. The non-VCL NAL units may contain data that conceptually does not belong to the video coding layer. Taking H.266 as an example, they can be classified as:
[0070] (1) Parameter set, which includes information that may be necessary for the decoding process and can be applied to more than one encoded picture. Parameter sets and conceptually similar NAL units can be NAL unit types such as DCI_NUT (Decoding Capability Information (DCI)), VPS_NUT (Video Parameter Set (VPS), which, among other things, establishes inter-layer relationships, etc.), SPS_NUT (Sequence Parameter Set (SPS), which, among other things, establishes parameters used and kept constant throughout the CVS, etc.), PPS_NUT (Picture Parameter Set (PPS), which, among other things, establishes parameters used and kept constant within the encoded picture, etc.), and PREFIX_APS_NUT and SUFFIX_APS_NUT (Prefix and Suffix Adaptation Parameter Sets). Parameter sets can include information required by the decoder to decode VCL NAL units and are thus referred to here as "specification" NAL units.
[0071] (2) Picture header (PH_NUT), which is also a "specification" NAL unit.
[0072] (3) NAL units that mark certain positions in the NAL unit stream. These NAL units include NAL units with NAL unit types AUD_NUT (Access Unit Delimiter), EOS_NUT (End of Sequence), and EOB_NUT (End of Bitstream). These NAL units are non-normative and are also called informative in the sense that a compliant decoder does not require these NAL units in its decoding process, although it needs to be able to receive these NAL units in the NAL unit stream.
[0073] (4) Prefix and suffix SEI NAL unit types (PREFIX_SEI_NUT and SUFFIX_SEI_NUT), which indicate NAL units containing prefix and suffix auxiliary enhancement information. In H.266 (04 / 2022), those NAL units are informative because they are not required for the decoding process.
[0074] (5) The filler data NAL unit type FD_NUT indicates filler data; the data can be random and can be used to "waste" bits in the NAL unit stream or bitstream, which may be necessary for transmission over certain synchronous transmission environments.
[0075] (6) Reserved and unspecified NAL unit types.
[0076] Still referring to Figure 5, which shows the layout of the NAL unit stream in decoding order (510), which contains encoded pictures (511), and the encoded pictures (511) contain some of the types of NAL units introduced previously. Somewhere early in the NAL unit stream, the DCI (512), VPS (513), and SPS (514) can jointly establish the parameters that the decoder can use to decode the encoded pictures (including the encoded pictures (511) of the NAL unit stream) of the CVS.
[0077] In the order depicted or any other order that conforms to the video coding technology or standard being used (here: H.266), the encoded pictures (511) can contain: Prefix APS (Prefix APS, 516), Picture header (PH, 517), prefix SEI (prefix SEI, 518), one or more VCL NAL units (519), and suffix SEI (suffix SEI, 520).
[0078] The prefix and suffix SEI NAL units (518 and 520) were motivated during standard development because for some SEI messages, the content of the message is known before the start of encoding of a given picture, while other content is only known after the picture has been encoded. Allowing certain SEI messages to appear earlier or later in the NAL unit stream of the encoded picture through the prefix and suffix SEIs can avoid buffering. In one or more examples, in the encoder, the sampling time of the picture to be encoded is known before the picture is encoded, so the picture timing SEI message can be a prefix SEI message (516). On the other hand, the decoded picture hash SEI message (which contains the hash of the sample values of the decoded picture and can be used, for example, to debug the encoder implementation) is a suffix SEI message (518) because the encoder cannot compute the hash of the reconstructed samples before the picture has been encoded. The positions of the prefix and suffix SEI NAL units are not limited to their positions in the NAL unit stream. The terms "prefix" and "suffix" can imply which encoded pictures or NAL units the prefix / suffix SEI messages can belong to, and details of such applicability can be specified, for example, in the semantic description of a given SEI message.
[0079] Still referring to Figure 5, which shows a simplified syntax diagram of a NAL unit containing a prefix SEI message 518 or a suffix SEI message 520. This syntax can be a container format for multiple SEI messages that can be carried in one NAL unit. For clarity, the details of the emulation prevention syntax specified in H.266 are omitted here. As with other NAL units, the SEI NAL unit starts with a NAL unit header (521). After the header is one or more SEI messages; two SEI messages (530, 531) are depicted and described below. Each SEI message within the SEI NAL unit includes an 8-bit payload type byte (payload_type_byte) (522) that specifies one of 256 different SEI types; an 8-bit payload size byte (payload_size_byte) (523) that specifies the number of bytes of the SEI payload, and a byte payload (524) of the number of payload_size_byte. This structure can be repeated until a payload_type_byte equal to 0xff is observed, which indicates the end of the NAL unit. The syntax of the payload (524) depends on the SEI message, which can be any length between 0 and 255 bytes.
[0080] Some current and previous video compression standards, including for example H.266, characterize SEI messages as not being required for the decoding process of luma and chroma samples. Some system standards (especially those that specify the transmission of video over bandwidth-constrained links) include language stating that when network congestion requires the network or an intermediate box therein to discard data from a NAL unit stream, it can be advantageous to discard SEI messages before discarding NAL units required for the luma / chroma sample decoding process.
[0081] Certain system environments may require the decoder to receive certain SEI messages, and system standards can therefore specify that such SEI messages must be created and sent by the encoder or sending system (at specific intervals or under specific conditions that may depend on the nature of the SEI message), must be transported by the network, and must be received, interpreted, and processed by the receiving system. As an example, this may be true for picture timing SEI messages, recovery point SEI messages, and user data T.35 SEI messages when using H.265 in the DVP TS101 154 specification. If the sending and receiving systems are not operating in their native environments specified in the system standard, but rather in a network environment where intermediate boxes are configured to discard SEI messages once congestion occurs, unexpected and potentially inconvenient or even fatal reactions may occur at the receiver or decoder because such systems may rely on the content of the discarded SEI messages. As network convergence progresses, such scenarios may become increasingly important.
[0082] In MPEG, certain enhancements to encoder and decoder technologies are being developed, for example, those related to Video Coding for Machines (VCM) or the adoption of certain guided post-filters, such as Neural Network based Post Filter (NNPF). Such systems may adopt certain SEI messages under the assumption that the SEI messages are delivered to the receiving system. For example, the NNVCSEI message as present in H.274 and its associated activation SEI message guide the neural network based post-filter to optimize the decoder output. If those SEI messages are not received in the NAL unit stream, the result may be sub-optimal post-filtered video due to the unavailability of the post-filter's guidance information. This is an example where the dropping of SEI messages may have an undesirable but non-fatal impact on the overall system environment. Also considered in MPEG in the context of VCM is SEI message control of neural network based intra coding, which can replace the intra codec built into the H.266 encoder. If such SEI messages are dropped by the network, the result will be fatal, as there will be no decoded intra pictures and even if reconstruction is possible (which may not be achievable), the remaining reconstructed bitstream may be unavailable.
[0083] Many audiovisual service architectures rely on media-aware network elements that may drop SEI NAL units under certain conditions, as by definition, the luminance and chrominance decoding processes do not require SEI NAL units. However, some applications also require some of these SEI messages to ensure consistency in service quality and experience across users. Embodiments of the present disclosure provide a solution to signal the SEI messages that an application needs to maintain in the bitstream. Additionally, a mechanism is proposed that can also mark NAL units of types other than SEI as required.
[0084] Common service architectures for video delivery include network elements responsible for content filtering and adaptation. This can be illustrated by two examples.
[0085] First, in session services, a multi-point immersive teleconferencing system, as defined in 3GPP TS26.114 MTSI, relies on a Media Resource Function (MRF) as a central media control unit (MCU).
[0086] In a large conference topology supported by the Multi-Stream Multimedia Telephony Service for IMS (MSMTSI) defined in Appendix S of the MTSI standard, support for multiple receivers with different capabilities is provided by a media processing entity called the Media Resource Function (MRF) in the media path. In the case of an immersive conference scenario, the MRF can provide viewport-related 360-degree video processing and transcoding capabilities to multiple clients.
[0087] Another example relates to multimedia streams, where the video signal passes through several network entities before being received by the end user. In an audiovisual distribution scenario, the network elements include video content processing functions, storage, and adaptation in a cloud CDN.
[0088] In both of the above cases, some MANEs are responsible for interpreting the video signal. The purpose of the MANE can be to selectively forward portions of the media data in response to network congestion, media switching, media mixing, archiving, and similar tasks typically performed by service providers rather than end users.
[0089] In the case of bandwidth constraints or processing limitations, some NAL units identified as not relevant to video decoding can be discarded. In fact, some system standards, especially those specifying the transmission of video over bandwidth-constrained links (e.g., the RTP payload format for H.264), include language suggesting that the network needs to discard data from the NAL unit stream when faced with network congestion. By definition, SEI messages that are not required for the decoding of luminance or chrominance sample data are typically identified as the main messages to be discarded when faced with network congestion or computational capacity limitations.
[0090] However, recent advances in certain areas of video decoding, including machine-oriented video coding and neural network-based guided postfilters, enable the system to forward relevant SEI messages even in the presence of the above limitations.
[0091] If the sending and receiving systems operate not in their native environments specified in the system standard but in a network environment where the MANE is configured to discard SEI messages once congestion occurs, the DVB receiver will identify interoperability issues that can potentially lead to fatal errors or denial of service. Such scenarios may become increasingly important with the progress of network convergence.
[0092] Also being developed in MPEG are certain enhancements to encoder and decoder technologies, such as those related to VCM or employing guided postfilters such as NNPF. Such systems can adopt certain SEI messages under the assumption that the SEI messages are delivered to the receiving system.
[0093] For example, NNVC SEI messages as present in H.274 and their associated activation SEI messages direct a neural-network based post-filter to optimize the decoder output. If those SEI messages are not received in the NAL unit stream, the result may not be as good as the optimally post-filtered video because the guidance information for the post-filter is not available. Therefore, a mechanism is needed to inform network-based elements that certain SEI messages are needed from the perspective of the receiving system, even though those SEI messages are not needed from the perspective of decoding the luma / chroma samples.
[0094] Over the years, a general trend observable in the standardization of media transmission in places such as the IETF has been to implement end-to-end encrypted content. However, end-to-end encrypted content makes media-aware processing of the content difficult because media-aware network elements (MANEs) cannot "see" the encrypted syntax elements. For RTP-based transmission, this trend seems to expose certain (a small amount of) information outside the security context that MANEs can "see" and thus act upon. For ITU / MPEG-based video technologies, this part is typically the syntax elements present in the NAL unit header: layer and sublayer information and the NAL unit type. Anything deeper in the video syntax is invisible to the MANE because it is encrypted. Therefore, SEI message type information may not be a suitable solution. A secondary, potentially overcome implementation issue is that MANEs process thousands of streams in parallel on general-purpose processors and thus have more stringent computational complexity limitations than the decoder.
[0095] Therefore, what is needed is a mechanism that allows signaling to network elements and decoders certain SEI messages or certain types of NAL units that are not normally required for the decoding process, which cannot be discarded and must be decoded and interpreted, and which should be forwarded to elements in the receiver in response to such content (e.g., neural-network post-filter configuration information) if the SEI message or NAL unit specification indicates so. Embodiments include defining signaling at the same level as the NAL unit type to ensure the interpretation of scrambled content.
[0096] In a first embodiment, the profile specification of a video coding standard makes certain SEI messages mandatory to process. Referring Figure 6 to, the profile can be a specification that canonically describes the values of certain syntax elements (e.g., as documented in Annex A of H.266), thus allowing or prohibiting certain tools. Figure 6A NAL unit stream is shown. The profile indication (701) can be carried, for example, in a parameter set (e.g., a sequence parameter set (SPS) (702)). In an embodiment, as an example, the profile description can state that a certain first SEI message (703) must (e.g., at certain intervals) be present in the bitstream, or be associated with a key frame or a similar triggering event, while a second SEI message (704) may or may not be present in the bitstream. Note that the NAL unit stream can also contain other NAL units (705), which are not depicted in detail here for clarity. The profile description can also state that the decoder must receive the first SEI message (703) of a first SEI type (706), while the decoder can choose to receive or discard the second SEI message (704) of a second SEI type (707). This means that the decoder must parse all SEI messages (703, 704) at least to the extent that the SEI message can be identified, e.g., by combinatorially interpreting the NAL unit type and the SEI type (706, 707) of the NAL unit containing the SEI message, both of which will be described further below.
[0097] While such additions in video codec specifications and video encoders and decoders can be relatively easily specified and implemented, in system layer specifications and processing units such as middleboxes, such additions can lead to high complexity. Refer to Figure 1, consider a scenario where a terminal (130) communicates with a terminal (140) via a network (150) involving a Selective Forwarding Unit (SFU) middlebox (160). For example, the SFU (160) can be configured to discard certain NAL units if congestion is sensed on its outgoing link. To this end, the SFU (160) can easily interpret certain fields in the NAL unit header. However, due to complexity reasons, it may be cost-prohibitive or undesirable for the SFU to interpret relatively complex video syntax structures such as parameter sets at a given time or maintain the state of the parameter sets in use. However, in order for the SFU to decide whether it needs to forward a received SEI message (as identified by its NAL unit type, which can be easily parsed from the NAL unit header), the SFU will need to know which profile is in use at the time of receiving the SEI NAL unit (position in the NAL unit stream), then interpret the NAL unit type to identify the SEI message, then interpret the SEI payload, then remove the discardable SEI messages while leaving the non-discardable SEI messages in the NAL unit, then rewrite the NAL unit (including steps such as emulation prevention), and then forward the NAL unit. The SEI payload can include any number of SEI messages, each with an 8-bit SEI type indication and information allowing to "skip" the SEI payload associated with that message to identify the next SEI payload. To this extent, according to the prior art, an SFU that does not need to interpret data outside the NAL unit header will now need to interpret the content of the SEI message; the message may be rewritten after removing unnecessary data but retaining necessary data, and interpret, maintain state, and at least associate the profile information from the parameter set with the SEI message. All of these greatly increase the complexity of the SFU, and due to the increased complexity of SFU processing, at least a software upgrade may be required, but an overall replacement of the SFU device may also be needed. From the perspective of a network provider, any solution that does not depend on the profile is preferred.
[0098] In a second embodiment, at least one (but advantageously two) previously unassigned NAL unit types can be allocated to indicate SEI messages that may be required for systems involving NNVC or VCM technologies such as those listed above. If the current available difference between the prefix and suffix NAL unit types is not necessary for the required SEI messages, a single NAL unit type may be sufficient. If it is necessary to preserve the difference between the prefix and suffix SEI messages, two NAL unit types may need to be allocated.
[0099] As an example of a practical implementation, the allocation of NAL unit types for non-VCL NAL units according to H.266 is reproduced below.
[0100] To implement the above mechanism, the currently reserved NAL unit types 26 and 27 can be allocated as follows:
[0101] With such a design, no major changes are required to the SFU, middleboxes, and MANE; of course, this does not apply to middleboxes that conform to the system layer standard, which operate under the assumption of forwarding anything they do not understand, and theoretically doing so creates a more robust network. Middleboxes and MANE that are configured to discard previously undefined or unspecified NAL unit types may need to be upgraded, but since the changes are minimal, it is conceivable that a software upgrade should be sufficient in many cases. The main drawback of this approach is the limited number of available NAL unit types. Specifically, the design change according to this embodiment will fill at least one (and possibly both) of the remaining two reserved NAL unit types.
[0102] Brief reference Figure 5 , a variant of the above design would be to allocate the currently reserved nuh_reserved_zero_bit(504) to indicate that signaling of the NAL unit may be "required" for NAL unit types 26 or 27 (SEI messages) and / or for other VCL or non-VCL NAL units. This design alternative has minimal overhead and may have no drawbacks other than that once the bit is allocated, it cannot be used for any other purpose. Given the limited code point space remaining in the NAL unit header, this allocation may be too costly from the perspective of standard writing.
[0103] In a third embodiment, preferably a single NAL unit type (e.g., the previously reserved NAL unit type 26) is allocated to the new required container NAL unit. The required container NAL unit can be used as a container for prefix or suffix SEI messages and other NAL units that do not need to be processed by the decoder, to signal to the decoder and middleboxes that the NAL units carried within the container are required by the system in use. Examples of other NAL units can include, for example, AUD NAL units, EOS NAL units, or EOB NAL units that are known to be relied upon in some systems.
[0104] The NAL unit header, with its fields filled as follows:
[0105] Forbidden_zero_bit is equal to 0;
[0106] nuh_reserved_zero_bit is equal to 0;
[0107] The nuh_layer_id is equal to the minimum layer_id of all NAL units carried within the required container SEI, or is equal to 0 in the same or another embodiment;
[0108] The nal_unit_type is equal to an unassigned NAL unit type, preferably equal to an unassigned reserved non-VCL NAL unit type, e.g., equal to 26 or 27;
[0109] The nuh_temporal_id_plus1 is equal to 0.
[0110] After the NAL unit header, there may be one or more NAL units that the receiving system may require, even if they may not be needed for the decoding process of chrominance or luminance samples. There are many options to construct this syntax, some of which are described below:
[0111] Reference Figure 7 , in the same or another embodiment, the payload of the required container NAL unit may be empty; i.e., the required container NAL unit consists only of the above NAL unit header. The required container NAL unit may be related to the NAL unit that immediately follows the required container NAL unit. In this case, the single NAL unit that immediately follows the required container NAL unit may be interpreted as being required to be processed by the receiving system. To illustrate this, a stream of NAL units in decoding order is shown, which includes any number of leading NAL units (801), a required container NAL unit (802), a NAL unit (803) (i.e., the NAL unit marked as required by the previously existing required container NAL unit (802)), another NAL unit (804) that is not marked as required because there is no required container NAL unit immediately following it, and any number of other NAL units (805).
[0112] This relatively easy-to-implement mechanism has advantages, including not requiring maintaining any state other than processing the NAL unit immediately following the required container NAL unit. This may be beneficial from the perspective of implementation complexity as well as from the perspective of error recovery. In some cases, the number of NAL units that need to be marked as required may be small (e.g., one per picture), because even if multiple SEI messages need to be marked as "required", the above can be used in Figure 5The mechanisms described in the context of include these SEI messages into a single prefix or suffix SEI NAL unit. Marking a single NAL unit as requiring the desired overhead can be the size of the required container NAL unit, which is 16 bits in the H.266 syntax. Another advantage can be that the mechanism is backward compatible, because if a decoder or an intermediate box does not understand the required container NAL unit, it can be discarded and the overall system design is no worse than without the required container NAL unit.
[0113] To reduce the overhead in cases where more than one (consecutive or spaced-apart) NAL unit needs to be marked as required, several options can be considered.
[0114] Referring again to Figure 7 , in the fourth embodiment, the required container NAL unit can be used to mark one or more NAL units as required. After the NAL unit header of the required container NAL unit, a single syntax element (e.g., a single byte) can be used to indicate the number of NAL units or the number of bytes to which the required container NAL unit applies. In the NAL unit stream, there can be any number of preceding NAL units (801), followed by the required container NAL unit (812), and the required container NAL unit includes a length field (813), which, for example, indicates the subsequent NAL units marked as "required" and is set to 2 here. Thus, the next two NAL units (814, 815) are marked as "required". Unless another required container SEI is encountered in the NAL unit stream (not shown), any number of subsequent NAL units (805) are not required. The disadvantages of this mechanism are that the minimum size of the required container NAL unit is 24 bits instead of 16 bits (of the third embodiment), and the state of "required" needs to be maintained for more consecutive NAL units. In addition, although processing this type of required container NAL unit is relatively simple, the encoder / transmitter implementation needs to know the number of subsequent required NAL units before writing the required container NAL unit, which can be difficult in some implementations.
[0115] Still referring to Figure 7, in the fifth embodiment, a required container NAL unit may encapsulate one or more NAL units. The NAL unit stream may consist of any number of preceding NAL units (801), followed by a required container NAL unit (822) consisting of its NAL unit header (823) and length field (824) that have been introduced. The required container NAL unit may encapsulate, for example, an access unit delimiter NAL unit (825) and a prefix SEI NAL unit (826). The length field (824) may be interpreted, for example, in units of subsequent bytes or subsequent NAL units. The advantage of this mechanism is that it can save bitrate when the required container NAL unit carries two or more NAL units. The disadvantage may be that if the encapsulated NAL units are not redundantly included in the NAL unit stream, this mechanism may not be easily implemented in a backward-compatible manner, which may in turn offset the effect of bitrate savings.
[0116] In the sixth embodiment, two new NAL unit types may be used, one for the start of a required container NAL unit (833) and the other for the end of a required container NAL unit (835). Any NAL unit (834) in the NAL unit stream that is located between these start and end tags in decoding order may be marked as "required". This mechanism requires the allocation of two NAL unit types and there are issues with error recovery, but it is easy to implement and efficient if there are many consecutive NAL units that need to be marked as "required".
[0117] In one or more examples, a NAL unit type that was previously unassigned can be assigned to indicate an empty NAL unit (e.g., a NAL unit with a zero-length RBSP), or a NAL unit with minimal control information in its RBSP, in order to indicate that potentially type-independent subsequent NAL unit(s) are "required". In such a case, no mandatory decoder action for a NAL unit that marks the required NAL unit prefix may be specified, which could be a fundamental change to the concept of optional NAL units (including SEI). However, the semantics can indicate that when NAL units marked as "required" are removed from the NAL unit stream by the action of the MANE, it can have an adverse impact on the user experience, even to the extent that the decoded bitstream becomes useless for the receiving system. In one or more examples, certain TVs and set-top boxes in response to HEVC bitstreams encapsulated by the DVB protocol do not act correctly on bitstreams that lack an Access Unit Delimiter NAL unit. This NAL unit is optional in HEVC (as well as AVC and VVC), but its use is mandated by the DVB specification. In a heterogeneous transmission system involving, for example, feeding from web real-time communication (webrtc) to DVB broadcast, if a DVB-compliant encoder includes an AUD, the coupled webrtc-based transmission chain (including the MANE) may discard the AUD for bandwidth or any reason (and can freely do so according to its specification), and if the bitstream transmitted by webrtc is fed into the DVB transmission for delivery to the user, the user equipment may fail because the webrtc device has discarded the AUD.
[0118] In one or more examples, the required NAL unit prefix can be as follows: Syntax:
[0119] In one or more examples, when a REQ_NU NAL unit is present, the NAL unit immediately following it can be considered required for the application. From the perspective of the decoder, the reception of this NAL unit can be a no-op. In other words, while a smart decoder may take cues from the fact that the encoder or sending system placed this NAL unit in the bitstream and acted accordingly, from a standards perspective, the decoder can ignore it.
[0120] In one or more examples, the syntax of the required NAL unit can be as follows:
[0121] In one or more examples, when there is a REQ_NU NAL unit, the req_nu_count_minus1 + 1 NAL units immediately following it can be considered necessary for the application.
[0122] The embodiments provide the following advantages. There is no problem from the perspective of a narrow video coding standard because the nu-req NAL unit type is unspecified and will thus be ignored by the decoder. If a traditional transmission chain does not understand the nu_req NAL unit type, by its design, it will either forward the strange NAL unit or discard it. If the NAL unit is forwarded, the mechanism works well further downstream, although that particular transmission chain will not specifically act on the NAL unit but perform what it always does (including possibly discarding required NAL units).
[0123] If nu_req is correctly interpreted in a modern transmission chain but there is a loss of NAL units, the following sub-cases can be envisioned. First, nu_req is lost. The transmission chain can freely remove the unmarked (due to loss) required NAL units. However, the encoder can freely include multiple redundant copies of nu_req in the bitstream to increase the statistical probability that at least one of those nu_req messages passes through. Second, a required NAL unit after nu_req is lost. In this case, the receiving application will encounter trouble. Moreover, the NAL unit is marked as "required", which is not really "required" in the sense of this design. From the perspective of the receiving application, there is nothing to be done except relying on general error recovery design considerations (e.g., building robustness and, in the worst case, re-syncing to the stream). As for the bitstream syntax, from the perspective of the standard, marking an incorrect NAL unit as required has no effect because the decoder will discard the information anyway.
[0124] In one or more examples, an AU includes one or more PUs arranged in increasing order of nuh_layer_id.
[0125] In one or more examples, there can be at most one AUD NAL unit in an AU. When an AUD NAL unit is present in an AU, it will be the first NAL unit of the AU and, thus, the first NAL unit of the first PU of the AU. When vps_max_layers_minus1 is greater than 0, there shall be and only be one AUD NAL unit in each IRAP or GDR AU.
[0126] In one or more examples, there can be at most one OPTIONAL unit in an AU. When an OPTIONAL unit is present in an AU, it shall be the first NAL unit after the AUD NAL unit (if any), otherwise it shall be the first NAL unit of the AU.
[0127] In one or more examples, there can be at most one EOB NAL unit in an AU.
[0128] In one or more examples, when an EOB NAL unit is present in an AU, it shall be the last NAL unit of the AU and thus the last NAL unit of the last PU of the AU.
[0129] A VCL NAL unit can be the first VCL NAL unit of an AU (and thus the PU containing the VCL NAL unit is the first PU of the AU) when the VCL NAL unit is the first VCL NAL unit of a picture and one or more of the following conditions are true: - The value of nuh_layer_id of the VCL NAL unit is less than or equal to the nuh_layer_id of the previous picture in decoding order. - The value of ph_pic_order_cnt_lsb of the VCL NAL unit is different from the ph_pic_order_cnt_lsb of the previous picture in decoding order. - The PicOrderCntVal derived for the VCL NAL unit is different from the PicOrderCntVal of the previous picture in decoding order.
[0130] In one or more examples, firstVclNalUnitInAu can be the first VCL NAL unit of an AU. The first NAL unit (if any) among the following NAL units before firstVclNalUnitInAu and after the last VCL NAL unit before firstVclNalUnitInAu specifies the start of a new AU: - The AUD NAL unit (when present), - The OPTIONAL unit (when present), - The DCI NAL unit (when present), - The VPS NAL unit (when present), - The SPS NAL unit (when present), - The PPS NAL unit (when present), - The prefix APS NAL unit (when present), –-PH NAL unit (when present), –-prefix SEI NAL unit (when present), –-NAL unit with nal_unit_type equal to RSV_NVCL_27 (when present), –-NAL unit with nal_unit_type in the range UNSPEC28..UNSPEC29 (when present).
[0131] In one or more examples, the first NAL unit (if present) before firstVclNalUnitInAu and after the last VCL NAL unit before firstVclNalUnitInAu is one of these types of NAL units. In one or more examples, the requirement for bitstream conformance is that, when present, the next PU of a particular layer after an EOS NAL unit belonging to the same layer shall be an IRAP or GDR PU.
[0132] In one or more examples, a PU consists of zero or one PH NAL unit, one coded picture including one or more VCL NAL units, and zero or more other non-VCL NAL units.
[0133] In one or more examples, when a picture consists of more than one VCL NAL unit, a PH NAL unit shall be present in the PU.
[0134] In one or more examples, a VCL NAL unit is the first VCL NAL unit of a picture when it has sh_picture_header_in_slice_header_flag equal to 1 or when the VCL NAL unit is the first VCL NAL unit after a PH NAL unit.
[0135] In one or more examples, the order of non-VCL NAL units (other than AUD, OPI, and EOB NAL units) within a PU shall follow the following restrictions: –-When a PH NAL unit is present in the PU, it shall be before the first VCL NAL unit of the PU.–-When any DCI NAL unit, VPS NAL unit, SPS NAL unit, PPS NAL unit, prefix SEI NAL unit, NAL unit with nal_unit_type equal to RSV_NVCL_27, or NAL unit with nal_unit_type in the range UNSPEC_28..UNSPEC_29 is present in the PU, they shall not follow the last VCL NAL unit of the PU. – When any DCI NAL unit, VPS NAL unit, SPS NAL unit, or PPS NAL unit is present in a PU, they shall be before the PH NAL unit (when present) of the PU and shall be before the first VCL NAL unit of the PU. – NAL units in a PU with nal_unit_type equal to SUFFIX_SEI_NUT, FD_NUT, or RSV_NVCL_27, or in the range UNSPEC_30..UNSPEC_31 shall not be before the first VCL NAL unit of the PU. – When any prefix APS NAL unit is present in a PU, they shall be before the first VCL NAL unit of the PU. – When any suffix APS NAL unit is present in a PU, they shall follow the last VCL NAL unit of the PU. – When an EOS NAL unit is present in a PU, it shall be the last NAL unit among all NAL units in the PU except other EOS NAL units (when present) or EOB NAL units (when present). – When a REQ_SEI_NUT NAL unit is present, the following PREFIX_SEI_NUT or SUFFIX_SEI_NUT is considered necessary for the application. The decoder's reaction to the REQ_NU NAL unit is unspecified. Or, – When a REQ_NU NAL unit is present, the following NAL unit is considered necessary for the application. The decoder's reaction to the REQ_NU NAL unit is unspecified. Or, – When a REQ_NU NAL unit is present, the following req_nu_count_minus1 + 1 NAL units are considered necessary for the application. The decoder's reaction to the REQ_NU NAL unit is unspecified.
[0136] In one or more examples, the common SEI RBSP syntax includes:
[0137] In one or more examples, information that is not necessary for decoding samples of an encoded picture from a VCL NAL unit may be included in the common auxiliary enhancement information RBSP. The VSEI RBSP may contain a VSEI message.
[0138] In one or more examples, vsei_importance being equal to 1 indicates that the general SEI message can be important or required. In one or more examples, vsei_importance being equal to 0 indicates that the general SEI message does not have a particular importance. In one or more examples, this information can be used when an entity that knows the flag needs to make a decision on whether to deliver / discard the general SEI message.
[0139] In one or more examples, the general SEI message syntax includes:
[0140] In one or more examples, the general supplementary enhancement information RBSP contains information that is not necessary for decoding the samples of the coded picture from the VCL NAL unit. The VSEI RBSP can contain one VSEI message.
[0141] In one or more examples, each general SEI message consists of a variable that specifies the importance and type payloadType of the SEI message payload. In one or more examples, the NAL unit byte sequence containing the SEI message can include one or more emulation prevention bytes (represented by the emulation_prevention_three_byte syntax element).
[0142] In one or more examples, the vsei_payload_type_byte is the byte of the payload type of the general SEI message.
[0143] Those skilled in the art can design various modifications and combinations of the foregoing techniques.
[0144] The techniques for signaling the required NAL units described above can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 8 A computer system (900) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0145] The computer software can be encoded in any suitable machine code or computer language, creating code including instructions through mechanisms such as assembly, compilation, and linking, which can be directly executed by a computer's central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.
[0146] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0147] Figure 8 The components of the illustrated computer system 900 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependency or requirement on any one component or combination thereof shown in the exemplary embodiments of the computer system (900).
[0148] The computer system (900) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or more human users through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as sound, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0149] The human-machine interface input device may include one or more of the following (only one is shown): keyboard (901), mouse (902), touchpad (903), touch screen (910), data glove (904), joystick (905), microphone (906), scanner (907), camera (908).
[0150] The computer system (900) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen (910), data glove (904), or joystick (905), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including cathode ray tube (CRT) screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touch screen input functionality and may or may not have tactile feedback functionality - some of which can output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0151] The computer system (900) may also include a human-accessible storage device and its associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) (920) or similar media (921), thumb drives (922), removable hard disk drives or solid state drives (923), conventional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated devices based on ROM / Application-Specific Integrated Circuit (ASIC) / Programmable Logic Device (PLD) such as security software protectors (not shown), and so on.
[0152] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0153] The computer system (900) may also include an interface to one or more communication networks. For example, the network may be wireless, wired, optical. The network may also be a LAN, wide area network, metropolitan area network, vehicular network, and industrial network, real-time network, delay-tolerant network, and so on. Examples of networks also include Ethernet, wireless local area network, cellular networks (Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc.) and other LANs, television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular networks and industrial networks (including Controller Area Network Bus (CANBus)), etc. Certain networks typically require an external network interface adapter for connection to certain common data ports or peripheral buses (949) (e.g., Universal Serial Bus (USB) ports of the computer system (900)). Other systems are typically integrated into the core of the computer system (900) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (900) can communicate with other entities. The communication may be one-way, only for receiving (e.g., wireless television), one-way only for sending (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., through a local or wide area digital network to other computer systems). Each of the above networks and network interfaces may use certain protocols and protocol stacks.
[0154] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (940) of a computer system (900).
[0155] The core (940) may include one or more CPUs (941), GPUs (942), dedicated programmable processing units in the form of Field Programmable Gate Areas (FPGAs) (943), hardware accelerators (944) for specific tasks, etc. These devices, as well as read-only memory (ROM) (945), random access memory (946), internal mass storage (such as internal non-user-accessible hard disk drives, solid-state drives, etc.) (947), etc., can be connected through a system bus (948). In some computer systems, the system bus (948) can be accessed in the form of one or more physical plugs for expansion with additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (848) of the core or connected through a peripheral bus (949). The structure of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc.
[0156] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute certain instructions, and these instructions combined can constitute the above-mentioned computer code. This computer code can be stored in the ROM (945) or RAM (946). Transitional data can also be stored in the RAM (946), while permanent data can be stored in, for example, the internal mass storage (947). Fast storage and retrieval of any memory device can be achieved through the use of a cache memory, which can be closely associated with one or more CPUs (941), GPUs (942), mass storage (947), ROM (945), RAM (946), etc.
[0157] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and computer code can be specially designed and constructed for the purposes of this application or be well-known and available to those skilled in the field of computer software.
[0158] By way of example and not limitation, a computer system having an architecture (900), particularly a core (940), can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be the media associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the non-volatile core (940), such as the core internal mass storage (947) or ROM (945). The software implementing the embodiments of the present application can be stored in such a device and executed by the core (940). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (940), particularly the processors therein (including CPU, GPU, FPGA, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in a Random Access Memory (RAM) (946) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in a circuit (e.g., an accelerator (944)), which can operate instead of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to computer-readable media can include circuits (such as ICs) that store the software for execution, circuits that contain the execution logic, or both. The present application encompasses any suitable combination of hardware and software.
[0159] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various equivalent substitutions that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and thus fall within the spirit and scope of the present application.
[0160] The above disclosure also encompasses the embodiments listed below:
[0161] (1)A method for video decoding, performed by at least one processor of a decoder, the method comprising: receiving (i) one or more encoded pictures and (ii) a NAL unit stream, the NAL unit stream including at least one first NAL unit of a first type; interpreting the first NAL unit; and decoding at least one of the one or more encoded pictures according to the interpretation of the first NAL unit; wherein the decoder is notified that the first NAL unit cannot be discarded by the decoder by at least one of: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicates a second NAL unit type, the second NAL unit type being a NAL unit in the NAL unit stream that is not required for the decoder to decode chrominance or luminance samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; the first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit is encapsulated by a required container NAL unit; and the first NAL unit is preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit.
[0162] (2) The method according to feature (1), wherein the profile indicates that the first NAL unit of the first type cannot be discarded by the decoder.
[0163] (3) The method according to feature (1) or (2), wherein the profile indicates that the NAL unit stream includes a first supplementary enhancement information (SEI) message that cannot be discarded and a second SEI message that can be discarded.
[0164] (4) The method according to any one of features (1)-(3), wherein the profile indicates that each supplementary enhancement information (SEI) message in the NAL unit stream that is (i) at a predetermined interval, (ii) associated with a key frame, or (iii) associated with a trigger event cannot be discarded by the decoder.
[0165] (5) The method according to any one of features (1)-(4), wherein the active parameter set is a sequence parameter set.
[0166] (6) The method according to feature (1), wherein the required container NAL unit is empty.
[0167] (7) The method according to feature (1) or (6), wherein the first NAL unit immediately follows the required container NAL unit, and wherein the NAL unit stream includes a second NAL unit immediately following the first NAL unit, and wherein the second NAL unit can be discarded by the decoder.
[0168] (8) The method according to feature (1), wherein the required container NAL unit includes a NAL unit header, and the NAL unit header includes a field indicating the number of subsequent NAL units, and among which, the immediately following NAL unit corresponding to the number of subsequent NAL units cannot be discarded by the decoder.
[0169] (9) The method according to feature (1) or (8), wherein the NAL unit stream includes a plurality of NAL units, and the plurality of NAL units includes the first NAL unit encapsulated by the required container NAL unit, and among which, each NAL unit encapsulated by the required container NAL unit cannot be discarded by the decoder.
[0170] (10) The method according to any one of features 1, wherein the NAL unit stream includes a plurality of NAL units, and the plurality of NAL units includes the first NAL unit between the start of the required container NAL unit and the end of the required container NAL unit, and among which, each NAL unit between the start of the required container NAL unit and the end of the required container NAL unit cannot be discarded by the decoder.
[0171] (11) A method for video encoding, executed by at least one processor of an encoder, the method comprising: generating a network abstraction layer unit (NAL unit) stream, the NAL unit stream including at least one first NAL unit of a first type; and encoding one or more encoded pictures according to the first NAL unit; wherein the NAL unit stream indicates that the first NAL unit cannot be discarded by at least one of the following: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicates a second NAL unit type, and the second NAL unit type is a NAL unit that is not required for processing chrominance or luminance samples in the NAL unit stream; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; the first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit is encapsulated by a required container NAL unit; and the first NAL unit is preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit.
[0172] (12) The method according to feature (11), wherein the profile indicates that the first NAL unit of the first type cannot be discarded.
[0173] (13) The method according to feature (11) or (12), wherein the profile indicates that the NAL unit stream includes a first supplementary enhancement information (SEI) message that cannot be discarded and a second SEI message that can be discarded.
[0174] (14) The method according to any one of features (11)-(13), wherein the profile indicates that each supplementary enhancement information (SEI) message in the NAL unit stream that is (i) located at a predetermined interval, (ii) associated with a key frame, or (iii) associated with a triggering event cannot be discarded.
[0175] (15) The method according to any one of features (11)-(14), wherein the active parameter set is a sequence parameter set.
[0176] (16) The method according to feature (11), wherein the required container NAL unit is empty.
[0177] (17) The method according to feature (11) or (16), wherein the first NAL unit immediately follows the required container NAL unit, and wherein the NAL unit stream includes a second NAL unit immediately following the first NAL unit, and wherein the second NAL unit can be discarded.
[0178] (18) The method according to feature (11), wherein the required container NAL unit includes a NAL unit header that includes a field indicating the number of subsequent NAL units, and wherein the immediately following NAL units corresponding to the number of subsequent NAL units cannot be discarded.
[0179] (19) The method according to feature (11) or (18), wherein the NAL unit stream includes a plurality of NAL units, the plurality of NAL units including the first NAL unit encapsulated by the required container NAL unit, and wherein each NAL unit encapsulated by the required container NAL unit cannot be discarded.
[0180] (20) A method, executed by at least one processor, the method comprising: receiving a Network Abstraction Layer unit (NAL unit) stream, the NAL unit stream including at least one first NAL unit of a first type; wherein, the decoder is notified that the first NAL unit cannot be discarded by at least one of the following: a profile indicated by a value of a profile identifier in an active parameter set; the first type indicates a second NAL unit type, the second NAL unit type being an NAL unit in the NAL unit stream that is not required for the decoder to decode chrominance or luminance samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; the first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; the first NAL unit is encapsulated by a required container NAL unit; and the first NAL unit is preceded by the start of a required container NAL unit and followed by the end of a required container NAL unit.
Claims
1. A method for video decoding, performed by at least one processor of a decoder, the method comprising: receiving (i) one or more coded pictures and (ii) a network abstraction layer (NAL) unit stream, the NAL unit stream comprising at least one first NAL unit of a first type; interpreting the first NAL unit; as well as decoding at least one coded picture of the one or more coded pictures based on the interpretation of the first NAL unit; The decoder is informed that the first NAL unit cannot be discarded by the decoder by at least one of the following: the profile indicated by the value of the profile identifier in the activity parameter set; The first type indicates a second NAL unit type, the second NAL unit type being a NAL unit in the NAL unit stream that is not required for the decoder to decode chroma or luma samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; The first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; The first NAL unit is encapsulated by a required container NAL unit; and The first NAL unit is preceded by a required container NAL unit start and followed by a required container NAL unit end.
2. The method according to claim 1, wherein: The profile indicates that the first NAL unit of the first type cannot be discarded by the decoder.
3. The method according to claim 1, wherein: The profile indicates that the NAL unit stream includes a first supplementary enhancement information (SEI) message that cannot be discarded and a second SEI message that can be discarded.
4. The method according to claim 1, wherein: The profile indicates that each supplementary enhancement information (SEI) message in the NAL unit stream that is (i) located at a predetermined interval, (ii) associated with a key frame, or (iii) associated with a triggering event cannot be discarded by the decoder.
5. The method according to claim 1, wherein: The activity parameter set is a sequence parameter set.
6. The method according to claim 1, wherein: The required container NAL unit is empty.
7. The method according to claim 1, wherein: The first NAL unit immediately follows the required container NAL unit, and wherein the NAL unit stream includes a second NAL unit immediately following the first NAL unit, and wherein the second NAL unit can be discarded by the decoder.
8. The method according to claim 1, wherein: The required container NAL unit includes a NAL unit header, wherein the NAL unit header includes the field indicating the number of subsequent NAL units, The NAL units that follow the number of subsequent NAL units cannot be discarded by the decoder.
9. The method according to claim 1, wherein: The NAL unit stream includes a plurality of NAL units including the first NAL unit encapsulated by the required container NAL unit, wherein each NAL unit encapsulated by the required container NAL unit cannot be discarded by the decoder.
10. The method according to claim 1, wherein: The NAL unit stream includes a plurality of NAL units including the first NAL unit located between the required container NAL unit start and the required container NAL unit end, wherein each NAL unit located between the required container NAL unit start and the required container NAL unit end cannot be discarded by the decoder.
11. A method of video encoding, performed by at least one processor of an encoder, the method comprising: generating a network abstraction layer unit (NAL unit) stream, the NAL unit stream comprising at least one first NAL unit of a first type; as well as encoding one or more coded pictures according to the first NAL unit; The NAL unit stream indicates that the first NAL unit cannot be discarded by at least one of the following: the profile indicated by the value of the profile identifier in the activity parameter set; The first type indicates a second NAL unit type, the second NAL unit type being a NAL unit in the NAL unit stream that is not required for processing chroma or luma samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; The first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; The first NAL unit is encapsulated by a required container NAL unit; and The first NAL unit is preceded by a required container NAL unit start and followed by a required container NAL unit end.
12. The method according to claim 11, wherein: The profile indicates that the first NAL unit of the first type cannot be discarded.
13. The method according to claim 11, wherein: The profile indicates that the NAL unit stream includes a first supplementary enhancement information (SEI) message that cannot be discarded and a second SEI message that can be discarded.
14. The method according to claim 11, wherein: The profile indicates that each supplementary enhancement information (SEI) message in the NAL unit stream that is (i) located at a predetermined interval, (ii) associated with a key frame, or (iii) associated with a triggering event cannot be discarded.
15. The method according to claim 11, wherein: The activity parameter set is a sequence parameter set.
16. The method according to claim 11, wherein: The required container NAL unit is empty.
17. The method according to claim 11, wherein: The first NAL unit immediately follows the required container NAL unit, and wherein the NAL unit stream includes a second NAL unit immediately following the first NAL unit, and wherein the second NAL unit can be discarded.
18. The method according to claim 11, wherein: The required container NAL unit includes a NAL unit header, wherein the NAL unit header includes the field indicating the number of subsequent NAL units, The NAL units that follow the number of subsequent NAL units cannot be discarded.
19. The method according to claim 11, wherein: The NAL unit stream includes a plurality of NAL units including the first NAL unit encapsulated by the required container NAL unit, wherein each NAL unit encapsulated by the required container NAL unit cannot be discarded.
20. A method, performed by at least one processor, comprising: receiving a network abstraction layer unit (NAL unit) stream, the NAL unit stream comprising at least one first NAL unit of a first type; The decoder is notified that the first NAL unit cannot be discarded by the decoder by at least one of the following: the profile indicated by the value of the profile identifier in the activity parameter set; The first type indicates a second NAL unit type, the second NAL unit type being a NAL unit in the NAL unit stream that is not required for the decoder to decode chroma or luma samples; the first NAL unit is preceded by a required container NAL unit of a third NAL unit type; The first NAL unit is preceded by a required container NAL unit including a field indicating the number of subsequent NAL units; The first NAL unit is encapsulated by a required container NAL unit; and The first NAL unit is preceded by a required container NAL unit start and followed by a required container NAL unit end.