Method, apparatus, and program for decoding a video bitstream
By determining the first VCL NAL unit within a PU and AU, the method optimizes video decoding processes, addressing inefficiencies in existing standards and enhancing decoding efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2023-05-25
- Publication Date
- 2026-05-18
AI Technical Summary
Existing video encoding standards like HEVC face challenges in efficiently decoding video bitstreams, particularly in identifying the first VCL NAL unit within a picture unit (PU) and access unit (AU), leading to inefficiencies in video decoding processes.
The method involves determining whether a VCL NAL unit is the first VCL NAL unit of a PU and AU, and decoding the AU based on this determination, using specific flags and conditions to optimize decoding processes.
This approach enhances decoding efficiency by accurately identifying the first VCL NAL unit, thereby improving the overall performance and efficiency of video decoding.
Smart Images

Figure 0007860931000001 
Figure 0007860931000002 
Figure 0007860931000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority under 35 U.S.C. § 119 to U.S. Provisional Application No. 63 / 005,640, filed on April 6, 2020, and U.S. Patent Application No. 17 / 096,168, filed on November 12, 2020, in the United States Patent and Trademark Office, and the disclosures of these are hereby incorporated by reference in their entireties.
[0002] The disclosed subject matter relates to video encoding and decoding, and more particularly, to signaling picture headers in an encoded video stream.
Background Art
[0003] The ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). In 2015, these two standardizing organizations jointly formed JVET (Joint Video Exploration Team) to explore the possibility of developing the next video encoding standard beyond HEVC. In October 2017, they issued a Joint Call for Proposals (CfP) for video compression with capabilities exceeding HEVC. By February 15, 2018, 22 CfP responses had been submitted regarding standard dynamic range (SDR), 12 CfP responses regarding high dynamic range (HDR), and 12 CfP responses regarding 360 video categories. In April 2018, all submitted Call for Papers (CfPs) were evaluated at the 122nd MPEG / 10th JVET meeting. As a result of this meeting, JVET officially began the standardization process for a next-generation video encoding that goes beyond HEVC. This new standard was named Versatile Video Coding (VVC), and JVET was renamed Joint Video Expert Team. [Overview of the Initiative]
[0004] In one embodiment, a method is provided for decoding a video bitstream encoded using at least one processor, comprising the steps of: obtaining a video encoding layer (VCL) network abstraction layer (NAL) unit; determining whether the VCL NAL unit is the first VCL NAL unit of a picture unit (PU) containing the VCL NAL unit; determining, based on the determination that the VCL NAL unit is the first VCL NAL unit of the PU, whether the VCL NAL unit is the first VCL NAL unit of an access unit (AU) containing the PU; and decoding the AU based on the VCL NAL unit, based on the determination that the VCL NAL unit is the first VCL NAL unit of the AU.
[0005] In one embodiment, a device is provided for decoding an encoded video bitstream, comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate as directed by the program code, wherein the program code comprises: a first acquisition code configured to cause at least one processor to acquire a video encoding layer (VCL) network abstraction layer (NAL) unit; a first determination code configured to cause at least one processor to determine whether a VCL NAL unit is the first VCL NAL unit of a picture unit (PU) containing a VCL NAL unit; a second determination code configured to cause at least one processor to determine, based on the determination that the VCL NAL unit is the first VCL NAL of the PU, that the VCL NAL unit is the first VCL NAL unit of an access unit (AU) containing a PU; and a decoding code configured to cause at least one processor to decode an AU based on a VCL NAL unit, based on the determination that the VCL NAL unit is the first VCL NAL unit of the AU.
[0006] In one embodiment, a non-temporary computer-readable medium for storing instructions is provided, the instructions comprising one or more instructions that, when executed by one or more processors of a device for decoding an encoded video bitstream, cause one or more processors to perform the following steps: acquire a video encoding layer (VCL) network abstraction layer (NAL) unit; determine whether the VCL NAL unit is the first VCL NAL unit of a picture unit (PU) containing the VCL NAL unit; determine whether the VCL NAL unit is the first VCL NAL unit of an access unit (AU) containing the PU, based on the determination that the VCL NAL unit is the first VCL NAL unit of the PU; and decode the AU based on the VCL NAL unit, based on the determination that the VCL NAL unit is the first VCL NAL unit of the AU. [Brief explanation of the drawing]
[0007] Further features, properties, and various advantages of the disclosed subject matter will become clearer from the following detailed description and accompanying drawings.
[0008] [Figure 1] This is a schematic diagram of a simplified block diagram of a communication system according to one embodiment.
[0009] [Figure 2] This is a schematic diagram of a simplified block diagram of a communication system according to one embodiment.
[0010] [Figure 3] This is a schematic diagram of a simplified block diagram of a decoder according to one embodiment.
[0011] [Figure 4] This is a schematic diagram of a simplified block diagram of an encoder according to one embodiment.
[0012] [Figure 5] This is a schematic diagram of an example syntax table according to one embodiment.
[0013] [Figure 6A] This is a flowchart of an exemplary process for decoding an encoded video bitstream according to one embodiment. [Figure 6B] This is a flowchart of an exemplary process for decoding an encoded video bitstream according to one embodiment. [Figure 6C] This is a flowchart of an exemplary process for decoding an encoded video bitstream according to one embodiment.
[0014] [Figure 7] This is a schematic diagram of a computer system according to one embodiment. [Modes for carrying out the invention]
[0015] Figure 1 shows a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The system (100) may include at least two terminals (110-120) interconnected via a network (150). For one-way data transmission, a first terminal (110) can encode video data at its local location for transmission to another terminal (120) via the network (150). A second terminal (120) can receive the encoded video data from the other terminal via the network (150), decode the encoded data, and display the restored video data. One-way data transmission is common in media delivery applications and the like.
[0016] Figure 1 shows a pair of second terminals (130, 140) provided to support the bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional data transmission, each terminal (130, 140) can encode video data captured at its local location for transmission to the other terminal over the network (150). Each terminal (130, 140) can also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the restored video data on a local display device.
[0017] In Figure 1, terminals (110-140) may be represented as servers, personal computers, and smartphones, but the principles of this disclosure are not limited to these. Embodiments of this disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (150) represents any number of networks that carry encoded video data between terminals (110-140), including, for example, wired and / or wireless communication networks. Communication network (150) can exchange data over circuit-switched and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network (150) are not important to the operation of this disclosure unless described below.
[0018] Figure 2 shows an example of the arrangement of a video encoder and decoder in a streaming environment for an application of the disclosed subject matter. The disclosed subject matter is similarly applicable to other video-enabling applications, including, for example, video conferencing, digital TV, and the storage of compressed video on digital media such as CDs, DVDs, and memory sticks.
[0019] A streaming system may include, for example, a video source (201) that generates an uncompressed video sample stream (202), and a capture subsystem (213) that may include, for example, a digital camera. This sample stream (202) is shown as a thick line to emphasize the high data volume when compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the camera (201). The encoder (203) can include hardware, software, or a combination thereof and enables or implements aspects of the disclosed subject matter as described in more detail below. The encoded bitstream (204) is shown as a thin line to emphasize the lower data volume when compared to the sample stream and can be stored in a streaming server (205) for future use. One or more streaming clients (206, 208) can access the streaming server (205) to obtain a copy (207, 209) of the encoded video bitstream (204). The client (206) can include a video decoder (210) that decodes an input copy of the encoded video bitstream (207) and generates an output video sample stream (211) that can be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 207, 209) can be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. One under development is a video encoding standard informally known as Versatile Video Coding or VVC. The disclosed subject matter can be used in the context of VVC.
[0020] FIG. 3 may be a functional block diagram of a video decoder (210) according to an embodiment of the present disclosure.
[0021] The receiver (310) can receive one or more encoded video sequences to be decoded by the decoder (210), and in the same or different embodiments, it can receive one encoded video sequence at a time, with the decoding of each encoded video sequence being independent of other encoded video sequences. The encoded video sequences may be received from a channel (312), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (310) may receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which may be transmitted using their respective entities (not shown). The receiver (310) may isolate the encoded video sequences from other data. To combat network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / analyzer (320) (hereinafter "analyzer"). When the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer (315) may be unnecessary or can be made small. For use in best-effort packet networks such as the Internet, the buffer (315) may be required and can be made relatively large and advantageously adaptable in size.
[0022] The video decoder (210) may include an analyzer (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols include information used to manage the operation of the decoder (210) and potential information for controlling rendering devices such as a display (212) that can be coupled to the decoder, as shown in FIG. 3, although not an essential part of the decoder. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a video user utility information (VUI) parameter set fragment (not shown). The analyzer (320) may analyze / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence can follow video encoding techniques or standards and can follow principles well-known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The analyzer (320) may extract a set of at least one subgroup parameter of a subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups include groups of pictures (GOP), pictures, sub-pictures, tiles, slices, bricks, macroblocks, coding tree units (CU), coding units (CU), blocks, transform units (TU), prediction units (PU), etc. A tile may indicate a rectangular region of CUs / CTUs within a specific tile column and row in a picture. A brick may indicate a rectangular region of a CU / CTU column within a specific tile. A slice may indicate one or more bricks of a picture included in a NAL unit. A sub-picture may indicate a rectangular region of one or more slices within a picture. The entropy decoder / analyzer may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence.
[0023] The analyzer (320) may perform an entropy decoding / analysis operation on the video sequence received from the buffer (315) in order to create a symbol (321).
[0024] The reconstruction of the symbol (321) may involve multiple different units, depending on the type of the encoded video picture or its portion (e.g., interframe and intraframe pictures, interframe and intraframe blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information analyzed from the video sequence encoded by the analyzer (320). The flow of such subgroup control information between the analyzer (320) and the following multiple units is not shown for clarity.
[0025] In addition to the functional blocks already described, the decoder 210 can be conceptually divided into several functional units, as will be discussed later. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and integrate with each other at least partially. However, in order to explain the disclosed subject matter, it is appropriate to conceptually subdivide it into the following functional units.
[0026] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantized transform coefficients as symbols (321) from the analyzer (320), along with control information including the transform to be used, block size, quantization factor, and quantization scaling matrix. It can output a block containing sample values that can be input to the aggregator (355).
[0027] In some cases, the output samples of the scaler / inverse transform (351) can be associated with encoded blocks within the frame, i.e., blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed portions of the current picture. Such predictive information can be provided by the in-frame picture predictive unit (352). In some cases, the in-frame picture predictive unit (352) generates blocks of the same size and shape as the block being reconstructed, using already reconstructed surrounding information extracted from the current (partially reconstructed) picture (358). The aggregator (355) may, for each sample, add the predictive information generated by the in-frame predictive unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0028] In other cases, the output samples of the scaler / inverse unit (351) may relate to encoded, potentially motion-compensated blocks between frames. In such cases, the motion-compensated prediction unit (353) can access the reference picture memory (357) to fetch samples to be used for prediction. After motion compensation of the fetched samples according to the symbols (321) relating to the blocks, these samples can be added by the aggregator (355) to the output of the scaler / inverse unit (referred to in this case to residual samples or residual signals) to generate output sample information. The address in the reference picture memory format from which the motion-compensated unit fetches prediction samples can be controlled by motion vectors available to the motion-compensated unit in the form of symbols (321), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when the precise motion vectors of subsamples are in use, motion vector prediction mechanisms, etc.
[0029] The output samples of the aggregator (355) can undergo various loop filtering techniques within the loop filter unit (356). The video compression technique is controlled by parameters contained in the encoded video bitstream and made available to the loop filter unit (356) as symbols (321) from the analyzer (320), but may include in-loop filtering techniques that can respond to metadata obtained during decoding of earlier parts (in decoding order) of the encoded picture or encoded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.
[0030] The output of the loop filter unit (356) can be a sample stream that can be output to the rendering device (212) and stored in a reference picture memory for future inter-frame picture prediction.
[0031] Once the encoded picture is fully reconstructed, it can be used as a reference picture for future predictions. Once the encoded picture is fully reconstructed and identified as a reference picture (e.g., by the analyzer (320)), the current reference picture (358) can be made part of the reference picture buffer (357), and fresh current picture memory can be reallocated before starting the reconstruction of the next picture to be encoded.
[0032] The video decoder 210 may perform decoding operations according to a predetermined video compression technique that may be documented in the ITU-T Rec.H.265 standard. The encoded video sequence may conform to the syntax defined by the video compression technique or standard being used, in the sense that it conforms to the syntax of the video compression technique or standard, as defined in the video compression technique documentation or standard, particularly in the profile documentation therein. Furthermore, compliance may require that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. Limitations set by the level may, in some cases, be further restricted through HRD specifications and metadata for virtual reference decoder (HRD) buffer management signaled in the encoded video sequence.
[0033] In one embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or SNR extension layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0034] Figure 4 may be a functional block diagram of a video encoder (203) according to one embodiment of the present disclosure.
[0035] The encoder (203) may receive video samples from a video source (201) (not part of the encoder) that can capture video images to be encoded by the encoder (203).
[0036] The video source (201) may provide a source video sequence encoded by an encoder (203) in the form of a digital video sample stream, which can be any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (201) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description will focus on samples.
[0037] According to one embodiment, the encoder (203) may encode and compress the pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Enforcing an appropriate encoding rate is one function of the controller (450). The controller controls and is functionally coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will be able to easily identify other functions of the controller (450) as they may relate to a video encoder (203) optimized for a particular system design.
[0038] Some video encoders operate in a manner that can be readily recognized by those skilled in the art as an "encoding loop." In very simple terms, the encoding loop may consist of an encoding portion of an encoder (430) (hereinafter, the "source coder") (responsible for generating symbols based on the input picture to be encoded and a reference picture), and a (local) decoder (433) embedded in the encoder (203) that reconstructs the symbols to create sample data that the (remote) decoder will also create (since any compression between the symbols and the encoded video bitstream is reversible in the video compression techniques considered in the disclosed subject). The reconstructed sample stream is input to the reference picture memory (434). Because decoding of the symbol stream yields bit-accurate results regardless of the decoder location (local or remote), the reference picture buffer content is also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using predictions during decoding. This basic principle of reference picture synchronization (for example, if synchronization cannot be maintained due to channel errors, drift will result) is well known to those skilled in the art.
[0039] The operation of the “local” decoder (433) can be the same as that of the “remote” decoder (210), as has already been described in detail above in conjunction with Figure 3. See also Figure 4, however, since symbols are available and the encoding / decoding of symbols to the encoded video sequence by the entropy coder (445) and analyzer (320) can be reversible, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer (315), and analyzer (320), may not be fully implemented in the local decoder (433).
[0040] At this point, it is understood that any decoder techniques other than parsing / entropy decoding present in the decoder must necessarily exist in the corresponding encoder in substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. A description of encoder techniques can be omitted, as it is the inverse of a comprehensive description of decoder techniques. More detailed explanations are necessary only in specific areas and are provided below.
[0041] As part of this operation, the source coder (430) may perform motion-compensated predictive coding, predictively coding the input frame by referencing one or more previously coded frames from a video sequence designated as “reference frames”. In this way, the coding engine (432) codes the difference between the pixel blocks of the input frame and the pixel blocks of the reference frame that may be selected as predictive references for the input frame.
[0042] The local video decoder (433) may decode the encoded video data of a frame that may be designated as a reference frame based on symbols generated by the source coder (430). The operation of the encoding engine (432) may, advantageously, be a lossy process. When the encoded video data can be decoded by the video decoder, the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder 433 may copy the decoding process that may be performed by the video decoder on the reference frame so that the reconstructed reference frame is stored in the reference picture cache (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has content in common with the reconstructed reference frame (without transmission errors) that will be obtained by the far-end video decoder.
[0043] The predictor (435) can perform a predictive search on the encoding engine (432). That is, for a new frame to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which serve as appropriate predictive references for the new picture. The predictor (435) can operate on a sample block vs. pixel block basis to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (434), as determined by the retrieval results obtained by the predictor (435).
[0044] The controller (450) can manage the encoding operation of the video coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.
[0045] All outputs of the aforementioned functional units can undergo entropy coding in the entropy coder (445). The entropy coder converts the symbols generated by the various functional units into coded video sequences by reversibly compressing the codes according to techniques known to those skilled in the art, such as Huffman coding, variable-length coding, and arithmetic coding.
[0046] The transmitter (440) can buffer the encoded video sequence generated by the entropy coder (445) and prepare it for transmission over the communication channel (460), which can be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) can merge the encoded video data from the video coder (430) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (not shown).
[0047] The controller (450) can manage the operation of the encoder (203). During encoding, the controller (450) can assign a specific encoded picture type to each encoded picture, which may affect the encoding technique that can be applied to each picture. For example, a picture is often assigned as one of the following frame types:
[0048] An in-frame picture (I-picture) may be one that can be encoded and decoded without using any other frames in the sequence as a source for prediction. Some video codecs allow different types of in-frame pictures, including, for example, an independent decoder refresh picture. Those skilled in the art will understand these variations of I-pictures, as well as their respective uses and characteristics.
[0049] A prediction picture (P-picture) may be encoded and decoded using intra-frame or inter-frame prediction, using at most one motion vector and reference index to predict the sample values of each block.
[0050] A bidirectional predictive picture (B-picture) may be encoded and decoded using intra-frame or inter-frame prediction, using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0051] A source picture can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be encoded. Blocks can be predictively encoded by referencing other (already encoded) blocks, as determined by the encoding assignment applied to each picture in the block. For example, blocks of picture I may be encoded unpredictably, or they may be encoded predictively by referencing already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of picture P may be encoded unpredictably, or by referencing a previously encoded reference picture, via spatial prediction or temporal prediction. Blocks of picture B may be encoded unpredictably, or by referencing one or two previously encoded reference pictures, via spatial prediction or temporal prediction.
[0052] The video coder (203) can perform encoding operations in accordance with a specified video encoding technique or standard, such as ITU-T Rec.H.265. In this operation, the video coder (203) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technique or standard being used.
[0053] In one embodiment, the transmitter (440) may transmit additional data along with the encoded video. The video coder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.
[0054] The embodiments may relate to modifications of syntax and semantics related to picture headers. At least one embodiment includes all_pic_coding_ in PPS as a gate flag. info_present_ in _p It may relate to signaling h_flag to save some bits for specifying whether picture-level coding tool information is present in PH or SH. At least one embodiment may relate to correcting the identification of the first VCLNAL unit in PU or AU. At least one embodiment may relate to modifying the meaning of gdr_or_irap_pic_flag to deal with the case where mixed_nalu_types_flag is equal to 1.
[0055] In an embodiment, the picture parameter set (PPS) can refer to a syntactic structure containing syntactic elements that apply to zero or more encoded pictures, as determined by the syntactic elements found within each slice header.
[0056] In the embodiment, the picture header (PH) can refer to a syntactic structure containing syntactic elements that apply to all slices of the encoded picture.
[0057] In an embodiment, the slice header (SH) may refer to a portion of an encoded slice that contains data elements relating to all tiles or encoded tree units within the tile represented in the slice.
[0058] The embodiment may relate to a video coding layer (VCL).
[0059] In an embodiment, a network abstraction layer (NAL) unit may refer to a syntactic structure that includes an indication of the type of subsequent data and bytes containing data in the form of a raw byte sequence payload interspersed with emulation-prevention bytes as needed.
[0060] In embodiments, VCL NAL unit can refer to a collective term for encoded slice NAL units and subsets of NAL units having the reserved value nal_unit_type, which are classified as VCL NAL units in this specification.
[0061] In the embodiment, a picture unit (PU) can refer to a set of NAL units that are related to each other according to a specified classification rule, are consecutive in decoding order, and contain exactly one encoded picture.
[0062] In an embodiment, an access unit (AU) may refer to a set of PUs that belong to different layers and contain encoded pictures related at the same time for output from a decoded picture buffer (DPB).
[0063] The embodiment may relate to sample adaptive offset (SAO).
[0064] In an embodiment, an adaptive loop filter (ALF) may refer to a filtering process that is applied as part of the decoding process and is controlled by parameters carried in an adaptive parameter set (APS).
[0065] The embodiments may relate to quantization parameters (QP).
[0066] The embodiment may relate to an intra-random access point (IRAP).
[0067] The embodiment may relate to gradual decoding refresh (GDR).
[0068] In this embodiment, the GDR picture may point to a picture in which each VCL NAL unit has a nal_unit_type equal to GDR_NUT.
[0069] The latest version of the VVC draft (JVET-Q2001-vE) allows the use of six flags to indicate whether picture-level encoding information is present in the picture header or slice header of the PPS syntax structure. Examples include rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, etc. In most cases, these values can be the same, either 0 or 1. It is unlikely that each xxx_info_in_ph flag will have a different value.
[0070] Therefore, in the embodiment, the gate flag all_pic_coding_info_present_ in4 The ph_flag indicates the presence of these flags in the PPS, which can save bits in the PPS. If the value of all_pic_coding_info_present_in_ph_flag is equal to 1, those xxx_info_in_ph flags are not signaled, and their values may be inferred to be equal to 1. This is because signaling picture-level coding information in the picture header may occur more frequently than signaling information in the slice header for slice-level control. An example of a syntax table consistent with the embodiment is shown in Figure 5.
[0071] In one embodiment, an all_pic_coding_info_present_in_ph_flag equal to 1 can specify that rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, alf_info_in_ph_flag, wp_info_in_ph_flag, and qp_delta_info_in_ph_flag are not present in the PPS. An all_pic_coding_info_present_in_ph_flag equal to 0 can specify that rpl_info_in_ph_flag, dbf_info_in_ph_flag, sao_info_in_ph_flag, alf_info_in_ph_flag, wp_info_in_ph_flag, and qp_delta_info_in_ph_flag are present in the PPS.
[0072] In one embodiment, an rpl_info_in_ph_flag equal to 1 can specify that the reference picture list information is present in the PH syntax structure but not in slice headers referencing PPS that do not contain the PH syntax structure. An rpl_info_in_ph_flag equal to 0 can specify that the reference picture list information is not present in the PH syntax structure but is present in slice headers referencing PPS that do not contain the PH syntax structure. When it is not present, the value of rpl_info_in_ph_flag may be inferred to be equal to 1.
[0073] In an embodiment, a dbf_info_in_ph_flag equal to 1 can specify that the unblock filter information is present in the PH syntax structure but not in slice headers that reference PPS that do not contain the PH syntax structure. A dbf_info_in_ph_flag equal to 0 can specify that the unblock filter information is not present in the PH syntax structure but is present in slice headers that reference PPS that do not contain the PH syntax structure. When it is not present, the value of dbf_info_in_ph_flag may be inferred to be equal to 0. When it is not present, the value of dbf_info_in_ph_flag may be inferred to be equal to 1.
[0074] A sao_info_in_ph_flag value equal to 1 can specify that SAO filter information exists in the PH syntax structure but not in slice headers referencing PPS that do not contain the PH syntax structure. A sao_info_in_ph_flag value equal to 0 can specify that SAO filter information does not exist in the PH syntax structure but is present in slice headers referencing PPS that do not contain the PH syntax structure. When it is not present, the value of sao_info_in_ph_flag may be inferred to be equal to 1.
[0075] A value of 1 for alf_info_in_ph_flag indicates that ALF information is present in the PH syntax structure but not in slice headers referencing PPS that do not contain the PH syntax structure. A value of 0 for alf_info_in_ph_flag indicates that ALF information is not present in the PH syntax structure but is present in slice headers referencing PPS that do not contain the PH syntax structure. When it is not present, the value of alf_info_in_ph_flag may be inferred to be equal to 1.
[0076] A wp_info_in_ph_flag equal to 1 can indicate that weighted prediction information exists in the PH syntax structure but not in slice headers referencing PPS that do not contain the PH syntax structure. A wp_info_in_ph_flag equal to 0 can indicate that weighted prediction information does not exist in the PH syntax structure but is present in slice headers referencing PPS that do not contain the PH syntax structure. When it does not exist, the value of wp_info_in_ph_flag may be inferred to be equal to 0. When it does not exist, the value of wp_info_in_ph_flag may be inferred to be equal to 1.
[0077] A qp_delta_info_in_ph_flag value equal to 1 can specify that QP delta information is present in the PH syntactic structure but not in slice headers referencing PPS that do not contain the PH syntactic structure. A qp_delta_info_in_ph_flag value equal to 0 can specify that QP delta information is not present in the PH syntactic structure but is present in slice headers referencing PPS that do not contain the PH syntactic structure. When it is not present, the value of qp_delta_info_in_ph_flag may be inferred to be equal to 1.
[0078] The latest VVC specification draft does not clearly define how to identify the first VCL NAL unit in the draft, PU, or AU. Embodiments may relate to the following modifications to the description of the order of NAL units.
[0079] In an embodiment, a VCL NAL unit is the first VCL NAL unit in an AU if it is the first VCL NAL unit following a PH NAL unit, or if it has a picture_header_in_slice_header_flag equal to 1 and one or more of the following conditions are true (thus, a PU containing a VCL NAL unit is the first PU in an AU): - The value of the nuh_layer_id of the VCL NAL unit is less than the nuh_layer_id of the previous picture in the decoding order. - The value of ph_pic_order_cnt_lsb for the VCL NAL unit is different from the ph_pic_order_cnt_lsb of the previous picture in the decoding order. - The PicOrderCntVal derived for a VCL NAL unit is different from the PicOrderCntVal of the previous picture in the decoding order.
[0080] In one embodiment, the flag gdr_or_irap_pic_flag in the picture header indicates whether the current picture is an IRAP or GDR picture. When the value of gdr_or_irap_pic_flag is equal to 1, the flag no_output_of_prior_pics_flag may also be present in the picture header. When the bitstreams of subpictures are merged, the value of no_output_of_prior_pics_flag for IRAP subpictures needs to be retained for subpicture extraction. To address this issue, embodiments may relate to the following modifications to the meaning of gdr_or_irap_pic_flag.
[0081] In this embodiment, a gdr_or_irap_pic_flag equal to 1 can specify that the current picture is a GDR or IRAP picture, or a picture having a VCL_NAL unit equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT and a mixed_nalu_types_in_pic_flag equal to 1. A gdr_or_irap_pic_flag equal to 0 can specify that the current picture may or may not be a GDR or IRAP picture.
[0082] In the embodiment, a gdr_or_irap_pic_flag equal to 1 can specify that the current picture is a GDR or IRAP picture, or a picture containing IRAP subpictures that have a mixed_nalu_types_in_pic_flag equal to 1. A gdr_or_irap_pic_flag equal to 0 can specify that the current picture may or may not be a GDR or IRAP picture.
[0083] In this embodiment, a gdr_or_irap_pic_flag equal to 1 can specify that the current picture is a GDR or IRAP picture. A gdr_or_irap_pic_flag equal to 0 can specify that the current picture may or may not be a GDR or IRAP picture.
[0084] It may be a bitstream compatibility requirement that when mixed_nalu_types_in_pic_flag is equal to 1, the value of gdr_or_irap_pic_flag is equal to 0.
[0085] Figures 6A to 6C are flowcharts of exemplary processes 600A, 600B, and 600C for decoding an encoded video bitstream. In some implementations, one or more process blocks in Figures 6A to 6C may be performed by the decoder 210. In some implementations, one or more process blocks in Figures 6A to 6C may be performed by another device or group of devices, either isolated from the decoder 210, such as the encoder 203, or including the decoder 210.
[0086] In this embodiment, one or more of the blocks shown in Figure 6A may correspond to one or more of the blocks in Figures 6B and 6C, or may be executed together with these blocks.
[0087] As shown in Figure 6A, process 600A may include acquiring a video coding layer (VCL) network abstraction layer (NAL) unit (block 611).
[0088] As further shown in Figure 6A, process 600A may include determining that the VCL NAL unit is the first VCL NAL unit of a picture unit (PU) containing the VCL NAL unit (block 612).
[0089] As further shown in Figure 6A, process 600A may include determining that a VCL NAL unit is the first VCL NAL unit of an access unit (AU) containing a PU, based on determining that the VCL NAL unit is the first VCL NAL unit of the PU (block 613).
[0090] As further shown in Figure 6A, process 600A may include decoding the AU based on the VCL NAL units (block 614), based on determining that the VCL NAL unit is the first VCL NAL unit of the AU.
[0091] In this embodiment, one or more of the blocks shown in Figure 6B may correspond to one or more of the blocks in Figures 6A and 6B, or may be executed together with these blocks.
[0092] As shown in Figure 6B, process 600B may include acquiring a VCL NAL unit (block 621).
[0093] As further shown in Figure 6B, process 600B may include determining whether the VCL NAL unit is the first VCL NAL unit following the picture header NAL unit (block 622).
[0094] As further shown in Figure 6B, process 600B may proceed to block 623 based on determining that the VCL NAL unit is the first VCL NAL unit following the picture header NAL unit (YES in block 622).
[0095] As further shown in Figure 6B, process 600B may proceed to block 624 based on determining that the VCL NAL unit is not the first VCL NAL unit following the picture header NAL unit (NO in block 623). In embodiments, process 600B may instead proceed to block 625.
[0096] As further shown in Figure 6B, process 600B may include determining whether a flag in the VCL NAL unit is set to indicate that the slice header contained in the VCL NAL unit contains a picture header (block 624). In an embodiment, the flag may correspond to picture_header_in_slice_header_flag.
[0097] As further shown in Figure 6B, process 600B may proceed to block 623 based on determining that a flag in the VCL NAL unit is set to indicate that the slice header contained in the VCL NAL unit contains a picture header (YES in block 624).
[0098] As further shown in Figure 6B, process 600B may proceed to block 625 based on determining that the flag in the VCL NAL unit is not set to indicate that the slice header contained in the VCL NAL unit contains a picture header (NO in block 624).
[0099] As further shown in Figure 6B, process 600B may include determining that the VCL NAL unit is the first VCL NAL unit in the PU containing the VCL NAL unit (block 623).
[0100] As further shown in Figure 6B, process 600B may include determining that the VCL NAL unit is not the first VCL NAL unit in the PU containing the VCL NAL unit (block 625).
[0101] In this embodiment, one or more of the blocks shown in Figure 6C may correspond to one or more of the blocks in Figures 6A and 6B, or may be executed together with these blocks.
[0102] As shown in Figure 6C, process 600C may include determining that the VCL NAL unit is the first VCL NAL unit of the PU (block 631).
[0103] As further shown in Figure 6C, process 600C may include determining whether the layer identifier of the VCL NAL unit is less than the layer identifier of the previous picture (block 632).
[0104] As further shown in Figure 6C, process 600C may proceed to block 633 based on determining that the layer identifier of the VCL NAL unit is less than the layer identifier of the previous picture (YES in block 632).
[0105] As further shown in Figure 6C, process 600C may proceed to block 634 based on determining that the layer identifier of the VCL NAL unit is greater than or equal to the layer identifier of the previous picture (NO in block 633). In embodiments, process 600C may instead proceed to block 635.
[0106] As further shown in Figure 6C, process 600C may include determining whether the picture order count of a VCL NAL unit is different from the picture order count of a previous picture (block 634). In an embodiment, this may be determined based on the least significant bit (LSB) of the picture order count.
[0107] As further shown in Figure 6C, process 600C may proceed to block 633 based on determining that the picture order count of the VCL NAL unit is different from the picture order count of the previous picture (YES in block 634).
[0108] As further shown in Figure 6C, process 600C may proceed to block 635 based on determining that the picture order count of the VCL NAL unit is no different from the picture order count of the previous picture (NO in block 634).
[0109] As further shown in Figure 6C, process 600C may include determining that the VCL NAL unit is the first VCL NAL unit in the AU containing the VCL NAL unit (block 633).
[0110] As further shown in Figure 6C, process 600C may include determining that the VCL NAL unit is not the first VCL NAL unit in the AU containing the VCL NAL unit (block 635).
[0111] In one embodiment, multiple flags corresponding to picture coding information do not necessarily need to be signaled based on a flag indicating that all picture coding information is present in the picture header. In one embodiment, the flag may correspond to all_pic_coding_info_present_in_ph_flag.
[0112] Figures 6A–6C show exemplary blocks of processes 600A–600C, and in some implementations, processes 600A–600C may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in Figures 6A–6C. Additionally or alternatively, two or more blocks of processes 600A–600C may run in parallel.
[0113] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium in order to perform one or more of the proposed methods.
[0114] The techniques described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 7 shows a computer system 700 suitable for implementing a particular embodiment of the disclosed subject matter.
[0115] Computer software can be encoded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms, and can produce code containing instructions that can be executed directly or through interpretation, microcode execution, etc., by a computer's central processing unit (CPU), graphics processing unit (GPU), etc.
[0116] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, and Internet of Things devices.
[0117] The components shown in Figure 7 for the computer system 700 are essentially illustrative and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or requirement on any one or combination of components shown in the exemplary embodiments of the computer system 700.
[0118] The computer system 700 may include certain human interface input devices. Such human interface input devices can respond to input from one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, applause), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., 2D video, 3D video including stereoscopic images).
[0119] The input human interface device may include one or more of the following: keyboard 701, mouse 702, trackpad 703, touchscreen 710 and associated graphics adapter 750, data glove, joystick 705, microphone 706, scanner 707, and camera 708.
[0120] The computer system 700 may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen 710, data glove, or joystick 705), audio output devices (e.g., speakers 709, headphones (not shown)), visual output devices (e.g., cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light-emitting diode (OLED) screen, each having or not having tactile feedback capability, some of which can output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic image output, such as screen 710, virtual reality glasses (not shown), holographic display, and smoke tank (not shown)), and printers (not shown).
[0121] The computer system 700 may also include human-accessible storage devices and associated media, such as optical media including a CD / DVD ROM / RW 720 having a CD / DVD or similar medium 721, a thumb drive 722, a removable hard drive or solid-state drive 723, legacy magnetic media such as tape and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0122] Those skilled in the art should also understand that the term “computer-readable medium” as used in relation to the subject matter currently disclosed does not include a transmission medium, carrier wave, or other transient signal.
[0123] The computer system 700 may also include interfaces to one or more communication networks (755). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicle and industrial, real-time, latency-tolerant, etc. Examples of networks include cellular networks such as Ethernet®, wireless LAN, global systems for mobile communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), and Long-Term Evolution (LTE), wired and wireless wide-area digital networks including cable television, satellite television, and terrestrial television, and vehicle and industrial networks including CANBus. Certain networks generally require an external network interface adapter (754) attached to a specific general-purpose data port or peripheral bus (749) (for example, the universal serial bus (USB) of the computer system 700), while others are generally integrated into the core of the computer system 700 by attachment to a system bus, as described later (for example, an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). As an example, network 755 may be connected to peripheral bus 749 using network interface 754. Using any of these networks, the computer system 700 can communicate with other entities. Such communication can be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., from CANbus to a specific CAN bus), or bidirectional to other computer systems using local or wide-area digital networks, for example. Specific protocols and protocol stacks may be used with each of those networks and network interfaces (754), as described above.
[0124] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core 740 of the computer system 700.
[0125] The core 740 may include one or more Central Processing Units (CPUs) 741, Graphics Processing Units (GPUs) 742, specialized programmable processing units in the form of Field Programmable Gate Areas (FPGAs) 743, and hardware accelerators 744 for specific tasks. These devices, along with internal mass storage 747 such as read-only memory (ROM), random-access memory (RAM), internal non-user-accessible hard drives, and solid-state drives (SSDs), can be connected via the system bus 748. In some computer systems, the system bus 748 can be made accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the core's system bus 748 or via the peripheral bus 749. The architecture of peripheral buses includes Peripheral Component Interconnect (PCI), USB, and others.
[0126] The CPU 741, GPU 742, FPGA 743, and accelerator 744 can execute specific instructions that, in combination, constitute the aforementioned computer code. This computer code can be stored in ROM 745 or RAM 746. Temporary data can also be stored in RAM 746, while permanent data can be stored, for example, in internal mass storage 747. High-speed storage and retrieval to any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more CPUs 741, GPUs 742, mass storage 747, ROM 745, RAM 746, etc.
[0127] Computer-readable media may have computer code on them for performing various computer-implemented operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and readily available to those skilled in computer software technology.
[0128] As an example, but not limited to, an architecture, specifically a computer system 700 having a core 740, can provide functionality as a result of a processor (including CPUs, GPUs, FPGAs, accelerators, etc.) that runs software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage described above, as well as specific storage of the core 740 of a non-transient nature, such as the core internal mass storage 747 or ROM 745. Software implementing various embodiments of this disclosure can be stored in such devices and executed by the core 740. The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause the core 740, in particular the processor (including CPUs, GPUs, FPGAs, etc.) therein, to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic wired within a circuit (e.g., accelerator 744) or otherwise embodied, which may operate in place of or with software to perform a particular process or a particular part of a particular process as described herein. References to software include logic, and where appropriate, vice versa. References to computer-readable media may include circuits that store software for execution (such as integrated circuits (ICs)), circuits that embody logic for execution, or both, where appropriate. This disclosure encompasses any suitable combination of hardware and software.
[0129] While this disclosure has described several exemplary embodiments, there are many modifications, substitutions, and alternative equivalents that fall within the scope of this disclosure. Therefore, those skilled in the art will understand that many systems and methods not expressly shown or described herein can be devised to embody the principles of this disclosure and thus fall within the spirit and scope of this disclosure.
Claims
1. A method for decoding a video bitstream encoded using at least one processor, The steps include obtaining a Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit, The steps include determining whether the VCL NAL unit is the first VCL NAL unit of a picture unit (PU) that includes the VCL NAL unit, A step to determine whether the VCL NAL unit is the first VCL NAL unit of an access unit (AU) including the PU, based on the determination that the VCL NAL unit is the first VCL NAL unit of the PU, wherein the VCL NAL unit is the first VCL NAL unit following a picture header NAL unit, or a first flag in the VCL NAL unit is set to indicate that a picture header is included in the slice header contained in the VCL NAL unit, and one or more of the conditions of a plurality of conditions are true, then the VCL NAL unit is determined to be the first VCL NAL unit of the AU, and the plurality of conditions are (1) The first value of nuh_layer_id of the VCL NAL unit is less than the second value of nuh_layer_id of the previous picture in the decoding order, (2) The first value of ph_pic_order_cnt_lsb of the VCL NAL unit is different from the second value of ph_pic_order_cnt_lsb of the previous picture in the decoding order, (3) A determination step including determining that the first picture order count value derived for the VCL NAL unit is different from the second picture order count value of the previous picture in the decoding order, The step of decoding the AU based on the VCL NAL unit based on determining that one or more of the above-mentioned conditions are true, If the second flag, which is a gate flag for picture encoding information, has a first value, then the multiple third flags corresponding to the picture encoding information are not signaled in the picture parameter set corresponding to the picture header. A method wherein the plurality of third flags include at least one from among rpl_info_in_ph_flag and wp_info_in_ph_flag.
2. The method according to claim 1, wherein the VCL NAL unit is determined to be the first VCL NAL unit of the PU based on the determination that the VCL NAL unit is the first VCL NAL unit following the picture header NAL unit.
3. The method according to claim 1, wherein the VCL NAL unit is determined to be the first VCL NAL unit of the PU, based on the determination that a fourth flag in the VCL NAL unit is set to indicate that the slice header included in the VCL NAL unit includes the picture header.
4. A device for decoding an encoded video bitstream, At least one memory configured to store program code, Apparatus comprising: a processor configured to read program code and to operate as instructed by the program code to perform the method described in any one of claims 1 to 3.
5. The steps include encoding the pictures from the source video sequence into the encoded video sequence, The steps include obtaining a video coding layer (VCL) network abstraction layer (NAL) unit in relation to the aforementioned video sequence, The steps include determining whether the VCL NAL unit is the first VCL NAL unit of a picture unit (PU) that includes the VCL NAL unit, A step to determine whether the VCL NAL unit is the first VCL NAL unit of an access unit (AU) including the PU, based on the determination that the VCL NAL unit is the first VCL NAL unit of the PU, wherein the VCL NAL unit is the first VCL NAL unit following a picture header NAL unit, or a first flag in the VCL NAL unit is set to indicate that a picture header is included in the slice header contained in the VCL NAL unit, and one or more of the conditions of a plurality of conditions are true, then the VCL NAL unit is determined to be the first VCL NAL unit of the AU, and the plurality of conditions are (1) The first value of nuh_layer_id of the VCL NAL unit is less than the second value of nuh_layer_id of the previous picture in the decoding order, (2) The first value of ph_pic_order_cnt_lsb of the VCL NAL unit is different from the second value of ph_pic_order_cnt_lsb of the previous picture in the decoding order, (3) A determination step including determining that the first picture order count value derived for the VCL NAL unit is different from the second picture order count value of the previous picture in the decoding order, The step of decoding the AU based on the VCL NAL unit based on determining that one or more of the above-mentioned conditions are true, If the second flag, which is a gate flag for picture encoding information, has a first value, then the multiple third flags corresponding to the picture encoding information are not signaled in the picture parameter set corresponding to the picture header. A method wherein the plurality of third flags include at least one from among rpl_info_in_ph_flag and wp_info_in_ph_flag.