METHOD FOR VIDEO ENCODING, NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM, PROGRAM, AND METHOD FOR TRANSMITTING BITSTREAM - Patent application
By improving syntax element signaling in video encoding, particularly for advanced standards like VVC, the method enhances encoding efficiency and video quality by accurately identifying IRAP and GDR pictures, addressing limitations in existing video encoding techniques.
Patent Information
- Application Number
- JP2023184964
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-20
- Filing Date
- 2023-10-27
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing video encoding techniques struggle with efficient signaling of syntax elements, leading to suboptimal video quality and bit rate management, particularly in advanced standards like VVC, where flexibility in block partitioning and syntax signaling is limited.
The method involves signaling syntax elements such as NAL unit types and picture headers to enhance decoding processes, allowing for better identification of Intra Random Access Points (IRAP) and Gradual Intra Refresh (GDR) pictures, enabling more efficient video encoding and decoding by determining NAL unit types and generating syntax elements based on picture characteristics.
This approach improves video encoding efficiency by optimizing bit rate usage and maintaining video quality, addressing limitations in existing standards like VVC through enhanced syntax signaling and block partitioning techniques.
Smart Images

Figure 0007783237000038 
Figure 0007783237000039 
Figure 0007783237000040
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 63 / 027,718, filed May 20, 2020. Priority to No. 10 / 2009, entitled "Signaling of Syntax Elements in Video Coding" No. 6,239,693, which is incorporated by reference in its entirety.
[0002] This disclosure relates to video encoding and compression, and in particular, but not exclusively, to video encoding. The present invention relates to a method and apparatus for signaling syntax elements in a mobile communication system. [Background technology]
[0003] Various video encoding techniques may be used to compress the video data. The video encoding is performed according to one or more video coding standards. The encoding standard is Versatile Video Coding (VVC). :VVC), Joint Exploration Test Model (Joint Exploration test st Model:JEM), High-Performance Video Encoding (H.265 / High-Effic iency Video Coding: HEVC), Advanced Video Coding (H.264 / Advanced Video Coding:AVC), MPEG(Moving P Video coding is a method for encoding a video signal. ,Generally, prediction methods exploit the redundancy present in a ,video image or sequence, e.g. The main purpose of video coding techniques is to Use lower bit rates while avoiding or minimizing video quality degradation The purpose of JPEG compression is to compress video data into a format. Summary of the Invention
[0004] This disclosure provides example techniques for signaling syntax elements during video encoding. Provide.
[0005] According to a first aspect of the present disclosure, there is provided a method for video encoding, the method comprising: The decoder uses a Picture Parameter Set (P A picture corresponding to a PS is sent to one or more Network Abstraction Layers (NALs). whether it contains a NAL (Non-Abstractive Abstraction Layer) unit, and whether one or more NAL units have the same NAL unit type receiving a first syntax element in the PPS that identifies the PPS; The picture corresponding to the Picture Header (PH) is Intra Random Access Point (Intra Random Access Point: IRAP picture or Gradual Intra Refresh The second sequence in the PH that identifies whether the picture is a GDR (Grayscale Refreshing) picture. Further, the decoder receives a second syntax element. , determine the value of the first syntax element.
[0006] According to a second aspect of the present disclosure, there is provided a method for video encoding, the method comprising: The decoder determines whether the picture corresponding to the PPS contains one or more NAL units. and one or more NAL units have the same NAL unit type receiving a first syntax element in the PPS that identifies whether This method is used when the decoder detects that the picture corresponding to PH is an IRAP picture or a GDR picture. and receiving a second syntax element in the PH that specifies whether the The method further comprises: a decoder generating a second syntax element based on the value of the first syntax element; This includes determining the value of the element.
[0007] According to a third aspect of the present disclosure, there is provided a method for video encoding, the method comprising: The decoder determines whether a picture contains one or more NAL units and Identifying whether multiple NAL units have the same NAL unit type Further, the decoder receives a first syntax element that corresponds to the first syntax element. Determine the second syntax element in the PH associated with the picture based on the syntax element. do.
[0008] According to a fourth aspect of the present disclosure, there is provided a method for video encoding, the method comprising: A decoder receives the syntax element and executes a decoding process based on the value of the syntax element. Furthermore, if the syntax element for picture is equal to 0 and If any slice of a picture has a nal_unit_type equal to GDR_NUT, then the GDR picture The value of the syntax element for this slice is equal to 0, and all other slices of the picture have the same value na l_unit_type and the picture is a GDR picture after the first slice of the picture is received It is recognized as a
[0009] According to a fifth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus comprising: one or more processors and executable by the one or more processors The one or more processors include a memory configured to store instructions. is configured to perform any method according to the first aspect of the present disclosure when executed.
[0010] According to a sixth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus comprising: one or more processors and executable by the one or more processors The one or more processors include a memory configured to store instructions. is configured to perform any method according to the second aspect of the present disclosure when executed.
[0011] According to a seventh aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus comprising: one or more processors and executable by the one or more processors The one or more processors include a memory configured to store instructions. is configured to perform any method according to the third aspect of the present disclosure.
[0012] According to an eighth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus comprising: one or more processors and executable by the one or more processors The one or more processors include a memory configured to store instructions. is configured to perform any method according to the fourth aspect of the present disclosure.
[0013] According to a ninth aspect of the present disclosure, there is provided a method for implementing a method for programming a computer program, the method comprising: and then causing the one or more computer processors to execute any of the methods according to the first aspect of the present disclosure. A non-transitory storage device for video encoding storing computer executable instructions for implementing the method of A computer-readable storage medium is provided.
[0014] According to a tenth aspect of the present disclosure, there is provided a method for implementing the method by one or more computer processors. When executed, the one or more computer processors perform tasks according to the second aspect of the present disclosure. A non-transitory storage device for video encoding that stores computer executable instructions for implementing any method. A computer-readable storage medium is provided.
[0015] According to an eleventh aspect of the present disclosure, there is provided a method for implementing the method by one or more computer processors. When executed, the one or more computer processors perform tasks according to the third aspect of the present disclosure. A non-transitory storage device for video encoding that stores computer executable instructions for implementing any method. A computer-readable storage medium is provided.
[0016] According to a twelfth aspect of the present disclosure, there is provided a method for implementing the method by one or more computer processors. When executed, the one or more computer processors perform tasks according to the fourth aspect of the present disclosure. A non-transitory storage device for video encoding that stores computer executable instructions for implementing any method. A computer-readable storage medium is provided. [Brief explanation of the drawings]
[0017] A more particular description of the examples of the present disclosure can be made by reference to specific examples that are illustrated in the accompanying drawings. These drawings depict only some examples and therefore do not limit the scope. The following examples are to be considered non-limiting and should be taken as illustrative only and should not be construed as limiting. This will provide more specific and detailed explanations and commentary.
[0018] [Figure 1] FIG. 1 is a block diagram illustrating an example video encoder according to some implementations of this disclosure.
[0019] [Figure 2] FIG. 2 is a block diagram illustrating an example video decoder according to some implementations of this disclosure.
[0020] [Figure 3] FIG. 3 is a diagram illustrating an example of a picture divided into multiple coding tree units (CTUs) according to some implementations of the present disclosure.
[0021] [Figure 4A] FIG. 4A is a diagram illustrating a multi-type tree partitioning model according to some implementations of the present disclosure. [Figure 4B] FIG. 4B is a diagram illustrating a multi-type tree partitioning model according to some implementations of the present disclosure. [Figure 4C] FIG. 4C is a diagram illustrating a multi-type tree partitioning model according to some implementations of the present disclosure. [Figure 4D] FIG. 4D is a diagram illustrating a multi-type tree partitioning model according to some implementations of the present disclosure.
[0022] [Figure 5] FIG. 5 is a diagram illustrating intra-coded regions between multiple inter-pictures according to some implementations of the present disclosure.
[0023] [Figure 6] FIG. 6 is a block diagram illustrating an example apparatus for video encoding according to some implementations of this disclosure.
[0024] [Figure 7]FIG. 7 is a flow diagram illustrating an example process for video encoding according to some implementations of this disclosure.
[0025] [Figure 8] FIG. 8 is a flow diagram illustrating an example process for video encoding according to some implementations of the present disclosure.
[0026] [Figure 9] FIG. 9 is a flow diagram illustrating an example process for video encoding according to some implementations of this disclosure.
[0027] [Figure 10] FIG. 10 is a flow diagram illustrating an example process for video encoding according to some implementations of this disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0028] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the detailed description, non-limiting examples are provided to aid in understanding the subject matter presented herein. Although numerous specific details are provided, it is understood that various modifications may be used. For example, the subject matter presented herein may be used in conjunction with digital video capabilities. It will be apparent to those skilled in the art that the present invention may be implemented in many types of electronic devices including:
[0029] Throughout this specification, the terms "one embodiment," "an embodiment," "example," and "some implementations" are used interchangeably. References to "some embodiments," "some examples," or similar terms are to the particular features being described. means that a property, structure, or feature is included in at least one embodiment or example Furthermore, any feature, structure, element, or feature described with respect to one or more embodiments may be used interchangeably. Features are also applicable to other embodiments unless expressly stated otherwise.
[0030] Throughout this disclosure, the terms "first," "second," "third," etc., all refer to, e.g., Instructions for reference to related elements such as devices, components, compositions, steps, etc. Unless otherwise expressly stated, no spatial or temporal order is implied. For example, a "first device" and a "second device" are not separate entities. Two connected devices, or two parts, components or movements of the same device It may refer to a state in which something is operational, or it may be arbitrarily named.
[0031] The terms "module," "submodule," "electrical circuit," "sub-electrical circuit," and "circuit" "," "subcircuit," "unit," or "subunit" means one or more processes. Memory (shared, dedicated, or group) that stores code or instructions that can be executed by the processor. A module may or may not contain stored code or instructions. A module or electrical circuit may contain one or more electrical circuits. may contain one or more indirectly connected components. The components may or may not be physically connected to each other, or may be located near each other. It may or may not be placed there.
[0032] As used herein, the terms "when" or "when" may be used interchangeably depending on the context. These terms may be understood as "in the course of" or "in response to" in the claims. Although stated, it does not imply that any related limitations or characteristics are conditional or optional. For example, the method may be such that: i) when or if condition X exists, Function or operation X' is performed and ii) condition Y exists; A method may include a step in which a function or an action Y' is performed. It is performed using both the ability to perform action X' and the ability to perform function or operation Y'. Thus, both functions X' and Y' may be executed at different times in multiple steps of the method. It may be implemented on a practice basis.
[0033] A unit or module can be either purely software-based or purely hardware-based. It may be implemented purely by hardware or by a combination of hardware and software. In a software implementation, for example, a unit or module may perform a particular function. functionally related coding blocks or blocks connected directly or indirectly to each other for May include software components.
[0034] FIG. 1 may be used with many video coding standards that use block-based processing. 1 is a block diagram illustrating an exemplary block-based hybrid video encoder 100. In encoder 100, a video frame is divided into multiple video blocks for processing. For each given video block, either an inter prediction approach or an intra prediction approach is used. In inter prediction, prediction is made based on one or more predictors. is calculated through motion estimation and motion compensation based on pixels from previously reconstructed frames. In intra prediction, the predictor is formed based on the reconstructed pixels in the current frame. Through mode decision, the best predictor is selected to predict the current block. It may be selected.
[0035] A prediction residual, which represents the difference between the current video block and its predictor, is sent to the transform circuit 102. The transform coefficients are then passed from the transform circuit 102 to the quantization circuit 104 for entropy reduction. 4. The quantized coefficients are then processed to generate the compressed video bitstream. As shown in FIG. information, motion vectors, reference picture indexes, and intra prediction modes, Prediction-related information 110 from inter-prediction and / or intra-prediction circuits 112 may also be used. Also, the compressed video bitstream is fed through an entropy coding circuit 106. It is stored in memory 114.
[0036] In the encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through an inverse quantization circuit 116 and an inverse transform circuit 118. This reconstructed prediction residual is used to find the unfiltered reconstructed pixels for the current video block. The block predictor 120 is combined with the block predictor 120 to generate the block predictor 120 .
[0037] Intra prediction (also called "spatial prediction") is a method for predicting the time between images in the same video picture and / or sequence. Samples from nearby blocks in the RICE that have already been coded (called reference samples) The spatial prediction predicts the current video block using pixels from the video signal The spatial redundancy inherent in the
[0038] Inter prediction (also called "temporal prediction") is a method of predicting a video picture that has already been coded. Temporal prediction uses reconstructed pixels from the current video block to predict the current video block. The temporal redundancy inherent in the audio signal is reduced. The temporal prediction signal for a CU or coding block is usually the current CU and the One or more motion vectors indicating the amount and direction of motion between the inter-reference n Vector (MV). In addition, multiple reference pictures are signaled by If supported, one reference picture index is additionally transmitted, which is , to identify from which reference picture in the reference picture store the temporal prediction signal comes. Used for.
[0039] After spatial and / or temporal prediction is performed, the intra / inter prediction in the encoder 100 The mode decision circuit 121 determines the best predictive mode, for example, based on a rate-distortion optimization method. The block predictor 120 is then subtracted from the current video block, The resulting prediction residuals are decorrelated using a transform circuit 102 and a quantizer circuit 104. The resulting quantized residual coefficients are inversely quantized by the inverse quantization circuit 116 and then inversely transformed. The reconstructed residual is then inverse transformed by the transform circuit 118 to form a reconstructed residual, which is then The deblocking signal is then added back to the measured block to form the reconstructed signal for the CU. Filter, Sample Adaptive Offset SAO, and / or Adaptive in-Loop Filter An in-loop filter 115, such as an Automated Filter (ALF), may be applied to the reconstructed CU. ,The reconstructed CU is then placed in the reference picture store of the picture buffer 117, Used to encode further video blocks. Output Video Bitstream 1 14, coding mode (inter or intra), prediction mode information, motion The information and the quantized residual coefficients are all sent to the entropy coding unit 106. are then compressed and packed to form a bitstream.
[0040] For example, the deblocking filter supports not only AVC and HEVC, but also the latest version of VVC. In HEVC, it is called SAO (Sample Adaptive Offset). An additional in-loop filter is specified to further improve coding efficiency. The latest version of the VC standard introduces yet another loop filter called ALF (Adaptive Loop Filter). In-group filters are being actively investigated and are likely to be included in the final standard.
[0041] These in-loop filter operations are optional. The implementation of these operations may affect coding efficiency. They also help improve the coding rate and visual quality to save computation. It may also be turned off if determined by the device 100.
[0042] Intra prediction is usually based on unfiltered reconstructed pixels, while inter prediction is based on If these filter options are turned on by the encoder 100, filtered reconstruction is performed. It should be noted that the method is based on pixel-by-pixel.
[0043] FIG. 2 illustrates an exemplary block-based video coding scheme that may be used with many video coding standards. 2 is a block diagram illustrating a video decoder 200. The decoder 200 is the same as the decoder 100 of FIG. In the decoder 200, the input video bits are The stream 201 is first decoded through entropy decoding 202 to obtain the quantization coefficients The quantized coefficient levels are then dequantized 204. and processed through an inverse transform 206 to obtain the reconstructed prediction residual. The block predictor mechanism implemented in the code selector 212 uses the decoded prediction information configured to perform either intra prediction 208 or motion compensation 210 based on The set of unfiltered reconstructed pixels is then added to the inverse transform 206 using adder 214. sums the reconstructed prediction residual from the block and the prediction output generated by the block predictor mechanism. This is obtained by:
[0044] The reconstructed block is then passed through an in-loop filter 209 and then filtered using the reference picture The picture is stored in the picture buffer 213, which functions as a storage device. The reconstructed video in the image is sent to drive a display device and is then used to generate future video streams. Used to predict locking when the in-loop filter 209 is on. A filtering operation is performed on these reconstructed pixels to produce the final reconstructed video An output 222 is derived.
[0045] Generic Video Coding (VVC) The 10th J.S.A. International Conference was held in San Diego, USA from April 10th to 20th, 2018. At the VET conference, JVET presented VVC and First draft of VVC Test Model 1 (VTM1) The first new coding feature in VVC was nested multitype trees. It was decided to include quad trees with 2-way and 3-way splits. The coding block partition structure includes both the coding process and the decoding process. The reference software VTM was developed, which implemented both the fusion and fusion processes. Updated at the meeting.
[0046] In VVC, the input video picture is divided into blocks called CTUs. U uses a quad tree with a nested multi-type tree structure to achieve the same prediction. along with CUs that define regions of pixels that share a mode (e.g., intra or inter) , CU. The term "unit" covers all components such as luma and chroma. The term "block" may refer to a region of an image that contains a particular component (e.g., luminance). may be used to define a region covering different components (e.g., luminance vs. chrominance). The chroma (degree) block is empty when considering chroma sampling formats such as 4:2:0. The interim location may vary.
[0047] Dividing a picture into CTUs FIG. 3 illustrates a picture divided into multiple CTUs 302 according to some implementations of the present disclosure. 1 is a diagram illustrating an example of a sensor 300. FIG.
[0048] In VVC, a picture is divided into a series of CTUs. The CTU concept is different from that of HEVC. For a picture with three sample arrays, the CTU is the same as the chroma sample It consists of an NxN block of luma samples together with the corresponding two blocks of
[0049] The maximum allowed size of a luminance block in a CTU is specified as 128x128 (however, (The maximum size of a degree transformation block is 64x64.)
[0050] Splitting CTUs using a tree structure In HEVC, CTUs are divided into CUs using a 4-element tree structure called a coding tree. It is divided and adapted to various local characteristics. Interpicture (time) or Into The decision whether to encode a picture area using spatial prediction is made by the Each leaf CU is further divided into one or two PUs according to the PU split type. The prediction process is the same for each PU. The relevant information is transmitted to the decoder on a PU-by-PU basis. After obtaining the residual block by performing Divide into Transform Units (TUs) according to the element tree structure One of the key characteristics of the HEVC structure is that it contains CUs, PUs, and TUs. It has multiple division concepts.
[0051] VVC uses nested multi-part segmentation structures with two-part and three-part segmentation structures. The quadtree with type tree replaces the concept of multiple partition unit types. , i.e., except as necessary for CUs whose size is too large for the maximum transform length. Removed the separation between CU, PU and TU concepts, supporting more flexibility in CU division shapes In the coding tree structure, a CU may have either a square or a rectangular shape. The CTU is first divided into four-element trees (also known as quadtrees). Then, the 4-element tree leaf nodes are further divided by the multitype tree structure. It can be done.
[0052] 4A to 4D illustrate a multi-type tree partitioning model according to some implementations of the present disclosure. As shown in Figures 4A to 4D, the multi-type tree structure has four Split type, vertical 2 split 402 (SPLIT_BT_VER), horizontal 2 split 404 (SPLIT_BT_HOR), vertical There are three-way split 406 (SPLIT_TT_VER) and three-way split 408 (SPLIT_TT_HOR). A Type Tree leaf node is called a CU, and as long as the CU is not too large for the maximum transformation length, This segmentation allows prediction and transformation processes without any further division. This is mostly used for nested multi-type tree coding blocks. In a quadtree with block structure, CU, PU, and TU have the same block size. If the maximum supported transform length is smaller than the width or height of the color components of the CU, If so, an exception occurs.
[0053] VVC Syntax In VVC, the first layer of syntax signaling is the bitstream. A stream is a NAL stream divided into a set of NAL units. The bit signals common control parameters such as SPS and PPS to the decoder. The other layer contains video data. Layer (VCL) NAL units contain slices of coded video A coded picture is called an access unit and is divided into one or more slices. It may be encoded as
[0054] The encoded video sequence is then decoded using instantaneous decoder refresh. All subsequent IDR (Decoder Refresh) pictures Video pictures are coded as slices. A new IDR picture is a slice of the previous video. Each N signal indicates that a video segment has ended and a new video segment has begun. An AL unit begins with a one-byte header and contains a raw byte sequence payload (R The RBSP is followed by a byte sequence payload (RBSP). The slices are binary coded so that they , may be padded with 0 bits to ensure that the length is an integral number of bytes. A slice consists of a slice header and slice data. is defined as a set of CUs.
[0055] The picture header concept was adopted at the 16th JVET conference and was the first VCL for pictures. Now transmitted once per picture as a NAL unit. Previously, slice headers It also proposes grouping some syntax elements that were previously in the picture header into this picture header. Syntax elements that functionally only need to be transmitted once per picture are: Moved to the picture header instead of being transmitted multiple times in slices for a particular picture I was able to do it.
[0056] In the VVC standard, a syntax table is provided to show all the allowed bitstream symbols. Other constraints on syntax are specified directly in other clauses. Tables 1 and 2 below show the slides in VVC. The syntax table for the header and PH. The meaning of some syntax is also explained. Examples are provided after the syntax table. [Table 1] JPEG0007783237000002.jpg252170JPEG0007783237000003.jpg191170 [Table 2] JPEG0007783237000005.jpg251170JPEG0007783237000006.jpg247170JPEG0007783237000007.jpg253170JPEG0007783237000008.jpg203170
[0057] The meaning of selected syntax elements ph_temporal_mvp_enabled_flag is the flag for the slice associated with the picture header (PH). Specifies whether a temporal motion vector predictor can be used for inter prediction of . If ral_mvp_enabled_flag is equal to 0, the syntax of the slice associated with the PH The element is constrained so that the temporal motion vector predictor is not used when decoding the slice. Otherwise (if ph_temporal_mvp_enabled_flag is equal to 1), the time The inter-motion vector predictor may be used during decoding of the slice associated with the PH. If not present, the value of ph_temporal_mvp_enabled_flag is inferred to be equal to 0. Decoded Picture Buffer (DP) If there is no reference picture in B) that has the same spatial resolution as the current picture, then ph_tempo The value of ral_mvp_enabled_flag shall be equal to 0.
[0058] The maximum number of subblock-based merge MVP candidates, MaxNumSubblockMergeCand, is It is derived as follows:
number
[0059] slice_collocated_from_l0_flag equal to 1 is used for temporal motion vector prediction Specifies that the co-located picture is derived from reference picture list 0. sl equal to 0 ice_collocated_from_l0_flag specifies the number of co-located pictures used for temporal motion vector prediction. Specifies that the image is derived from reference picture list 1. slice_type is equal to B or P. ph_temporal_mvp_enabled_flag is equal to 1 and slice_collocated_from_l0_flag is equal to 1. If g is not present, the following applies: - If rpl_info_in_ph_flag is equal to 1, slice_collocated_from_l0_flag is equal to ph_coll Inferred to be equal to ocated_from_l0_flag. - Otherwise (rpl_info_in_ph_flag is equal to 0 and slice_type is equal to P), The value of slice_collocated_from_l0_flag is inferred to be equal to 1.
[0060] slice_collocated_ref_idx is the collocated picture used for temporal motion vector prediction Identify the reference index of
[0061] slice_type is equal to P, or slice_type is equal to B and slice_collocated If _from_l0_flag is equal to 1, slice_collocated_ref_idx is the reference picture list 0. The value of slice_collocated_ref_idx is 0 to NumRefIdxActiv The range is e[0]-1 or less.
[0062] If slice_type is equal to B and slice_collocated_from_l0_flag is equal to 0, then ce_collocated_ref_idx means an entry in reference picture list 1, and slice_collocation The value of cated_ref_idx must be in the range from 0 to NumRefIdxActive[1]-1. If slice_collocated_ref_idx is not present, the following applies: - If rpl_info_in_ph_flag is equal to 1, the value of slice_collocated_ref_idx is Inferred to be equal to ted_ref_idx. - Otherwise (if rpl_info_in_ph_flag is equal to 0), slice_collocated_ref_idx The value is inferred to be equal to 0.
[0063] The picture referenced by slice_collocated_ref_idx is the entire coded picture. It is a bitstream conformance requirement that the slices are the same.
[0064] pic_width_in_luma_samp of the reference picture referenced by slice_collocated_ref_idx The values of pic_height_in_luma_samples and pic_height_in_luma_samples are the sum of the pic_width_i Equal to the values of n_luma_samples and pic_height_in_luma_samples, and RprConstraintsAc tive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] is equal to 0. These are the requirements for bitstream conformance.
[0065] The values of RprConstraintsActive[i][j] are specified in Section 8 of the VVC standard as summarized below. Note that this is derived in 3.2.
[0066] Decoding Process for Reference Picture List Construction This process is invoked at the beginning of the decoding process for each slice of a non-IDR picture. A reference picture is addressed through a reference index, which is , is an index into the reference picture list. When decoding an I slice, No reference picture list is used to decode the data. When decoding a P slice, the slice Only reference picture list 0 (i.e., RefPicList[0]) is used to decode the device data. When decoding a B slice, reference picture list 0 and Both reference picture list 1 (i.e., RefPicList[1]) and reference picture list 2 (i.e., RefPicList[1]) are used.
[0067] At the beginning of the per-slice decoding process of a non-IDR picture, the reference picture list Re fPicList[0] and RefPicList[1] are derived. The reference picture list is defined in 8.3.3. Used to create reference pictures as specified in the will be done.
[0068] For an I-slice of a non-IDR picture that is not the first slice of the picture, RefPicList RefPicList[0] and RefPicList[1] may be derived for bitstream conformance checking. However, their derivation is based on the current picture or the picture that follows it in decoding order. For P slices that are not the first slice of a picture, RefP icList[1] may be derived for bitstream conformance checking, The derivation is necessary for decoding the current picture or a picture that follows the current picture in decoding order. isn't it.
[0069] Reference picture lists RefPicList[0] and RefPicList[1], reference picture scaling ratio RefPicScale[i][j][0] and RefPicScale[i][j][1], and the reference picture scale flags The RprConstraintsActive[0][j] and RprConstraintsActive[1][j] are derived as follows: will be done.
number
[0070] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset and scaling_win_bottom_offset and scaling_win_bottom_offset are applied to the picture size for scaling ratio calculation. If not present, scaling_win_left_offset, scaling_win_rig The values of ht_offset, scaling_win_top_offset, and scaling_win_bottom_offset are , pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset , and is inferred to be equal to pps_conf_win_bottom_offset.
[0071] The value of SubWidthC*(scaling_win_left_offset+scaling_win_right_offset) is pic_width_ in_luma_samples and SubHeightC*(scaling_win_top_offset+scaling_win_bottom _offset) is less than pic_height_in_luma_samples.
[0072] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
number
number
[0073] NAL unit syntax Similar to HEVC, the VVC standard uses a byte array to identify basic information about each NAL unit. At the beginning of the NAL unit, one NAL unit header table with a total length of 2 bytes is included. Table 3 shows the syntax elements present in the current NAL unit header. is doing. [Table 3]
[0074] In Table 3, the first bit is the forbidden_zero_bit, which indicates that there is no error during transmission. Used to determine whether an error occurred. 0 indicates that the NAL unit is normal. 1 means that there is a syntax violation. For streams, the corresponding value shall be equal to 0. The next bit is nuh_rese rved_zero_bit, which is reserved for future use and is equal to 0. 6 bits contain the value of the syntax nuh_layer_id, which identifies the layer to which the NAL unit belongs. The value of nuh_layer_id ranges from 0 to 55. Other values of h_layer_id are reserved for future use. unit_type is the NAL unit type, i.e., the NAL Used to identify the type of RBSP data structure contained in the unit. [Table 4] JPEG0007783237000016.jpg74168
[0075] Gradual Intra Refresh Low latency and error resilience are two important considerations in practical video transmission systems. Intra-refresh, which periodically inserts IRAP pictures, is a key factor in determining whether a time period is to limit error propagation between frames and to enhance the error resilience of the bitstream. However, the coding efficiency of inter-coding is lower than that of intra-coding. It is much better than the conventional method of sending data over a network at a fixed transmission rate. The relatively large size of intra pictures can sometimes cause delay problems. This can lead to undesirable network congestion and packet loss. To solve this problem, we use the inter-picture method shown in Figure 5. Gradual Intra Refresh (GDR), which distributes the intra-coding area, is adopted in the VVC standard. As shown in Figure 5, two regions are defined. Part 2 represents the clean region. The clean area corresponds to the pixels that have been refreshed during the current GDR period and are not dark. Part 1 corresponds to the one area that has not been refreshed. Part 2 corresponds to the intra-coded area. The principle of GDR is that the time division within the same GDR period The clean region is created using pixels derived only from the refreshed region of the reference picture. The goal of VVC is to ensure that all pixels in the picture are reconstructed. Three GDR-related syntax elements signaled by ph_gdr_or_irap_pic_flag, ph_ gdr_pic_flag and ph_recovery_poc_cnt are present. Table 5 shows the correspondence in the picture header. 1 shows the GDR signaling and associated semantics. [Table 5]
[0076] ph_gdr_or_irap_pic_flag equal to 1 indicates that the current picture is a GDR or IRAP picture. ph_gdr_or_irap_pic_flag equal to 0 indicates that the current picture is It may or may not be a GDR picture and may be an IRAP picture. Identify something.
[0077] ph_gdr_pic_flag equal to 1 indicates that the picture associated with the PH is a GDR picture. ph_gdr_pic_flag equal to 0 specifies that the picture associated with the PH is G Specifies that it is not a DR picture. If it does not exist, the value of ph_gdr_pic_flag is equal to 0. If sps_gdr_enabled_flag is equal to 0, the value of ph_gdr_pic_flag is shall be equal to 0.
[0078] If ph_gdr_or_irap_pic_flag is equal to 1 and ph_gdr_pic_flag is equal to 0, the PH-related The linked pictures are IRAP pictures.
[0079] ph_recovery_poc_cnt identifies the recovery point of a decoded picture in output order If the current picture is a GDR picture, the variable recoveryPointPocVal is set to It is derived as follows.
number
[0080] The current picture is a GDR picture, and the current GDR picture in CLVS decoding order Following this, there exists a picture picA with PicOrderCntVal equal to recoveryPointPocVal. If so, picture picA is referred to as the recovery point picture. The first picture in output order with a PicOrderCntVal greater than recoveryPointPocVal is referred to as a recovery point picture. A recovery point picture is a picture that is located immediately before the current It must not precede a GDR picture. It must be associated with the current GDR picture and must not be re Pictures with PicOrderCntVal less than coveryPointPocVal are not GDR picture rotations. The value of ph_recovery_poc_cnt is between 0 and MaxPicOrderCntL. The range is sb-1 or less.
[0081] sps_gdr_enabled_flag is equal to 1 and the current picture's PicOrderCntVal is If the recoveryPointPocVal of the decoded GDR picture is greater than or equal to the recoveryPointPocVal of the currently decoded GDR picture in output order, The decoded picture and the next decoded picture are related in decoding order and are placed in a different order than the GDR picture. by starting the decoding process from the previous IRAP picture (if any) before the The image will match exactly the corresponding picture created by
[0082] Mixed NAL types in one picture The HEVC standard requires that the NAL types of slices within a picture must be the same. Unlike the JPEG2000 standard, it is not possible to mix IRAP and non-IRAP NAL unit types within a single picture. The purpose of such a function is to provide region-based random access using subpictures. For example, in the case of 360-degree video streaming, Some regions may be viewed by more users than others. More frequent IRAP pictures are displayed to improve the tradeoff with viewpoint switching delay. A pixel can be used to encode areas that are more viewed than other areas. For this reason, one flag, pps_mixed_nalu_types_in_pic_flag, is introduced in PPS. If the flag is equal to 1, the flag indicates that each picture that references the PPS has two or more NAL units and no NAL units have the same value of nal_unit_type Otherwise (flag equals 0), each picture that references the PPS is 1. NAL units of each picture that has one or more NAL units and references a PPS In addition, the flag pps_mixed_nalu_types_in_pic_fla If g is equal to 1, then for any particular picture, some NAL units are Some have a specific IRAP NAL unit type, others have one specific non-IRAP NAL unit type. One additional bitstream conformance constraint applies: to have L unit types. In other words, the NAL units of any particular picture, as specified below, An entity cannot have more than one IRAP NAL unit type and cannot have more than one It is not possible to have non-IRAP NAL unit types.
[0083] For any particular picture's VCL NAL units, the following applies: - If pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type is It is the same for all VCL NAL units, and a picture or PU is a picture. has the same NAL unit type as the coded slice NAL unit of the It is considered to be - Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), the following applies: can be. A picture shall have at least two sub-pictures. - The VCL NAL units of a picture have two or more different nal_unit_type values. - There are no VCL NAL units of the picture with nal_unit_type equal to GDR_NUT. - The VCL NAL unit of at least one subpicture of a picture has IDR_W_RADL, I In a picture with a specific value of nal_unit_type equal to DR_N_LP, or CRA_NUT All other subpicture VCL NAL units have nal_unit_type equal to TRAIL_NUT. It shall have.
[0084] In the current VVC, mvd_l1_zero_flag is used for pictures without any conditional constraints. Signaled in the header (PH), but controlled by the flag mvd_l1_zero_flag The described characteristics are applicable only if the slice is a bidirectionally predicted slice (B slice). Therefore, flag signaling is performed for the slice associated with the picture header. If is not a B slice, it is redundant.
[0085] Similarly, in another example, the sequence parameter set (SPS) signaled Only the corresponding valid flags (sps_bdof_pic_present_flag, sps_dmvr_pic_present_flag) Only when true, ph_disable_bdof_flag and ph_disable_dmvr_flag respectively are enabled in PH. However, as shown in Figure 6, the flags ph_disable_bdof_flag and ph_d The characteristic controlled by isable_dmvr_flag determines whether the slice is a bidirectionally predicted slice (B-slice). Therefore, the signaling of these two flags The tag is redundant or invalid if the slice associated with the picture header is not a B slice. It is beneficial. [Table 6]
[0086] Another example is the syntax element ph_collocated_from_l0_fl, as shown in Table 7. ag, it is possible to indicate that the co-located picture is from list 0 or list 1. As another example, as shown in FIG. 8, the weighting table for bidirectional prediction is The syntax element pred_weight_table() is a related syntax element. [Table 7] [Table 8] JPEG0007783237000022.jpg255169
[0087] The problem is related to the syntax ph_temporal_mvp_enabled_flag. In the current VVC, The resolution of the co-located picture selected for TMVP derivation is the same as the resolution of the current picture. Since they are the same, the following checks the value of ph_temporal_mvp_enabled_flag: There are stream conformance constraints. If there is no reference picture in the DPB that has the same spatial resolution as the current picture, ph_tem The value of the local_mvp_enabled_flag shall be equal to 0.
[0088] However, in the current VVC, the resolution of the co-located picture does not affect the enablement of TMVP. not only the image size, but also the offset applied to the picture size for scaling ratio calculations. It also affects the enablement of TMVP. However, in the current VVC, ph_temporal_mvp_en The offset is not taken into account in the bitstream adaptation of abled_flag.
[0089] Furthermore, the picture referenced by slice_collocated_ref_idx is the coded picture. It is a requirement of bitstream conformance that the bitstream size is the same for all slices of a However, if a coded picture has multiple slices and there is commonality among all these slices, If there are no common reference pictures, this bitstream conformance is unlikely to be met. Furthermore, in such cases, ph_temporal_mvp_enabled_flag is constrained to 0. It is necessary.
[0090] According to the current VVC standard, an IRAP picture is a picture containing all of the associated NAL units. all are treated as one picture with the same nal_unit_type belonging to an IRAP NAL type. Specifically, the following defines an IRAP picture in the VVC standard: It is used for Intra Random Access Point (IRAP) picture: All VCL NAL units The data is coded with the same value of nal_unit_type in the range from IDR_W_RADL to CRA_NUT. This is a picture.
[0091] An IRAP picture does not use inter prediction in its decoding process and is not a CRA picture or It may be an IDR or IDR picture. The first picture in the bitstream in decoding order is , IRAP or GDR picture. Required parameters if reference is required. If the data set is available, the IRAP pictures and their corresponding frames in decoding order in CLVS. All non-IRAP pictures following the IRAP picture are considered to be part of any picture preceding the IRAP picture in decoding order. It is possible to accurately decode the data without performing the decoding process of the decoder.
[0092] The value of pps_mixed_nalu_types_in_pic_flag for an IRAP picture is equal to 0. pps_mixed_nalu_types_in_pic_flag for the picture is equal to 0 and any If a slice has a nal_unit_type in the range from IDR_W_RADL or greater to CRA_NUT or less, the picture All other slices in the picture have the same value of nal_unit_type, and the picture is an IRAP picture. It is recognized as Kutcha.
[0093] As can be seen from the above, for each IRAP picture, the The corresponding PPS must have its pps_mixed_nalu_types_in_pic_flag equal to 0. Similarly, in the current VVC standard, a GDR picture requires all One picture where the nal_unit_type of this NAL is equal to GDR_NUT as specified below. It is referred to as a Gradual Decoding Refresh (GDR) picture: Each VCL NAL unit is assigned to a GDR_NUT Pictures with equal nal_unit_type.
[0094] All NAL units in a GDR picture must have the same NAL type. Therefore, the flag pps_mixed_nalu_types_i in the corresponding PPS referenced by the GDR picture n_pic_flag cannot be equal to 1.
[0095] On the other hand, there are two flags, namely ph_gdr_or_irap_pic_flag and ph_gdr_pic_flag However, whether a picture is an IRAP picture or a GDR picture The flag ph_gdr_or_irap_pic_fla is signaled in the picture header to indicate whether If g is equal to 1 and the flag ph_gdr_pic_flag is equal to 0, the current picture is a single I It is a RAP picture. The flag ph_gdr_or_irap_pic_flag is equal to 1 and the flag ph_gdr_pi If c_flag is equal to 1, the current picture is a GDR picture. According to the C standard, the value of the flag pps_mixed_nalu_types_in_pic_flag in the PPS must be taken into account. Instead, these two flags are allowed to be signaled as 1 or 0. However, as mentioned above, the NAL units of a picture must have the same nal_unit_type. that is, the corresponding pps_mixed_nalu_types_in_pic_flag is 0. Only one picture can be one IRAP picture or one GDR picture. Therefore, the existing IRAP / GDR signaling in the picture header is dr_or_irap_pic_flag and ph_gdr_pic_flag or both are equal to 1 (i.e. indicates that the current picture is either an IRAP picture or a GDR picture. ), and the corresponding pps_mixed_naly_types_in_pic_flag is equal to 1 (i.e., the current pixel This is problematic when multiple NAL types exist in a channel.
[0096] The flags mvd_l1_zero_flag, ph_disable_bdof_flag, and ph_disable_dmvr_flag The characteristics controlled by the slice are only available if the slice is a bidirectionally predicted slice (B slice). Therefore, according to the method of the present disclosure, the associated slice is a B slice. We propose to signal these flags only if the reference picture list is P If signaled within H (e.g., rpl_info_in_ph_flag=1), it is coded All slices of a given picture use the same reference picture signaled in the PH. Note that this means that the reference picture list is signed in the PH. nulled and signaled that the current picture is not bi-predictive If the configuration list indicates, the flags mvd_l1_zero_flag, ph_disable_bdof_flag and ph_dis In the first embodiment, the picture header redundant signaling or signaling due to incorrect values being sent for some of the syntax within These are transmitted in the picture header (PH) to prevent undefined decoding behavior. Some conditions are added to the syntax. Some examples based on the embodiment are as follows: where the variable num_ref_entries[i][RplsIdx[i]] is the number of reference pixels in list i. It represents the number of cha.
number
[0097] Alternatively, these conditions can be written in a more compact form to achieve the same result. A bidirectionally predicted slice (B-slice) or bidirectionally predicted picture must be included in at least one list. Since the current slice / picture always has a reference picture in list 1, You only need to check whether the item has a char. Below are some examples of alternative condition checks: Shown below.
number
[0098] The meaning of mvd_l1_zero_flag is also explained to handle the case where it is not signaled. It will be corrected.
[0099] mvd_l1_zero_flag equal to 1 indicates that the mvd_coding(x0,y0,1) syntax structure is parsed. and MvdL1[x0][y0][compIdx] and MvdCpL1[x0][y0][cpIdx][compIdx] are indicates that it is set equal to 0 for cpIdx=0..1 and cpIdx=0..2. The 1_zero_flag indicates that the mvd_coding(x0,y0,1) syntax structure is parsed. If not present, the value of mvd_l1_zero_flag is inferred to be equal to 0.
[0100] Below are some examples of conditional signaling of the syntax element ph_disable_dmvr_flag: Shown below.
number
[0101] Similarly, examples of alternative condition checks are given below.
number
[0102] The meaning of ph_disable_dmvr_flag has also been revised to handle the case where it is not signaled. It will be corrected accordingly.
[0103] ph_disable_dmvr_flag equal to 1 disables inter bi-prediction based decoder motion vector refinement. Specifies that is disabled for the slice associated with PH. ph_disable equal to 0 _dmvr_flag indicates whether inter bi-prediction based decoder motion vector refinement is associated with PH. Identify what may or may not be valid in a slice.
[0104] If ph_disable_dmvr_flag is not present, the following applies: -If sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 0, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 0. -If sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 1, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 1. - Otherwise (if sps_dmvr_enabled_flag is equal to 0), the value of ph_disable_dmvr_flag is assumed to be equal to 1.
[0105] If the value of ph_disable_dmvr_flag is not present, the alternative methods for deriving the value are shown below. . - All conditions are considered to derive the value of ph_disable_dmvr_flag and the value is clearly signed. Nulled or implicitly derived: sps_dmvr_enabled_flag equals 1 , and sps_dmvr_pic_present_flag is equal to 0, the value of ph_disable_dmvr_flag is 0 are assumed to be equal. - If sps_dmvr_enabled_flag is equal to 0 and sps_dmvr_pic_present_flag is equal to 0, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 1. -sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 1, Additionally, if rpl_info_in_ph_flag is equal to 0, the value of ph_disable_dmvr_flag is equal to X. It is speculated that (X is clearly signaled). -sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 1, Additionally, if rpl_info_in_ph_flag is equal to 1 and num_ref_entries[1][RplsIdx[1]]>0, If this is the case, the value of ph_disable_dmvr_flag is inferred to be equal to X (X is not explicitly signaled). (This is the case.) -Otherwise (sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 1, and rpl_info_in_ph_flag is equal to 1, and num_ref_entries[1][Rpls Idx[1]]==0), the value of ph_disable_dmvr_flag is inferred to be equal to 1.
[0106] The syntax element ph_disable_dmvr_flag explicitly signals the third and fourth conditions. Therefore, if ph_disable_dmvr_flag is not present, the derivation of ph_disable_dmvr_flag is These may be omitted from the If ph_disable_dmvr_flag is not present, the following applies: -If sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 0, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 0. - If sps_dmvr_enabled_flag is equal to 0 and sps_dmvr_pic_present_flag is equal to 0, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 1. -Otherwise (sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 1, and rpl_info_in_ph_flag is equal to 1, and num_ref_entries[1][Rpls Idx[1]]==0), the value of ph_disable_dmvr_flag is inferred to be equal to 1.
[0107] The condition can be rewritten simply as follows: If ph_disable_dmvr_flag is not present, the following applies: -If sps_dmvr_enabled_flag is equal to 1 and sps_dmvr_pic_present_flag is equal to 0, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 0. -Otherwise (sps_dmvr_enabled_flag is equal to 0 or sps_dmvr_pic_present_f lag is equal to 1), the value of ph_disable_dmvr_flag is inferred to be equal to 1.
[0108] If the value of ph_disable_dmvr_flag is not present, another alternative method for deriving the value is given below. show. If ph_disable_dmvr_flag is not present, the following applies: -sps_dmvr_pic_present_flag is equal to 0), and the value of ph_disable_dmvr_flag is 1 Inferred to be equal to vr_enabled_flag. - If sps_dmvr_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is equal to 0 , the value of ph_disable_dmvr_flag is inferred to be equal to 1-sps_dmvr_enabled_flag. -sps_dmvr_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is equal to 1, and Furthermore, if num_ref_entries[1][RplsIdx[1]]>0, the value of ph_disable_dmvr_flag is 1-sps_dmvr Inferred to be equal to _enabled_flag. -Otherwise (sps_dmvr_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is 1 and num_ref_entries[1][RplsIdx[1]]==0), ph_disable_dmvr_fla The value of g is assumed to be equal to 1.
[0109] The syntax element ph_disable_dmvr_flag explicitly signals the second and third conditions. Therefore, if ph_disable_dmvr_flag is not present, the derivation of ph_disable_dmvr_flag is These may be omitted from the If ph_disable_dmvr_flag is not present, the following applies: - If sps_dmvr_pic_present_flag is equal to 0, then ph_disable_dmvr_flag has a value of 1 Inferred to be equal to _enabled_flag. Otherwise, the value of ph_disable_dmvr_flag is inferred to be equal to 1.
[0110] Some examples of conditional signaling of the syntax element ph_disable_bdof_flag are given below: Shown below.
number
[0111] Similarly, examples of alternative condition checks are given below.
number
[0112] The meaning of ph_disable_bdof_flag has also been revised to handle the case where it is not signaled. It will be corrected accordingly.
[0113] ph_disable_bdof_flag equal to 1 disables bidirectional directional optical flow interprediction Specifies that inter bi-prediction is disabled for the slice associated with PH. Equal to 0. The new ph_disable_bdof_flag disables bidirectional directional optical flow prediction-based interferometry. Bi-prediction may or may not be enabled for the slice associated with the PH. Identify the following:
[0114] If ph_disable_bdof_flag is not present, the following applies: - If sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 0 , the value of ph_disable_bdof_flag is inferred to be equal to 0. - If sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 1, In this case, the value of ph_disable_dmvr_flag is inferred to be equal to 1. - otherwise (if sps_bdof_enabled_flag is equal to 0), the value of ph_disable_bdof_flag is assumed to be equal to 1.
[0115] If the value of ph_disable_bdof_flag is not present, an alternative method for deriving the value is shown below. . All conditions are considered to derive the value of ph_disable_bdof_flag, and the value is clearly signaled. When linked or implicitly derived: - If sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 0 , the value of ph_disable_bdof_flag is inferred to be equal to 0. - If sps_bdof_enabled_flag is equal to 0 and sps_bdof_pic_present_flag is equal to 0, In this case, the value of ph_disable_bdof_flag is inferred to be equal to 1. -sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 1, Additionally, if rpl_info_in_ph_flag is equal to 0, the value of ph_disable_bdof_flag is equal to X. It is speculated that (X is clearly signaled). -sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 1, Additionally, if rpl_info_in_ph_flag is equal to 1 and num_ref_entries[1][RplsIdx[1]]>0, If this is the case, the value of ph_disable_bdof_flag is inferred to be equal to X (X is not explicitly signaled). (This is the case.) -Otherwise (sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 1, and rpl_info_in_ph_flag is equal to 1, and num_ref_entries[1][Rpls Idx[1]]==0), the value of ph_disable_bdof_flag is inferred to be equal to 1.
[0116] The syntax element ph_disable_bdof_flag explicitly signals the third and fourth conditions. Therefore, if ph_disable_bdof_flag is not present, the derivation of ph_disable_bdof_flag is These may be omitted from the If ph_disable_bdof_flag is not present, the following applies: - If sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 0 , the value of ph_disable_bdof_flag is inferred to be equal to 0. - If sps_bdof_enabled_flag is equal to 0 and sps_bdof_pic_present_flag is equal to 0, In this case, the value of ph_disable_bdof_flag is inferred to be equal to 1. -Otherwise (sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 1, and rpl_info_in_ph_flag is equal to 1, and num_ref_entries[1][Rpls Idx[1]]==0), the value of ph_disable_bdof_flag is inferred to be equal to 1.
[0117] The condition can be rewritten simply as follows: If ph_disable_bdof_flag is not present, the following applies: - If sps_bdof_enabled_flag is equal to 1 and sps_bdof_pic_present_flag is equal to 0 , the value of ph_disable_bdof_flag is inferred to be equal to 0. -Otherwise (sps_bdof_enabled_flag is equal to 0 or sps_bdof_pic_present_f lag is equal to 1), the value of ph_disable_bdof_flag is inferred to be equal to 1.
[0118] If the value of ph_disable_bdof_flag is not present, another alternative method for deriving the value is given below. show. If ph_disable_bdof_flag is not present, the following applies: - If sps_bdof_pic_present_flag is equal to 0, then ph_disable_bdof_flag has a value of 1. Inferred to be equal to _enabled_flag. - If sps_bdof_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is equal to 0 , the value of ph_disable_bdof_flag is inferred to be equal to 1-sps_bdof_enabled_flag. -sps_bdof_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is equal to 1, and Furthermore, if num_ref_entries[1][RplsIdx[1]]>0, the value of ph_disable_bdof_flag is 1-sps_bdof Inferred to be equal to _enabled_flag. -Otherwise (sps_bdof_pic_present_flag is equal to 1 and rpl_info_in_ph_flag is equals 1 and num_ref_entries[1][RplsIdx[1]]==0), ph_disable_bdof_fla The value of g is assumed to be equal to 1.
[0119] The syntax element ph_disable_bdof_flag explicitly signals the second and third conditions. Therefore, if ph_disable_bdof_flag is not present, the derivation of ph_disable_bdof_flag is These may be omitted from the If ph_disable_bdof_flag is not present, the following applies: - If sps_bdof_pic_present_flag is equal to 0, then ph_disable_bdof_flag has a value of 1. Inferred to be equal to _enabled_flag. Otherwise, the value of ph_disable_bdof_flag is inferred to be equal to 1.
[0120] Additionally, the syntax elements ph_collocated_from_l0_flag and weight_table() The signaling condition is that the slices with which the two types of syntax elements are associated are B-slices. This is corrected because it is only available if the syntax element is a rice. Examples of signaling are shown in Tables 9 to 11 below. [Table 9]
[0121] The meaning of ph_collocated_from_l0_flag also addresses the case where it is not signaled. is corrected to
[0122] ph_collocated_from_l0_flag equal to 1 specifies the same ph_collocated_from_l0_flag used for temporal motion vector prediction. Specifies that the position picture is derived from reference picture list 0. ph_col equal to 0 located_from_l0_flag specifies the co-located picture referenced by the picture used for temporal motion vector prediction. Specifies that it is derived from picture list 1.
[0123] If ph_collocated_from_l0_flag is not present, the following applies: -If num_ref_entries[0][RplsIdx[0]] is greater than 1, ph_collocated_from_l0_flag The value is inferred to be 1. -Otherwise (num_ref_entries[1][RplsIdx[1]] is greater than 1), ph_collocated The value of _from_l0_flag is inferred to be 0. [Table 10] [Table 11]
[0124] Similarly, examples of alternative condition checks are given below.
number
[0125] The meaning of the syntax elements in pred_weight_table() also depends on how they are signaled. will be modified to address cases where
[0126] If both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1, then num_l1 _weights specifies the number of weights signaled for an entry in reference picture list 1. The value of num_l1_weights must be between 0 and Min(15,num_ref_entries[1][RplsIdx[1]]). The range is as below.
[0127] The variable NumWeightsL1 is derived as follows:
number
[0128] The value of num_l1_weights in the sense of the syntax element in pred_weight_table() exists. If not present, an alternative method for deriving its value is given below. If both pps_weighted_bipred_flag and wp_info_in_ph_flag are equal to 1, then num_l1_w eights specifies the number of weights signaled for entries in reference picture list 1 The value of num_l1_weights is between 0 and Min(15,num_ref_entries[1][RplsIdx[1]]). If not present, the value of num_l1_weights is inferred to be equal to 0.
[0129] The variable NumWeightsL1 is derived as follows:
number
[0130] The value of num_l1_weights in the sense of the syntax element in pred_weight_table() exists. If not present, another alternative method of deriving its value is given below.
number
[0131] Conceptually, any bit available only in the B slice to avoid signaling redundant bits. References from both the List 0 and List 1 reference picture lists for syntax elements A signaling condition is used to check whether the current picture has a reference picture. The check condition is to check the reference picture list (e.g., list 0 / list The above method is limited to checking the size of both the reference picture list and the image. The check condition is that the current picture is a reference picture in both list 0 and list 1. It may be any other way to indicate whether a picture has a reference from the picture list. For example, the flag indicates that the current picture uses reference pictures from both list 0 and list 1. The metric may be signaled to indicate whether it has
[0132] No syntax elements are signaled within the PH, but reference picture list information is signaled. If nulled, the value of the syntax element is the current picture in list 0 and list It has both reference pictures in list 0 and list 1, or it has only reference pictures in list 0 or list 1. In one example, the ph_collocated_from_l0_flag is derived using information about whether the If not signaled, the value is the only reference picture that the current picture has. In another example, if sps_bdof_enabled_flag is equal to 1 and sps_bdof_pi If c_present_flag is equal to 1 but ph_disable_bdof_flag is not signaled, then This is done according to the proposed signaling condition for ph_disable_bdof_flag. s[0][RplsIdx[0]] is equal to 0 or num_ref_entries[1][RplsIdx[1]] is equal to 0 Therefore, ph_disable_bdof_flag is not signaled under this condition. In the current VVC, the resolution of the co-located picture is TMVP This may affect the validity of the picture for scaling ratio calculations. An offset applied for size may also affect TMVP enablement. However, in the current VVC, the bitstream adaptation of ph_temporal_mvp_enabled_flag In the second embodiment, as shown below, the offset is not taken into account in ph_t The value of emporal_mvp_enabled_flag is applied to the picture size for scaling ratio calculation. Add a bitstream conformance constraint to the current VVC that requires it to depend on the offset We propose that:
[0133] In the DPB, select a picture with the same spatial resolution and scaling ratio calculation as the current picture. If there is no reference picture with the same offset applied to the pixel size, The value of p_enabled_flag shall be equal to 0.
[0134] The above sentence can also be written in another way:
[0135] Reference pictures in the DPB that have the associated variable value RprConstraintsActive[i][j] equal to 0 If not present, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.
[0136] In the current VVC, the picture referenced by slice_collocated_ref_idx is coded. The bitstream conformance requirement is that the bitstream However, if a coded picture has multiple slices and all of these slices This bitstream conformance is met if there are no common reference pictures between the In the third embodiment of the present disclosure, the ph_temporal_mvp_enabled_flag is The bitstream conformance requirement is that there must be a common reference picture between all slices of the current picture. Based on the embodiment, the VVC Some exemplary variations to the standard are illustrated below.
[0137] ph_temporal_mvp_enabled_flag specifies whether inter prediction for the slice associated with PH is performed temporally. Specifies whether the temporal motion vector predictor is enabled. If ag is equal to 0, the syntax elements of the slice associated with PH are It is assumed that the decoding is constrained to not use temporal motion vector predictors. Otherwise (if ph_temporal_mvp_enabled_flag is equal to 1), temporal motion vector prediction is performed. The parameter may be used in decoding the slice associated with PH. If not present The value of ph_temporal_mvp_enabled_flag is inferred to be equal to 0. If there is no reference picture with the same spatial resolution as the picture, ph_temporal_mvp_enabled_ The value of flag shall be equal to 0. The common reference pin for all slices associated with the PH If no architecture is present, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.
[0138] ph_temporal_mvp_enabled_flag specifies whether inter prediction for the slice associated with PH is performed temporally. Specifies whether the temporal motion vector predictor is enabled. If ag is equal to 0, the syntax elements of the slice associated with PH are It is assumed that the decoding is constrained to not use temporal motion vector predictors. Otherwise (if ph_temporal_mvp_enabled_flag is equal to 1), temporal motion vector prediction is performed. The parameter may be used in decoding the slice associated with PH. If not present The value of ph_temporal_mvp_enabled_flag is inferred to be equal to 0. If there is no reference picture with the same spatial resolution as the picture, ph_temporal_mvp_enabled_ The value of flag shall be equal to 0. Common to all inter-slices associated with PH If no reference picture exists, the value of ph_temporal_mvp_enabled_flag shall be equal to 0. Let's say.
[0139] ph_temporal_mvp_enabled_flag specifies whether inter prediction for the slice associated with PH is performed temporally. Specifies whether the temporal motion vector predictor is enabled. If ag is equal to 0, the syntax elements of the slice associated with PH are It is assumed that the decoding is constrained to not use temporal motion vector predictors. Otherwise (if ph_temporal_mvp_enabled_flag is equal to 1), temporal motion vector prediction is performed. The parameter may be used in decoding the slice associated with PH. If not present The value of ph_temporal_mvp_enabled_flag is inferred to be equal to 0. If there is no reference picture with the same spatial resolution as the picture, ph_temporal_mvp_enabled_ The value of flag shall be equal to 0. It is shared among all non-intra slices associated with PH. If no common reference picture exists, the value of ph_temporal_mvp_enabled_flag shall be equal to 0. Let's say.
[0140] In one example, the bitstream adaptation for slice_collocated_ref_idx is as follows: It is simplified. pic_width_in_luma_sample of the reference picture referenced by slice_collocated_ref_idx The values of pic_height_in_luma_samples and pic_height_in_luma_samples are the sum of the pic_width_in_luma_samples of the current picture, respectively. luma_samples and pic_height_in_luma_samples values, and RprConstraintsActi ve[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] is equal to 0 is a requirement for bitstream conformance.
[0141] If the value of pps_mixed_nalu_types_in_pic_flag is equal to 1, each picture that references a PPS If a packet has two or more NAL units, but those NAL units have the same nal_unit_type On the other hand, the current picture header signaling does not have the associated Even if the value of the flag pps_mixed_nalu_types_in_pic_flag in the PPS is equal to 1, The value of ph_gdr_or_irap_pic_flag and ph_gdr_pic_flag shall be signaled as 1. NAL units of one IRAP picture or one GDR picture are allowed. Since all the signals must have the same nal_unit_type, such a signaling scenario Rio should not be tolerated.
[0142] In one example, the picture header is We propose to condition on the presence of the flag ph_gdr_or_irap_pic_flag in the data. , ph_gdr_or_irap_pic_f only if the value of pps_mixed_nalu_types_in_pic_flag is equal to 0 lag is signaled. Otherwise, the flag pps_mixed_nalu_types_in_pic_flag is If equal to 1, the flag ph_gdr_or_irap_pic_flag is not signaled and is 0 Table 12 shows an example of the proposed modifications. [Table 12]
[0143] In one example, if the flag pps_mixed_nalu_types_in_pic_flag is equal to 1, the signaling to require the corresponding value of the flag ph_gdr_or_irap_pic_flag to be equal to 1. To achieve this, we propose a bitstream conformance constraint. The frame compatibility constraints can be specified as follows:
[0144] ph_gdr_or_irap_pic_flag equal to 1 indicates that the current picture is a GDR or IRAP picture. ph_gdr_or_irap_pic_flag equal to 0 indicates that the current picture is It may or may not be a GDR picture and may be an IRAP picture. If the value of pps_mixed_nalu_types_in_pic_flag is equal to 1, then ph_gdr The value of _or_irap_pic_flag shall be equal to 0.
[0145] In one example, the signaling of pps_mixed_nalu_types_in_pic_flag is changed from the PPS level to Propose to move to picture level, slice level, or other coding level. For example, suppose a flag is moved to the picture header, and the flag's name is ph_mix. May be changed to ed_nalu_type_in_pic_flag. Additionally, ph_gdr_or_irap_pic We propose to use flags to condition the signaling of _flag. Specifically: , ph_gdr_or_rap_pic_fl only if the flag ph_mixed_nalu_type_in_pic_flag is equal to 0 ag is signaled. Otherwise, the flag ph_mixed_nalu_type_in_pic_flag is set to 1. In this case, the flag ph_gdr_or_rap_pic_flags is not signaled and is inferred to be 0. In another example, if the value of ph_mixed_nalu_type_in_pic_flag is equal to 1, then ph_gdr_or_ir Adds a bitstream conformance constraint, such as the value of ap_pic_flag must be equal to 0. In another example, we propose to use the ph_mixed_nalu_type_in_pic_flag flag as a condition. We suggest using ph_gdr_or_irap_pic_flag to set the The flag ph_mixed_nalu_type_in_pic_flag is true if and only if the value of or_rap_pic_flag is equal to 0. Otherwise, if the value of ph_gdr_or_rap_pic_flag is equal to 1, the flag The flag ph_mixed_nalu_type_in_pic_flag is not signaled and is always assumed to be 0. can be.
[0146] In one example, pps_mixed_na is set only for pictures that are neither IRAP nor GDR pictures. We propose to apply the value of lu_types_in_pic_flag. Therefore, the meaning of pps_mixed_nalu_types_in_pic_flag needs to be modified as follows: pps_mixed_nalu_types_in_pic_flag equal to 1 indicates that the IRAP picture refers to a PPS. Each picture that is neither a VCL nor a GDR picture has two or more VCL NAL units, and Specifies that no VCL NAL units have the same value of nal_unit_type. The same pps_mixed_nalu_types_in_pic_flag is used for IRAP pictures that refer to PPSs. Each picture that is not a DR picture has one or more VCL NAL units, and The VCL NAL units of each picture that references a PPS shall have the same value of nal_unit_type. Identify what you are doing.
[0147] On the other hand, in the current VVC standard, all NAL units of one GDR picture are GDR. It is required that the nal_unit_type must be equal to R_NUT. Definition of GDR pictures so that the corresponding value of xed_nal_types_in_pic_flag is equal to 0 The following bitstream conformance constraints apply:
[0148] Gradual Decoding Refresh (GDR) picture: Each VCL NAL unit is a GDR_NUT pps_mixed_nal for GDR pictures. The value of u_types_in_pic_flag is equal to 0. pps_mixed_nalu_types_in_pic for the picture _flag is equal to 0 and any slice of the picture has a nal_unit_type that is GDR_NUT. If so, all other slices of the picture must have the same value of nal_unit_type and the picture must After the first slice of the picture is received, the picture is recognized as a GDR picture.
[0149] Another embodiment involves removing the GDR NAL unit type from the NAL unit header. At the same time, the syntax element ph indicates whether the current picture is a GDR picture. I suggest using only _gdr_or_irap_pic_flag and ph_gdr_pic_flag.
[0150] The pps_mixed_nalu_types_in_pic_flag constraint applies to IRAP and GDR pictures. Unlike the above method, which applies to both, in the following the constraints are applied to IRAP pictures only. We propose three methods that are applicable to GDR pictures but do not apply to GDR pictures.
[0151] In one example, the picture header is We propose to condition on the presence of the flag ph_gdr_pic_flag in the pps_mix parameter. The flag ph_gdr_pic_flag is signaled only if the value of ed_nalu_types_in_pic_flag is equal to 0. Otherwise, if the flag pps_mixed_nalu_types_in_pic_flag is equal to 1, In this case, the flag ph_gdr_pic_flag is not signaled and is inferred to be 0, i.e. The current picture cannot be a GDR picture. The formula (Table 13) is modified as follows after the proposed signaling conditions are applied: . [Table 13]
[0152] ph_gdr_pic_flag equal to 1 indicates that the picture associated with the PH is a GDR picture. ph_gdr_pic_flag equal to 0 specifies that the picture associated with the PH is G Specifies that it is not a DR picture. If not present, the value of ph_gdr_pic_flag is pps_mix If ed_nalu_types_in_pic_flag is 0, it is equal to 0, and pps_mixed_nalu_types_in_pic_flag If sps_gdr_enabled_flag is 1, it is assumed to be equal to the value of ph_gdr_or_irap_pic_flag. If g is equal to 0, the value of ph_gdr_pic_flag shall be equal to 0.
[0153] In one example, ph_gdr_or_irap_pic_flag is 1 and pps_mixed_nalu_types_in_pic_f A single bit: if lag is 1, then ph_gdr_pic_flag must be equal to 1. We propose to introduce a stream conformance constraint, which is specified as follows: ph_gdr_pic_flag equal to 1 means the picture associated with PH is a GDR picture ph_gdr_pic_flag equal to 0 specifies that the picture associated with the PH is a GD If not present, the value of ph_gdr_pic_flag is equal to 0. If sps_gdr_enabled_flag is equal to 0, the value of ph_gdr_pic_flag is shall be equal to 0. ph_gdr_or_irap_pic_flag shall be equal to 1 and pps_mixed_nalu_ty If pes_in_pic_flag is equal to 1, the value of ph_gdr_pic_flag must be equal to 1. If ph_gdr_or_irap_pic_flag is equal to 1 and ph_gdr_pic_flag is equal to 0, the PH-related The linked pictures are IRAP pictures.
[0154] In one example, the flag pps_mixed_nalu_types_in_pic_flag is applied only to non-IRAP pictures. Specifically, this method proposes that the pps_mixed_nalu_types_in_pic_flag The meaning needs to be modified to: pps_mixed_nalu_types_in_pic_flag equal to 1 specifies that each non-IRAP picture that references a PPS If a packet has two or more VCL NAL units and the VCL NAL units have the same value, pps_mixed_nalu_types_in_pic equal to 0. _flag specifies that each non-IRAP picture that references a PPS contains one or more VCL NAL units. The VCL NAL units of each picture that has a PPS and references a PPS have the same value of NA Identifies that it has l_unit_type.
[0155] The method is implemented in an application specific integrated circuit (ASIC). Integrated Circuit (ASIC), Digital Signal Processor (Digi Digital Signal Processor (DSP), Digital Signal Processing Device (D Digital Signal Processing Device (DSPD), Programmable Logic Device (PL D), Field Programmable Gate Arrays (Field Programmable FPGA, controller, microcontroller, micro Comprising one or more circuits, including a processor or other electronic components The method can be performed using an apparatus. The apparatus can be used in conjunction with other hardware to perform the method. Circuitry in combination with hardware or software components may also be used. Each module, sub-module, unit, or sub-unit disclosed herein may be used in one or more may be implemented at least in part using multiple circuits.
[0156] FIG. 6 illustrates an example apparatus for video encoding according to some implementations of the present disclosure. The device 600 may be a mobile phone, a tablet computer, a digital book, or the like. The terminal may be a broadcast terminal, a tablet device, or a personal digital assistant.
[0157] As shown in FIG. 6, the device 600 includes a processing component 602, a memory 604, a power supply a supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor controller 614, and one or more of a communication component 616. Good too.
[0158] The processing component 602 typically handles display, telephony, data communication, camera operation, and audio recording. Controls the overall operation of the device 600, including operational operations. 2 is one or more instructions for completing all or part of the steps of the above method The processing component 602 may include multiple processors 620. One or more components that facilitate interaction between the management component 602 and other components. For example, the processing component 602 may include a multimedia Multimedia component 608 and processing component 602. It may also include a multimedia module.
[0159] Memory 604 stores various types of data to support the operation of device 600. Examples of such data include any data that operates on device 600. instructions for applications or methods, contact data, phone book data, messages; Photos, videos, etc. The memory 604 may be any type of volatile or non-volatile memory. The memory 6 may be implemented by a nonvolatile storage device or a combination thereof. 04 is a static random access memory (Static Random Access Memory emory: SRAM, Electrically Erasable Programmable Read-Only Memory (Elec trically Erasable Programmable Read-Only Memory:EEPROM, erasable programmable read-only memory (Era sable Programmable Read-Only Memory: EPRO M), Programmable Read-Only Memory PROM, Read-Only Memory ry:ROM), magnetic memory, flash memory, magnetic disk or compact disk It may also be a
[0160] The power supply component 606 provides power to the different components of the device 600. The power supply component 606 includes a power supply management system, one or more power supplies. power supply and other components involved in generating, managing, and distributing power for device 600. The composition may also contain other components.
[0161] The multimedia component 608 provides an output interface between the device 600 and the user. In some examples, the screen is a liquid crystal display. Liquid Crystal Display (LCD) and Touch Panel (T If the screen includes a touch panel, The screen may be implemented as a touch screen to receive input signals from a user. The touch panel detects touches, slides, and gestures on the touch panel. The touch sensor may include one or more touch sensors. It not only detects the boundaries of the motion, but also the duration and time associated with the contact or sliding motion. Pressure may also be detected. In some examples, the multimedia component 608 The device 600 may include a front camera and / or a rear camera. In operation modes such as mode or video mode, the front camera and / or rear The camera is capable of receiving external multimedia data.
[0162] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 may be a microphone (MICrophone). The device 600 can be configured in various modes, such as a call mode, a recording mode, and a voice recognition mode. In this mode of operation, the microphone is configured to receive an external audio signal. The received audio signal may then be stored in memory 604 or transmitted to a communication component. 616. In some examples, the audio component 610 may be The audio signal may further include a speaker for outputting an audio signal.
[0163] The I / O interface 612 interfaces with the processing component 602 and peripheral devices. The peripheral device interface forms an interface between the base module and the peripheral device. The module may be a keyboard, a click wheel, a button, etc. Buttons include the Home button, Volume button, Start button, and Lock button. These include, but are not limited to:
[0164] The sensor component 614 provides different aspects of the condition assessment for the device 600. For example, the sensor component 614 may include one or more sensors for detecting the It can detect the on / off state of the 00 and the relative position of the components. For example, the components are the display and keypad of the device 600. Component 614 may also be used to change the arrangement, placement, or placement of device 600 or components of device 600. the presence or absence of a user's contact with the device 600, the direction or acceleration / deceleration of the device 600, It can also detect deceleration and temperature changes of the device 600. Sensor component 61 4 is configured to detect the presence of nearby objects without any physical contact. The sensor component 614 may include a proximity sensor. It includes optical sensors such as CMOS or CCD image sensors used in In some examples, the sensor component 614 may include an acceleration sensor, a gyro sensor, or the like. The sensor may further include a loop sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0165] The communication component 616 may be configured to communicate wired or wirelessly between the apparatus 600 and other devices. The device 600 is configured to facilitate the use of a wireless network such as WiFi, 4G, or a combination thereof. Based on the communication standard, a wireless network can be accessed. The broadcast signal or broadcast related information is transmitted to the broadcast component 616. broadcast channel from an external broadcast management system. The communication component 616 may include a near field communication (NFC) to facilitate short-range communication. It also includes an NFC (Near Field Communication) module. For example, the NFC module may transmit radio frequency identification information (RFID). RFID technology, Infrared Data Association (Infr IrDA (Internal Data Association) technology, Ultra Wideband (Ultra -Wide Band (UWB) technology, Bluetooth (registered trademark) h:BT) technology, as well as other technologies.
[0166] In one example, the device 600 may include an application specific integrated circuit (AS) for implementing the above-described method. IC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGs) A) a controller, microcontroller, microprocessor or other electronic It may be implemented by one or more of the elements.
[0167] The non-transitory computer-readable storage medium may be, for example, a hard disk drive. Disk Drive (HDD), Solid-State Drive Drive:SSD), flash memory, hybrid drive or solid state drive Solid-State Hybrid Drive (S SHD), read-only memory (ROM), compact disk read-only memory (C Compact Disc Read-Only Memory (CD-ROM), magnetic tape The storage medium may be a tape, a floppy disk, etc.
[0168] FIG. 7 is a flowchart illustrating an example process for video encoding according to some implementations of the present disclosure. FIG.
[0169] In step 702, processor 620 refers to, corresponds to, or or whether the picture associated with it contains one or more NAL units. and whether one or more NAL units have the same NAL unit type. The first syntax element in the PPS specifies whether the PPS is to be
[0170] In step 704, processor 620 refers to or corresponds to PH, or indicates whether the associated picture is an IRAP picture or a GDR picture. The second syntax element in the PH specifies whether the
[0171] In step 706, processor 620 determines the first syntax element based on the value of the second syntax element. Determine the value of the syntax element.
[0172] In some examples, the processor 620 is implemented in a decoder.
[0173] In some examples, the first syntax element equal to 1 refers to or is a PPS. Each picture corresponds to or is associated with two or more VCL NAL units. contains bits and two or more VCL NAL units have the same NAL unit type The first syntax element equal to 0 refers to the PPS. Each picture that corresponds to or is associated with one or more It contains a VCL NAL unit, and one or more VCL NAL units Identify that they have the same NAL unit type.
[0174] In some instances, the second syntax element, equal to 1, refers to or is associated with PH. The corresponding or associated picture is an IRAP picture or a GDR picture. The second syntax element, which specifies that the PH is a , the picture corresponding to it or associated with it is an IRAP picture or G It is also determined that the picture is not a DR picture.
[0175] In some examples, the processor 620 may determine whether the picture is an IRAP picture or a GDR picture. Upon determining that the object is a texture, the value of the first syntax element is required to be 0. To do this, we apply constraints to the first syntax element to create a second syntax element. Constrain the value of the first syntax element based on the value of the tax element.
[0176] FIG. 8 is a flowchart illustrating an example process for video encoding according to some implementations of the present disclosure. FIG.
[0177] In step 802, the processor 620 determines whether a picture is one or more NAL units. and whether one or more NAL units contain the same NAL unit type and receiving a first syntax element that specifies whether the first syntax element has a
[0178] In step 804, the processor 620 generates a picture based on the first syntax element. The second syntax element in the PH associated with the object is determined.
[0179] In some examples, the processor 620 is implemented in a decoder.
[0180] In some examples, the second syntax element indicates whether the picture is a GDR picture or an IR picture. Identify whether it is an AP picture.
[0181] In some examples, the first syntax element is a sequence within a PPS associated with a picture. It is being gunned down.
[0182] In some examples, the first syntax element is pps_mixed_nalu_types as described above. _in_pic_flag.
[0183] In some examples, the processor 620 determines that the first syntax element is equal to 0. determining that a second syntax element is signaled within the PH in response to the and, in response to determining that the first syntax element is equal to 1, If it is determined that the element is not signaled in the PH and the second syntax element is 0, and inferring the associated picture based on the first syntax element. Constrain the second syntax element in PH, where the first syntax element equals 1 Each picture that references, corresponds to, or is associated with a PPS is It contains two or more NAL units, and two or more NAL units are the same NAL unit. The first syntax element, which specifies that the unit has no unit type and is equal to 0, Each picture that references, corresponds to, or is associated with a PPS is one or contains multiple NAL units, and one or more of the NAL units are the same Additionally, the second syntax equal to 1 specifies that the NAL unit type is The hex element identifies the picture as a GDR or IRAP picture, and The second syntax element, equal to 0, indicates whether the picture is an IRAP picture or a GDR picture. Identify whether it is a
[0184] In some examples, the processor 620 may It determines that both syntax elements are received and that the first syntax element is equal to 1. In response to the and applying a constraint to the picture based on the first syntax element. Constrains the second syntax element in the given PH. The first syntax element equals 1. Each picture that references, corresponds to, or is associated with a PPS is It contains two or more NAL units, and two or more NAL units are the same NAL unit. The second syntax element, which specifies that the unit has no unit type and is equal to 0, Identifies that the picture is neither an IRAP picture nor a GDR picture. The flag element may be pps_mixed_nalu_types_in_pic_flag as described above.
[0185] In some examples, the first syntax element is a symbol in a PH associated with a picture. In some examples, the first syntax element is ph_m as described above. It may also be mixed_nalu_types_in_pic_flag.
[0186] In some examples, the processor 620 determines that the first syntax element is equal to 0. determining that a second syntax element is signaled within the PH in response to the and, in response to determining that the first syntax element is equal to 1, The second syntax element is determined to be 0 and the second syntax element is determined to be not signaled in the PH. Inferring a second syntax element in PH based on the first syntax element The first syntax element equal to 1 refers to or corresponds to the PH. Each picture corresponding to or associated with a picture contains two or more NAL units. and no two or more NAL units have the same NAL unit type. The first syntax element that is equal to 0 refers to or corresponds to the PH. Each picture associated with it contains one or more NAL units. and one or more NAL units have the same NAL unit type The second syntax element, equal to 1, specifies whether the picture is a GDR picture or an I A second syntax element that identifies a RAP picture and is equal to 0 indicates that the picture Identifies that the picture is neither an IRAP picture nor a GDR picture. The second syntax element may be ph_gdr_or_rap_pic_flag as previously described.
[0187] In some examples, the processor 620 may It determines that both syntax elements are received and that the first syntax element is equal to 1. In response to the By applying a constraint, a second syntax element in the PH is generated based on the first syntax element. The first syntax element equal to 1 refers to or is a PH. Each picture that corresponds to or is associated with two or more NAL units and no two or more NAL units have the same NAL unit type The second syntax element, equal to 0, specifies that the picture is a GDR picture. Identify the picture as neither an IRAP nor an IRAP picture.
[0188] In some examples, the second syntax element indicates whether the picture is a GDR picture. The first syntax element specifies whether a picture is to be streamed within the PPS associated with the picture. Additionally, processor 620 may process the first syntax element equal to 0. In response to determining that the second syntax element is signaled in the PH, and in response to determining that the first syntax element is equal to 1, The syntax element is not signaled in the PH, and the value of the second syntax element is 0 and associates it with a picture based on the first syntax element. Constrains the second syntax element signaled within the attached PH. Equal to 1 The first syntax element refers to, corresponds to, or is associated with a PPS. Each framed picture contains two or more NAL units and two or more NAL The first NAL unit is equal to 0 and specifies that the units do not have the same NAL unit type. The syntax elements in this document refer to, correspond to, or are associated with a PPS. Each picture contains one or more NAL units, and one or more Identifies NAL units with the same NAL unit type. The syntax element is used when the picture associated with the PH is either an IRAP picture or a GDR picture. Identify that it is not Cha.
[0189] In some examples, the processor 620 may determine whether the second syntax element is signaled in the PH. further determining a value of the first syntax element in response to determining that the first syntax element is not matched; In response to determining that the value of the syntax element is 0, the value of the second syntax element is 0, and in response to determining that the value of the first syntax element is 1, The value of the second syntax element is the value of the third syntax element signaled within the PH. The third syntax element indicates whether the picture is a GDR picture or an IRA picture. Identifies whether it is a P picture.
[0190] In some examples, the processor 620 may further include a valid PPS signaled. The value of the second syntax element in the PH is determined according to the value of the flag. It is used to identify whether a picture is valid as a GDR picture. Then, processor 620, in response to determining that the value of the valid flag is equal to 0, The value of the tax element is determined to be 0. In some instances, the valid flag is set to 0 as described above. It may be sps_gdr_enabled_flag as shown above.
[0191] In some examples, the second syntax element indicates whether the picture is a GDR picture. The first syntax element specifies whether a picture is to be streamed within the PPS associated with the picture. Further, the processor 620 performs the first syntax element and the PH. The P associated with the picture based on the third syntax element signaled in By constraining the second syntax element to be signaled in H, The first picture signaled in the PH associated with the picture based on the syntax element The third syntax element specifies whether the picture is a GDR picture. In some cases, the third syntax The flag element may be ph_gdr_or_irap_pic_flag as described above.
[0192] In some examples, the processor 620 further determines whether a third syntax element is equal to 1. and in response to determining that the first syntax element is equal to 0, The first syntax element equal to 0 refers to the PPS. Each picture that corresponds to or is associated with one or more NAs It contains an L unit and one or more NAL units have the same NAL unit type. The second syntax element, equal to 1, specifies that the PH has a type. The third syntax bit equal to 1 identifies the captured picture as a GDR picture. The _PICTURE_EXPRESS element specifies whether the picture is a GDR picture or an IRAP picture.
[0193] FIG. 9 is a flowchart illustrating an example process of video encoding according to some implementations of the present disclosure. In step 902, processor 620 receives syntax elements. The syntax elements are signaled within the PPS associated with the picture as described above. In step 904, the processor 6 20 performs the decoding process based on the values of the syntax elements.
[0194] In some cases, the syntax element for a picture is equal to 0 and the picture's If any slice has nal_unit_type equal to GDR_NUT, then the syntax of the GDR picture is The value of the slice element is equal to 0, and all other slices of the picture have the same value of nal_unit_type and recognizes the picture as a GDR picture after receiving the first slice of the picture. It is recognized.
[0195] In some examples, the processor 620 may select a syntax element whose value is equal to 0 and To determine if a slice of image contains a NAL unit type equal to GDR_NUT, Accordingly, all other slices of the picture contain the same NAL unit type, and A picture may be determined to be a GDR picture after the first slice of the picture is received. .
[0196] In some examples, the processor 620 is implemented in a decoder.
[0197] FIG. 10 illustrates an example process for video encoding according to some implementations of the present disclosure. In step 1002, the processor 620 selects a picture corresponding to the PPS. contains one or more NAL units, and one or more N The first NAL unit in a PPS that specifies whether the NAL units have the same NAL unit type Receive syntax element 1.
[0198] In step 1004, the processor 620 determines whether the picture corresponding to PH is an IRAP picture. The second syntax element in the PH specifies whether the picture is a GDR or GDR picture. Believe.
[0199] In step 1006, processor 620 determines the first syntax element based on the value of the first syntax element. Determine the value of the syntax element 2.
[0200] In some examples, the processor 620 may be implemented in a decoder.
[0201] In some examples, the first syntax element equal to 1 indicates that each picture corresponding to the PPS The data contains two or more VCL NAL units and two or more VCL NAL units. The first NAL unit is also equal to 0. The syntax element specifies that each picture corresponding to a PPS is represented by one or more VCL NAL units. unit and one or more VCL NAL units are the same NAL unit. Identify that the device has the same type.
[0202] In some examples, the second syntax element equal to 1 indicates that the picture corresponding to PH is Identifies an IRAP or GDR picture and has a second symbol equal to 0. The tax element is used whether the picture corresponding to the PH is an IRAP picture or a GDR picture. It has been determined that this is not the case.
[0203] In some examples, the processor 620 may be configured to determine whether the NAL types of the other slices are the same as the NAL types of the slices. to the NAL type of other slices in the picture to require that the NAL type be the same. applies the first constraint to the A second constraint may be applied to the second syntax element, such as:
[0204] In some examples, a non-transitory computer-readable storage medium for video encoding is provided. The computer-executable instructions on the non-transitory computer-readable storage medium may include one or more When executed by the computer processor 620, one or more computers The processor 620 performs the method illustrated in FIG.
[0205] In some examples, a non-transitory computer-readable storage medium for video encoding is provided. The computer-executable instructions on the non-transitory computer-readable storage medium may include one or more When executed by the computer processor 620, one or more computers The processor 620 performs the method illustrated in FIG.
[0206] In some examples, a non-transitory computer-readable storage medium for video encoding is provided. The computer-executable instructions on the non-transitory computer-readable storage medium may include one or more When executed by the computer processor 620, one or more computers The processor 620 performs the method illustrated in FIG.
[0207] In some examples, a non-transitory computer-readable storage medium for video encoding is provided. The computer-executable instructions on the non-transitory computer-readable storage medium may include one or more When executed by the computer processor 620, one or more computers The processor 620 performs the method illustrated in FIG.
[0208] The description of the present disclosure has been presented for purposes of illustration, but is not intended to be exhaustive or limiting of the disclosure. Several modifications, variations, and alternative implementations are possible in accordance with the above description and related These and other objects, advantages and modifications will be apparent to one skilled in the art having the benefit of the teachings presented in the accompanying drawings.
[0209] These examples illustrate the principles of the disclosure and allow those skilled in the art to understand the disclosure in various implementations and The underlying principles and various implementations with various modifications suited to the particular intended use The disclosure has been chosen and described in order to make the best possible use thereof. The present invention is not limited to the specific implementations disclosed, and modifications and other implementations may be made. are intended to be included within the scope of this disclosure.
Claims
1. 1. A method for video encoding, comprising: receiving, by a decoder, a first syntax element in a Picture Parameter Set (PPS) that specifies whether a picture corresponding to the PPS includes two or more Network Abstraction Layer (NAL) units and whether the two or more NAL units have the same NAL unit type; receiving, by the decoder, a second syntax element in a Picture Header (PH) that specifies whether a picture corresponding to the PH is an Intra Random Access Point (IRAP) picture or a Gradual Decoding Refresh (GDR) picture; determining, by the decoder, a value of the second syntax element to be 1 in response to determining, based on the value of the first syntax element in the PPS, that each picture corresponding to the PPS includes two or more VCL NAL units, the two or more VCL NAL units have the same NAL unit type, and the NAL unit type is a GDR NAL unit type or a specific IRAP NAL unit type; wherein the value of the first syntax element equal to 0 specifies that each picture corresponding to the PPS includes two or more VCL NAL units, and the two or more VCL NAL units have the same NAL type, and the value of the second syntax element equal to 1 specifies that the picture corresponding to the PH is an IRAP picture or a GDR picture; and determining, by the decoder, the value of the second syntax element to be 1
2. 2. The method of claim 1, wherein the value of the first syntax element equal to 1 specifies that each picture corresponding to the PPS includes two or more video coding layer (VCL) NAL units, and the two or more VCL NAL units do not have the same NAL unit type, and the value of the second syntax element equal to 0 specifies that the picture corresponding to the PH is not a GDR picture.
3. Determining the value of the second syntax element based on the first syntax element includes:
3. The method of claim 2, further comprising: determining, by the decoder, in response to determining, based on the value of the first syntax element in the PPS, that each picture corresponding to the PPS has two or more VCL NAL units, the two or more VCL NAL units do not have the same NAL unit type, and the NAL unit type is neither a GDR NAL unit type nor an IRAP NAL unit type, that the value of the second syntax element is 0.
4. The method of claim 3 , wherein the specific IRAP NAL unit type is IDR W RADL or CRA NUT.
5. 1. A method for video encoding, comprising: receiving, by a decoder, syntax elements; performing, by the decoder, a decoding process based on the values of the syntax elements; 1. A method according to claim 1, wherein the syntax element specifies whether Network Abstraction Layer (NAL) units of a picture have the same NAL unit type, and the value of the syntax element is 0 when the picture is a Gradual Decoding Refresh (GDR) picture or an Intra Random Access Point (IRAP) picture.
6. The method of claim 5 , wherein the syntax element is a syntax element pps_mixed_nalu_types_in_pic_flag signaled in a Picture Parameter Set (PPS) associated with the picture.
7. 6. The method of claim 5, wherein if the value of the syntax element for a picture is equal to 0 and any slice of the picture has a nal_unit_type equal to GDR_NUT, then the value of the syntax element for a gradual decoding refresh (GDR) picture is equal to 0, all other slices of the picture have the same value of nal_unit_type, and the picture is recognized as a GDR picture after reception of the first slice of the picture.
8. 6. The method of claim 5, wherein if the syntax element for a picture is equal to 0 and any slice of the picture has nal_unit_type equal to a particular IRAP NAL unit type, all other slices of the picture have nal_unit_type of the same value, and the picture is recognized as an IRAP picture after reception of the first slice of the picture.
9. The method of claim 8 , wherein the specific IRAP NAL unit type is IDR_W_RADL or CRA_NUT.
10. 1. An apparatus for video encoding, comprising: one or more processors; a memory configured to store instructions and a bitstream that are executed by the one or more processors; 10. An apparatus, wherein the one or more processors are configured, when the instructions are executed, to perform a method according to any of claims 1 to 9 using the bitstream.
11. 10. A non-transitory computer-readable storage medium for video encoding, storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform a method according to any one of claims 1 to 9.
12. A program stored on a computer-readable storage medium having instructions which, when executed by a processor, perform the method according to any one of claims 1 to 9.
13. A method for transmitting a bitstream, said bitstream being decoded by a method according to any of claims 1 to 9.